Method and system for processing distributed workloads - Patents.com

JP2024521582A5Pending Publication Date: 2025-06-17SAILION INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023576348
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-10
Filing Date
2022-06-09
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing systems face inefficiencies in distributing computational workloads across heterogeneous devices, including underutilization of resources and high costs associated with building and maintaining cloud computing data centers, while portable devices like smartphones have limited network access and battery life, complicating data processing.

Method used

A system and method for distributing data processing jobs to a diverse set of network-attached computing devices, including smartphones and tablets, using a central management platform that estimates job execution speeds based on device similarity and allocates data chunks to ensure timely completion, with mechanisms for reallocating workloads and utilizing cloud services when necessary.

Benefits of technology

This approach optimizes workload distribution, maximizes device utilization, and ensures timely job completion by selecting appropriate devices and reallocating tasks, even in cases of network or battery limitations, thereby reducing costs and improving efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method and system for distributing a computational model and data to be processed to heterogeneous distributed computing devices is disclosed. The computational model and some of the data are processed in a benchmark system and timing is used to estimate the job execution speed for each computing device. Based on the estimation, a computing device is selected and assigned data chunks to complete the distributed processing within a predefined period. Using separate processes, such as a payload manager configured to transfer computation jobs to remote devices and a messaging engine configured to transfer data messages, the computational model and data chunks are sent to the separate computing devices, and the payload manager and messaging engine communicate with corresponding software engines in the computing devices.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims priority to U.S. Provisional Patent Application No. 63 / 209,002, filed June 10, 2021, the entire contents of which are incorporated herein by reference.

[0002] FIELD OF THE DISCLOSURE This disclosure generally relates to computer systems and methods for distributing and decomposing computational jobs and workloads to remote and mobile devices and collecting job processing results from these devices. [Background technology]

[0003] Often, individuals or companies need to perform resource-intensive calculations to analyze large data sets, but do not have directly available computing resources, and therefore turn to third-party on-demand services, such as various commercial cloud computing platforms built in one or more data centers with closed networks of servers. Certain types of such calculations can process data in individual chunks rather than operating on the entire data set as a whole. Such workloads include, but are not limited to, deep learning classification and predictive models, time series, multivariate and probabilistic models, etc. These types of calculations can be broken down into smaller jobs that are chunked evenly across devices. To achieve this, forced uniformity across compute node devices and / or resulting in underutilization of computing assets, which results in very inefficient workload and device utilization. Building and maintaining such cloud computing data centers is expensive, requires significant investment in computer hardware, and has large power requirements.

[0004] It would be advantageous to provide a system and method that can distribute execution workloads in a heterogeneous distributed computing environment, and even more advantageous if such a system and method could operate reliably using portable mobile devices, such as smart phones and tablets, as computational resources that can have limited and inconsistent network access to receive job data and return job results, and the operation of any given device is constrained by available battery life and the need for a user of a given mobile device to use the device for other purposes. Summary of the Invention [Problem to be solved by the invention]

[0005] These and other problems are solved by a method and system that distributes data processing jobs to multiple networked computing devices that are remotely accessible and include heterogeneous devices that are unrelated to each other, under the operation of a central management platform, which may be a dedicated server or a cloud-based system, where each computing device has a device type, e.g., a particular type of smartphone or IoT device, has a distinct network address, and is characterized by various available device resources, such as processing and storage resources.

[0006] When a job payload containing an executable computation model and associated data is submitted to the management platform, a portion of the received data is used to apply the received computation model to a benchmark system to determine a benchmark job computation speed for the received payload. The benchmark system may be a single platform with known characteristics or may be multiple different types of devices reflecting the types of computing devices usable by the system.

[0007] The benchmark computation speeds are used to estimate job execution speeds for a plurality of computing devices, and the estimated execution speeds are then used to select a cluster of computing devices from the plurality of computing devices on which to execute the computational model. When a plurality of different types of benchmark devices are used, the estimation of the job execution speed for a given computing device can be made based on the benchmark data of the benchmark device that is deemed to be essentially closest to the individual computing device, such as indicated by the value of a device feature similarity vector.

[0008] The device is selected and one or more data chunks are assigned such that the total time for a given computing device to complete processing of all assigned data chunks is less than a predetermined time. For example, if an entire computing job needs to be computed within time t, each computing device is selected and assigned data chunks that will be completed within time t based on an estimated job execution rate. A buffer (e.g., 10%) may be included to account for possible delays.

[0009] The selection of computing devices and allocation of data chunks is done to maximize desired operating conditions. In one embodiment, the selection is made to avoid using faster computing devices than are necessary to complete the assigned work by the deadline. If sufficient computing devices are not available to complete all jobs within the performance constraints, a commercial cloud computing service may be used to create or launch one or more virtual computing devices.

[0010] After the data chunks are assigned to the various computing devices, the assigned computational model and data are sent to the individual computing devices in the cluster. For any given computing device, all assigned data chunks may be sent together for processing, or the computing device may be given no data chunks, or only one or some of the data chunks, with the remaining data chunks being just data chunk identifiers. As other data chunks are completed, new data chunks are obtained from the management platform as needed. The processed data results generated by the computing devices are then sent back to the management platform for use by the requesters.

[0011] In one embodiment, the computational model and the data chunks are transmitted via separate communication transport methods. For example, a first transport mechanism for transmitting the computational model payload may be in the form of edge computing software of the management platform, configured to deliver the computational payload to a corresponding payload processing engine of the computation device that extracts and executes the payload on the specified data. A second transport mechanism, such as a streamlined message broker service configured for general event-driven and inter-application messaging without additional overhead of a payload processing system, transmits the data chunks using the second transport mechanism. Corresponding messaging software operating on the computation device operates to make the data chunks received on the computation device available to the payload processing engine. The messaging service may also be used to send back the results of the computation.

[0012] After the computation model and data are allocated and transmitted to the computing devices, the operation of the computing devices is monitored to detect potential error conditions. In one embodiment, a completion deadline for a given computing device is determined based on the device's estimated job execution speed. If the deadline passes without a computation result being received, the allocated data chunk is reallocated by the management platform to another computing device (which may be selected from multiple computing devices and may be a device already in the cluster or a computing device not in the cluster), requiring the computation model and reallocated data chunk to be transmitted.

[0013] The operation of the computing devices is also monitored. Software in the computing devices operates to detect health or other operational problems and send health alerts to the management platform as needed. The management platform also periodically polls the computing devices for response messages indicating the health status of the individual computing devices. If a health alert message or polling response is received that indicates an alert or health status of some kind such that the computing device is unable or likely to be unable to complete processing of an assigned chunk of data within a time period, then one or more of the assigned chunks are reallocated to another computing device. If the computing device is a battery-powered device, a particular health alert indicates a low battery condition in which insufficient power is available for the computing device to complete assigned work.

[0014] The computing device itself can also monitor its own health, and if it detects a health condition or alert, such as low battery, low memory, or low CPU resources, in addition to sending an alert to the central management platform, the computing device will pre-offload some or all of its unprocessed data chunks to another computing device for processing. Such a feature is particularly useful in operating environments where the network link to the management platform is not continuously available and health issues arise well before the management platform is expected to receive the results.

[0015] The target computing device selected for transferring the workload or job results may be a device that is accessible via a local connection such as WiFi, Bluetooth, NFC, etc., even if an Internet connection to the management platform is unavailable. The management platform obtains information from either one or both computing devices regarding the offloading of unprocessed or processed workloads permitted by the available network connections. In such a case, the management platform decides whether to keep the reallocation or whether to reallocate some or all of the data chunks to a third computing device, if deemed necessary, for example, to meet a job completion deadline. Similarly, if the computing device is unable to return the data results due to network connectivity issues, it operates to transfer the results to another computing device that has a network connection to the management platform and is able to return the data. [Brief description of the drawings]

[0016] Further features and advantages of the present invention, as well as the structure and operation of various embodiments of the present invention, are described in detail below with reference to the accompanying drawings.

[0017] [Figure 1] FIG. 1 is a schematic diagram of a system used for distributed workload processing. [Diagram 2]FIG. 2 is a schematic diagram of the software structure of the management platform. [Diagram 3] FIG. 3 is a schematic hardware block diagram of a representative computing device. [Figure 4] FIG. 4 is a schematic diagram of the software structure of application software for a representative computing device. [Figure 5A] FIG. 5A shows a schematic flow chart of a user submitting and validating a job payload containing a model and data. [Figure 5B] FIG. 5B shows a schematic flow chart of a user submitting and validating a job payload containing a model and data. [Figure 6] FIG. 6 is a schematic flow chart of the computational job benchmark. [Figure 7] FIG. 7 is a schematic flow chart of a general process for job placement executable within the management platform. [Figure 8] FIG. 8 is a schematic flow chart of selecting a computing device and disaggregating data to multiple computing devices. [Figure 9] FIG. 9 is a schematic flow chart of the operation of the application software 140 of the computing device. [Figure 10] FIG. 10 is a schematic flow chart of health and resource monitoring in a computing device. [Figure 11] FIG. 11 is a schematic flow diagram of a health monitoring function executable in the management platform. [Figure 12A] FIG. 12A is a schematic flow chart illustrating an alarm monitoring process that can transfer data and / or assigned computing jobs from one computing device to another. [Figure 12B] FIG. 12B is a schematic flow chart illustrating an alarm monitoring process that can transfer data and / or assigned computing jobs from one computing device to another. [Figure 12C]FIG. 12C is a schematic flow chart illustrating an alarm monitoring process that can transfer data and / or assigned computing jobs from one computing device to another. [Figure 13] FIG. 13 is a schematic diagram showing communication and data flow between various elements of the system during cloud computing implementation. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0018] FIG. 1 is a schematic diagram of a system 100 used for distributed workload processing. The system 100 includes a computing management platform 110 that is connected to a network 105 and communicates with a number of computing devices 115, such as devices 115a, 115b, via the network 105. The computing devices 115 are used to execute a computing job. The computing job includes an executable computational model and data to be processed by the model. As described further herein, the management platform 110 selects one or more computing devices 115 to be used to execute the computing job. A decomposition process is used to break the workload data into data chunks consisting of one or more individual records, each of which is processed by the computational model. Based on the capabilities of the device, the data chunks are sized and assigned to the computing devices 115, and the workload data chunks may vary in size depending on the resources and capabilities of the various computing devices used. One or more data chunks are assigned to the devices for processing. The management platform 100 distributes the models and chunks of job data to the selected computing devices 115. Application software 140 running on the computing devices 115 operates to receive and apply the models to the received chunks of data and transmit the processed data results back to the management platform 110. The results of each computing device 115 execution are obtained and aggregated within the management platform 110, and the final results are transmitted back to the computation requester.

[0019] A customer or system administrator accesses the management platform 110 through a remote computing device 120. An access interface 125 can be used to submit computing jobs and data, retrieve results, control system settings, monitor the operation of the management platform 110 or the various computing devices 115, or for other purposes. The interface 125 provides a direct connection to the management platform 110 through the network 105, such as in the form of an Internet web server interface, or through a suitable API.

[0020] Network 105 may be a distributed network such as the Internet, a LAN, a WAN, or other open or closed network. Management platform 110 and computing devices 115 include conventional hardware and software such as hardwired Ethernet, or wireless Wi-Fi, Bluetooth, cellular, or other data links capable of establishing an appropriate network connection. Intermediate network devices (not shown), such as access points and routers, cell towers, and intermediate networks that connect to network 105, may be used to establish a connection to network 105. Although FIG. 1 shows only a single network 105, in various embodiments, one or more networks may be used for communication between devices, as described further herein.

[0021] The management platform 110 is implemented in a computer server platform that includes one or more computer systems having suitable processing power, network connectivity, and data storage capabilities to maintain a registry of computing devices, receive job models and data, distribute the models and data to various computing devices, and receive and aggregate results to be returned to the user. Suitable hardware platforms are known to those skilled in the art, and specific attributes, such as operating systems and access methods, may vary from implementation to implementation. In a particular embodiment, the management platform 110 is implemented in a cloud computing system using one or more virtualization technologies, such as a virtual computer or cloud computer with cloud storage. This may facilitate the implementation and expansion of the capabilities of the management platform 110. A hybrid approach may also be used in which certain functions are implemented in one individual computer and other functions are offloaded to a cloud computing platform.

[0022] A test platform 130 may also be connected to the management platform 110. As described further herein, the test platform 130 is used for model benchmarking operations.

[0023] 2 is a schematic diagram of the software structure of the management platform 110. As described herein, the administration (Admin) engine 205 provides the main functionality of the management platform 110. It is implemented using software, which itself can be divided into various applications and processes. The Admin engine 205 communicates with memory 210, which is a combination of various individual storage, which may include a mix of short-term (RAM) and long-term storage, such as local network drives as well as cloud storage. In addition to storing the operational software used by the Admin engine 205, the memory 210 is also used to store one or more models 211 to be executed, data 212 to be processed by the models, payloads 213 to be delivered with or without chunks of data to be processed with the models, and result data 214 generated by executing the models on the data chunks and sent back to the management platform 110. Memory 210 is also used to store device registration data 218 for computing devices 115 that are called upon for use in system 100, and user information 216 for users of system 100. Users of system 100 include customers who submit jobs and data to be processed, persons associated with registered computing devices 115, and administrators of system 100. Additional information 217, such as operational data and system logs, may also be stored.

[0024] The payload manager engine 215 communicates with the software 140 of the computing device 115 to send payloads to the computing device 115, the payloads including the computational model and one or more data chunks to be processed using the model. A separate messaging engine 220 is also used to communicate with the software 140 of the computing device 115. As described further herein, in one embodiment, the model is sent to the computing device using the payload manager, and the messaging engine 220 is used to send the data chunks to be processed by the model and to receive the results. The payload manager 215 and the messaging engine 220 are connected to the network 105 using the same lower network interface 225. They may also be connected to other networks, such as the network 105'. The testing platform 130 can be accessed through a network, such as the network 105, or through a direct connection. General access to the management platform 110 by a user or administrator can be through a web browser connecting to a web page interface 230.

[0025] In system 100, there are typically multiple individual computing devices 115, each of which has computing capacity that can be called upon by management platform 110 to execute computing jobs at least part of the time. System 100 operates with a range of homogeneous or heterogeneous computing devices, allowing a variety of different types of computing devices to be used simultaneously.

[0026] A particular computing device 115 may be a portable mobile device, such as a battery-powered smartphone or tablet or laptop computer, an IoT device, or other network-connected computing device. A computing device 115 may also be a fixed location device having an external power source, a computer such as a desktop PC, or a computer server operating in a standalone environment or available as part of a data center. Other types of computing devices 115 may also be used.

[0027] FIG. 3 illustrates a schematic hardware block diagram of a representative computing device 115. The representative computing device 115 includes a microcontroller 305 connected to a memory 310, which may include both short-term and long-term storage, such as RAM, ROM, solid-state drives, and the like. The memory 310 stores software executed by the microcontroller 305. The software may include program storage for an operating system, application software 140 used by the system 100, software for other applications, and data storage used by the software. One or more network interfaces are also provided for connecting to the network 105 and other networks. Available network connections include a wired network connection 315, such as an Ethernet connection, and one or more types of wireless data connections 320, including Wi-Fi, Bluetooth, Near Field Communication (NFC), and cellular data. The network connections 315, 320 are used to communicate with the management platform 110 over the network 105. The network connections are also used to identify and communicate with other computing devices 115. The computing device may be externally powered or powered by a battery 325. Depending on the type of computing device, it may include a visual display 330 and a user input device 335 such as a keyboard or touch screen. Various sensors 340 may be provided to monitor device conditions such as microcontroller temperature and battery level.

[0028] Computing devices 115 may be used for other purposes and may be mobile devices. Thus, for any given computing device 115, computing resources may be available continuously or intermittently, and when the resources are available, they may substantially reflect the full capacity or only a portion of the capacity of the computing device 115. The available capacity of a given computing device 115 may change over time.

[0029] 4 is a schematic diagram of the software structure of the application software 140. The payload services 410 operate to receive payloads from the management platform 110, such as those sent by the payload manager 215, execute computation jobs on selected sets of data, and store the results. The payload services 410 may include an edge service 411 that manages communication with the payload manager 215, and a payload processing engine 412 that has the functionality to execute the received models and data. The received models are stored in the model storage 420 in the memory 310 of the computing device 115. The sets of data (data chunks) applied to the models are stored in the data storage area 425 of the memory 310. The computation data from the models applied to the received data chunks are stored in the result data area 430 of the memory 310.

[0030] As described further below, in one embodiment, data chunks processed by the model can be provided to a supervisor service 405 from a messaging engine 220 of the management platform 110. The supervisor service 405 includes a messaging engine 406 in communication with the messaging engine 220 of the management platform 110, and a supervisor engine 407. The supervisor engine 407 has operational functionality to store received data chunks in a data storage area 425 in a manner known and available to the payload service 410. The payload service 410 applies the model to the data. Similarly, the result data is accessed by the supervisor service 405, and the result data is transmitted by the supervisor engine 407 back to the management platform 110 via the messaging engine 406. The supervisor engine 407 performs various other functions as described further herein.

[0031] The application software 140 operates on the computing device 115 in communication with the device operating system, through which it has access to various system functions, inputs and outputs, etc., including software and firmware for data communication and other functions. FIG. 4 also shows a health resource monitor 440, which signals system status, including CPU temperature, battery level, available memory, processor, and other system resource status, as appropriate, to the supervisor services 405 and payload services 410. Although the health resource monitor 440 is shown as a single module, various separate modules may be used. The monitoring may be part of the application software 140 or may be implemented in the computing device hardware, firmware, and / or O / S 435 or other software. The health resource monitor 440 issues system queries to the O / S to obtain system status data and monitors for system interruptions. System interruptions signal critical issues, such as low battery warnings, excessive CPU temperature, errors due to low memory, etc.

[0032] Details about resources and capacity available from a given computing device 115 may be collected during initial registration of the computing device 115 with the management platform 110. During initial registration, information about the computing device 115 is collected and stored in memory 210, such as in device registration database 218. End-user applications 140 are also installed on the computing device during registration to interact with the system 100 to receive and execute assigned computing jobs and return results, as disclosed herein. The overall operation of the applications 140 may be the same for different types of computing devices, although specific implementation details may differ. The application software 140 may be installed in stages, with the initial installation verifying that the node device meets the minimum requirements of the hardware and software platform operating requirements.

[0033] Each computing device 115 has an associated unique IP or other network address that can be used to communicate with the computing device 115. Also, an individual device ID can be assigned during registration. Depending on the device capabilities, one or more communication addresses are available and these individual addresses are stored. The management platform 110 selects among the available addresses depending on the network availability and the type of communication. For example, it is expected that many communications will take place over the network 105, which may be the Internet. However, if the main network connection is unavailable but communication is still desired, some messages may be communicated over a cellular network using an individual protocol such as SMS messaging.

[0034] Due to the wide distribution and diversity of these computing resources, across networks of differing capabilities and persistence, various computing metadata attributes of each device may also be collected during the registration period. The metadata identifies the individual capabilities of the registered computing devices. Device attributes reflecting the device's capacity may include one or more of the following: CPU type and speed, RAM and software storage space, network latency, operating system, whether the system is or may be battery powered, and the type of network connection available. Other metadata may include limitations on the number of days / hours that a registered device is or is not available for use by the system 100. Some or all of such metadata may be used by the management platform 110 to identify available computing devices 115 to which a given computing job is sent for execution and to determine how to chunk and distribute data to the selected devices.

[0035] The available capacity of a computing device may change over time. For example, a given device may offer 100% of its computing power and memory outside of normal working hours, but only 25% during the weekdays from 9am to 5pm. Details about expected availability may also be collected during the registration process. Availability patterns may also be monitored over time to develop a device profile. The available capacity of a given device may also change dynamically. A device may have 100% capacity at time t while idling, but if the device is actively used for other reasons, only a portion of the capacity is offered to the current system. Thus, the capabilities of a given device, either initially set in the system registration or learned over time, are checked by the management platform 110 before assigning any computing job to the device. The computing device 115 may also periodically send messages to the management platform 110 reporting its available resources, or the management platform 110 may periodically query the computing device to obtain the information.

[0036] During general operation of the system 100, the management system 110 monitors the overall state of registered computing devices and maintains an internal record of the current computing device state. If the computing device is operating normally (which may be indicated by messages sent by the computing device, for example, periodically or in response to a system query), the device state is set to active. If the node is not operating at full capacity, for example, only a portion of the scheduled resources are available, but the computing node has enough resources remaining for use in at least some situations, the computing node state can be set to a degraded level. This degraded level can then be used when ranking and selecting nodes to receive workload packages. If the device is not operating or the available resources do not meet minimum requirements, the node device state can be set to suspended.

[0037] In some cases, it may be necessary to temporarily or permanently remove a computing device from the available inventory. If the computing device is currently running a job, it is instructed to send a message to the management platform 110 system when the job is complete, or the computing device can be queried periodically. Also, if a notification is received that the device is not in use, the device is removed from the inventory.

[0038] 5A and 5B show a schematic flow chart of a user submitting and validating a job payload including models and data. In an initial step, a user connects to the management platform 110 via, for example, the web interface 230 and signs in. Initial checks are performed to authenticate the user and check that the user is authorized to submit a job (steps 502-512). If the user has the appropriate permissions, in response to selecting the payload upload option, the user is prompted to provide information about the payload, such as the general type of processing job, the amount of data, the expected time / date of completion, etc. Checks are performed to ensure that the input data is correct and does not exceed user permissions. For example, a maximum job size may be set (steps 514-518). If there are no blocking issues, the customer uploads the payload using conventional mechanisms, such as tagging a job model file and one or more job data files (step 522).

[0039] After uploading the user's payload, an initial security check is performed: a virus scan is performed, and if the payload looks suspicious, the payload is quarantined and the problem is reported to the customer (steps 524-529). Also, checksums or similar validations are performed on the model and data (alone or in combination) to detect data errors by comparing the data with check data provided by the user or embedded in the uploaded payload. If an error is detected, the payload is rejected and the user is notified of the error (steps 530-536).

[0040] Different types of model analyses have different computational complexities. To aid in the selection of a set of computing devices that can execute the job in the target time, the job submitter may be required to identify the type of analysis the workload model is performing, for example, by selecting from a set of predefined analysis types. The type of analysis is used to estimate the speed at which a given processor will process a given amount of typical data, which in turn estimates the time it will take a given computing device 115 to process a chunk of data.

[0041] As part of the upload process, a payload model type check is performed to ensure that the model type or algorithm coded in the payload matches the type of model or algorithm specified by the customer (step 538). Pattern matching is used in the code to identify common types of routines used for different analysis types, as well as other metadata. For example, a simple classification model made with TensorFlow Lite will have different performance characteristics compared to an object detection model made with TensorFlow. Any model mismatches are reported to the customer (steps 540-542). AI-based pattern matching model type detection is used in conjunction with AI models learned over time as the system 100 operates.

[0042] In addition to, or instead of, obtaining job type information from the submitter, the management platform 110 may perform a benchmarking operation to determine a computation speed measurement for a given job (step 544). Benchmarking also advantageously provides validation that the computation model provided is functional and can operate on the data provided. When multiple different devices are used in the benchmarking, the benchmarking identifies platforms that do not fit the computation model and should not be selected for use with the job, even if they are otherwise available.

[0043] In one embodiment, the submitted computational model, using a portion of the actual data to be processed for the job, is executed in the test platform 130 on one or more computers with known computational characteristics to perform the benchmark. Referring to Figure 6, after receiving the submitted model and the data to be applied to it (step 602), one or more benchmark machines are selected (step 604) and representative data is extracted from the submitted data set to be applied to the submitted model (step 606). Benchmark tests using the actual model and representative data for the job are then sent to the benchmark machines for execution (steps 608, 610).

[0044] The management platform 110 is connected to the computers in the test platform 130 using a link separate from the network 105, such as a LAN or other direct or local connection. Alternatively, the computers in the test platform 130 may be accessible via the network 105. For purposes of sending and receiving benchmark results, the computers in the test platform 130 are treated by the management platform 110 in the same manner as the computing devices 115. If each benchmark machine runs on the same model and data set, a single common payload is prepared and sent to each benchmark machine, and the process runs substantially in parallel on each benchmark machine.

[0045] While a single test computer with known operating characteristics can be used to obtain benchmark measurements of processing speed, in one embodiment, multiple test computers are used that are representative of typical types of computing devices 115 that have been registered or will be registered. For example, the multiple test computers may include different brands and models of smartphones, such as several different iPhone® models running iOS, Android® based smartphones, and desktop computers with common CPU configurations running the Windows® operating system. The test computers 130 each run appropriate application software to communicate with the management platform 110 to receive and execute jobs and return results.

[0046] The size of the sample portion of the computational data used for the benchmark may be comparable to the size of the data chunks used during the actual processing of the job, but is usually assumed to be much smaller. This is because the objective of the benchmark is not to process a large amount of data, but to measure the computation speed of the model of interest using representative data. Nevertheless, if the processing speed depends on the value of the input data, choosing a sufficiently large sample will average out the difference in processing time due to data variability. For example, if the model is for an image processing algorithm, it is necessary to process several pictures of different complexity to make the benchmark more accurate, since the more detailed the source image is, the longer it may take to compute.

[0047] In one embodiment, data output from the model during benchmarking can be discarded. Alternatively, and especially if the volume of data used for benchmarking is large, the data can be stored and only the remaining source data can be chunked and distributed across computing devices for processing.

[0048] The time required to process the data in the benchmark is then used to determine the processing speed of the submitted model and data, step 612. The benchmark data is used to estimate the time it will take a given computing device 115 to process a chunk of data, and this information is used to select the computing device to be used for the job and to allocate the data chunks to the computing device.

[0049] More specifically, given benchmark data from one or more test machines, the time it takes a given computing device 115 to apply a submitted model to data chunks of the submitted data is estimated based on the benchmark times of the test machines. If multiple test devices are available, the benchmark time of a computer in the test platform 130 that is deemed to be closest to the expected operational characteristics of the computing device 115 may be used. The proximity of a computing device 115 to a particular test device is measured by determining the magnitude of a similarity vector having numeric dimensions based on the proximity of matches with device types or brands, given models or versions of types, operating systems and operating system versions, processor configurations, and other factors.

[0050] A scale factor may also be applied to the benchmark processing speed to reflect differences in computing power between the computing device 115 and the test platform device(s). Differences in power that affect processing time include CPU type and speed, availability of co-processing devices, operating system, amount and speed of available memory and storage. Information for scaling the benchmark time of a given computing device is obtained during the initial registration of the computing device with the management platform 110. The device characteristics are used to estimate the scale factor. Additionally or alternatively, an initial benchmark test is run on a new computing device upon registration or at other times. The same or equivalent benchmark test is run against the device in the test platform, and the results are used to provide a relative performance scale factor of the particular computing device 115 compared to one or more benchmark test devices. The initial benchmark test may be independent of any particular submitted computing job. If multiple test devices are available, the benchmark time of a computer in the test platform 130 that is deemed to be closest to the expected operating characteristics of the computing device 115 may be used to estimate the performance of the computing device 115. The proximity of a computing device 115 to a particular test device may be measured by device type or brand, operating system, processor architecture, and the like.

[0051] FIG. 7 is a schematic flow chart of a general process of job placement executable within the management platform 110, such as the admin engine 205 and the payload manager 215. When a computation job is submitted for execution, the system compiles a list of available devices and then identifies a set of devices to serve the aggregated computation request to meet predefined criteria. For example, selected to use a minimum number of devices and have a projected completion time that meets the job requirements. In one embodiment, first, job information is obtained (step 702). The minimum hardware requirements for the computation device to be used for the job are determined (step 704). The minimum hardware requirements are determined by the type of model to be executed, the model size, and the available storage space. For example, a computation device may have computation capacity that is suitable for some job types but not others. The available storage space should be sufficient to store the computation model, at least one data chunk to be applied to the model, and the results of the computation (see, for example, memory 310, 420, 425, 430 in FIG. 4). Available storage space is used as part of job placement to determine how much data can be allocated to a given computing device, how many chunks of data can be delivered in one batch, or how chunks of data can be streamed over time.

[0052] A maximum execution time may be calculated (step 706) and used for reference. The maximum execution time is an indication of the amount of time it takes to execute a computational workload across the entire data set. The execution time is calculated using the average data processing speed across the collection of computing devices. Alternatively, the maximum execution time may be viewed as the maximum time a job can be completed by a specified deadline.

[0053] A check is performed to determine which registered computing devices 115 are available (step 708). A computing device 115 may be considered available if it is available to receive payloads from the management platform 110. Availability criteria include factors such as device connectivity and device health measurements, which may consist of device storage state, CPU temperature, and charge state for battery powered devices. In a further embodiment, the expected future availability of a computing device is considered when the computing device is assumed to be available before the management platform 110 is scheduled to send new jobs and data to the computing device.

[0054] An initial allocation of the job to available computing devices can be made to ensure that there are enough available computing devices to complete the job within the allocated time period. The estimated data processing rates of the available computing devices, determined in the benchmarking process, are used to estimate the job completion time under various scenarios of device selection and data chunk allocation. If there are not enough computing devices (step 710), an insufficient capacity signal is generated (step 712). The management platform 110 signals the problem to the submitting user and adjusts the target completion deadline to match the total available capacity (step 716).

[0055] Alternatively or additionally, if the registered computing devices available for use by the management platform 110 do not have sufficient capacity to complete the total workload by the target completion time, a virtual computing device is created or launched (step 714) to provide the necessary computing resources to achieve the target. For example, the management platform 110 is configured to create a new virtual device in a conventional commercial cloud computing service, e.g., Amazon® Web Services, or launch an existing, possibly dormant, virtual device pre-configured in the service. The virtual device is then added to the set of available computing devices to be used for the job, and data chunks are allocated and distributed accordingly with the model. Since third-party cloud services typically charge a fee for the use of virtual devices, the use of virtual machines to replace or augment registered computing devices 115 is limited to cases where registered computing devices are insufficient to meet the computing requirements. The protocol for communicating with the virtual machine to transmit the computational model and the allocated data chunks needs to be done using a different protocol than that of the computing device 115 using the application software 140 for the system 100. Instead, the version of the application software 140 used on the computing device can be loaded onto the virtual machine, so that the virtual machine is treated by the management platform 110 in the same way as a registered computing device 115 .

[0056] After the initial device selection, the device availability and resources of the device are verified and the device is assigned to a cluster. In one embodiment, once the available resources from a given computing device 115 are verified and the device is selected, the application software 140 is instructed to reserve these resources for a particular period of time as far as possible given the operating system of the given computing device, etc., and other factors, so that the management platform 110 can assign and distribute jobs depending on the specified and verified resources. Device verification and resource reservation are performed at various stages during the selection of the computing device.

[0057] The available computing devices are assigned to the cluster and set to a reserved state (steps 718, 720). In one embodiment, available resources from a given computing device 115 are identified and once the device is selected, the application software 140 is instructed to reserve these resources for a particular period of time as far as possible given the operating system of the given computing device, etc., and other factors, so that the management platform 110 can assign and distribute jobs depending on the specified and identified resources. Device validation and resource reservation are performed at various stages during the selection period of the computing device.

[0058] Based on the requirements of the computation job, the management platform 110 system uses the available metadata of the computing devices to select a set of computing devices 115 from the list of available devices and assigns the data chunk or chunks to be processed to each computing device in a decomposition process (step 722). To increase security when sequential processing of data is not required, the data records can be randomized and then assigned to the devices sequentially (if the chunk size is a single record) or sequentially to the chunks assigned to the devices. A non-random stripe selection of data records may also be used. Distributing the data records in this way can reduce the value of hacking into individual devices by eliminating or reducing the information collected by consecutive data records. By storing the order of the shuffled records, the computation results are arranged to match the record order originally provided.

[0059] Selection factors may include available RAM, other storage space, network latency, computing platform hardware and software. Different computational models may require specific computing environments, so available devices are filtered to eliminate devices that do not have the capability to run the model being deployed. Also, estimates of job execution speeds based on benchmark data or the like may be used.

[0060] Selecting which computing devices 115 to use for a job and how to allocate data chunks between them prioritizes against one or more factors. The process can also select the devices and the amount of data provided for execution so that the workload assigned to a given computing device 115 uses at least a certain amount of maximum storage utilization, for example within 10-15% of the maximum storage utilization when taking into account the model size, the size and quantity of data chunks being sent, and the memory required to store the data processing results before they can be offloaded back to the management platform 110.

[0061] Given an estimate of the execution time speeds of the models on the available computing devices, a selection is made that provides an estimated job completion time that meets the target completion time requirement, selecting computing devices to meet the completion deadline while avoiding the use of faster computing devices than necessary.

[0062] In selecting the device, the network delay between the aggregation point and the computing device is further considered. Long delays can cause computation interruptions, data distribution interruptions, and the need to resend data to overcome potential losses. This can affect the estimated completion time. If the target job is completed at or before a set time or meets other criteria, a statistical estimate of the impact from the network delay can be included in the selection of the device.

[0063] Additionally, the location of a computing device may be an additional consideration in device selection. As part of validating the availability of a computing device, the device location may be queried and the current location recorded. Job submitters may also restrict where job data can be sent. For example, certain confidential information may be subject to export controls. Restrictions on the location of a computing device may also be used to filter and select available computing devices.

[0064] After selecting the set of computing devices 115 to be used for the job and allocating the data chunks, a payload is prepared including the model and the data chunks allocated to the selected computing devices 115 for execution (step 724).

[0065] The distribution of payload data may be in the form of transmitting a complete and cohesive payload over a network connection. The entire set of computational data to be executed by a given computing device 115 may be transmitted in advance, and all data of the workload is distributed during distribution of the initial job among the devices specified during payload distribution. This has the advantage that the initial payload distribution can be performed for a complete computing job. Alternatively, particularly if a given computing device 115 has limited storage, the device may be configured to receive or fetch multiple payload chunks for a given computing workload as needed, with the chunks being retrieved sequentially. For example, a device with storage available for one payload chunk may be assigned three payload chunks and have only the initial chunk provided for processing. The device automatically requests the next subsequent chunk when the existing chunk is completed, or when the management platform repeal the existing chunk in response to an indication that processing of the data chunk is completed, e.g., after receiving the processed data for the completed chunk. Other methods may also be used. Models and specific distribution methods for distributing data are further described below.

[0066] Before sending the model and data chunks to the computing device, a connection is made to the selected computing device 115 to check for readiness (step 726). Readiness is determined by verifying that there is sufficient memory and other resources available to run the computational workload. For battery-powered devices, a determination of whether there is sufficient battery power to run the computational workload may be used. If the device is not ready, the device is rejected and another device is requested (steps 728, 730). Replacing the device requires at least partial reallocation to be performed, and each computing device 115 is instructed to reserve resources that will currently be used for the next job if no resource reservation has been made. If the device is ready, the model and data payload are delivered (step 732). If there are repeated errors preventing delivery, the computing device is rejected and replaced (step 734).

[0067] In one embodiment, different transport mechanisms are used to transmit the computational model and the assigned data chunks from the management platform 110 to the selected computing device 115. Each transport mechanism has a corresponding software engine or agent in the management platform 110 in the computing device 115 to facilitate the transmission. The first transport mechanism may be in the form of traditional edge computing software designed to deliver the payload to the edge device and run within the edge device to extract and execute the payload, e.g., payload manager 215 in the management platform 110, payload service 410 in the computing device. The second transport mechanism may be a streamlined message broker service, e.g., messaging engine 220 in the management platform 110 and messaging engine 406 in the computing device, used for data chunk transfer and more general event-driven and other messaging between the computing device and the management platform. Although distinct transport mechanisms are used, they differ primarily at the application layer and possibly the protocol layer, and use the same lower communication layers to send and receive data to and from designated network addresses over the network 105.

[0068] A suitable platform for the first forwarding mechanism is the Open Horizon (OH) platform. A suitable message broker service for the second forwarding mechanism is Apache Kafka. The OH management hub in the management platform 110 is used as a payload manager 215 to communicate with an OH edge agent acting as an edge service 411 installed on the compute node 115. The OH edge agent 411 in the compute node 115 provides the computational model to the OH edge service, which acts as a payload processing engine 412 in the compute node and executes the model on the provided data. The OH platform can include the data in the distribution package, but since the data chunks to apply to the model are different for each computing device, each computing device needs a separate package and each package has its own copy of the computational model. As the number of computing devices used for a given job increases, the rate at which system processor and memory resources are consumed increases. Also, the distribution process becomes cumbersome to reallocate data chunks from one device while continuing the same overall job.

[0069] In this embodiment, a single package is generated and stored that includes the computational model but does not include the data. A first transfer mechanism, such as OH, is then used to send the package to each of the selected computing devices. A second transfer mechanism is used to transfer the data chunks assigned to each of the computing devices to each of the computing devices. Additional functionality of the application software 140 in the computing devices receives the data chunks and stores them in memory so that the edge service can apply the data to the received computational model. Similarly, the second transfer mechanism can be used to transmit computational results from the computing devices back to the management platform.

[0070] An embodiment of selecting a computing device and disaggregating data to multiple computing devices will be described in more detail with reference to the flowchart of FIG. 8. First, a list of available devices is created (step 802). Then, the list is sorted based on the desired optimization. For storage-optimized disaggregation, the list of devices is ordered based on the size of the available data storage (step 804), and the storage available for payload applications is determined from the list (step 806). To ensure there is enough storage, only a percentage of the actual available storage is considered, e.g., 80%, so if a device has 1 GB of actual available storage, only 0.8 GB is considered for payload allocation.

[0071] Other constraints such as working RAM and CPU type, speed, or estimated data processing speed based on benchmark data of the current job are used to select the device. For example, the devices are ranked based on their suitability for running the model, taking into account whether they have a GPU or only a CPU and whether this difference affects the execution of the job, whether the execution of the model will cause the processor demand to exceed a certain threshold that can be set during the registration period of the computing device or at other times, the percentage of the total data payload that can be allocated to the device and does not exceed the device threshold of a set remaining capacity, and network latency. A data chunk size is also selected (step 808). The chunk size is determined based on the initial job information provided by the user, the size of the file or data segment uploaded with the job, or by other means.

[0072] A computing device is selected from the sorted list in order, and chunks of input data are assigned to the device and added to the payload sent to the selected device (steps 810-814). Devices are considered in order until all data chunks are assigned (steps 816-818). The assignment takes into account the estimated benchmark data processing speed for each computing device, with a given buffer, e.g., 10% or 15% buffer, to complete the final assignment before the target deadline. When optimizing storage, the maximum number of data chunks that can be assigned to the selected device in one delivery is selected. For example, when sending multiple chunks of data initially, the maximum number of chunks that can be stored and have room to store the output data is selected. Data chunks are assigned to avoid splitting records. For example, if the record length is 1K (1024 bytes) and 5000 bytes are available, only 4096 bytes of data are assigned. Also, the device selection is done to use the slowest devices first, leaving faster devices available for other jobs for which the slower devices are not suitable. It should be noted that this process may result in some of the selected computing devices not being used. In such a case, the computing devices can be freed up for use by subsequent jobs. If all selected devices have been allocated data and additional data is still available, the remaining data is allocated to containers applicable to virtual devices and to activated devices, if not already done so (step 820).

[0073] During the allocation process, instead of directly allocating data to create a payload, an initial decomposition process determines the number of data chunks to be allocated to a given device. The actual data chunks are then allocated after all allocations are completed. The actual data chunks are allocated to be processed by a given device. The allocated data chunks may be sent in batches. When streaming data chunks over time, the computing device is informed of the number of data chunks allocated and the allocation of a particular chunk when the next chunk needs to be sent. This simplifies reallocating computing devices that may be required.

[0074] In an example of sizing an individual device, a computing device may have 2 GB of native storage and 1 GB of starting disk space. The available space is reduced by buffers for data output and reservations. A 20% buffer results in 800 MB of disk space. Assuming model and algorithm storage requirements of 5 MB, 795 MB of storage is available for input data.

[0075] In an example of device selection, the computational workload includes a 256MB model and a total data payload of 1GB. The system determines the supportability of the workload based on the specific payload requirements, which may include both a minimum size requirement of the computational model or algorithm and the available data storage availability of the device. Portions of the data and the model are distributed from the available devices to optimize the desired parameters. Devices that do not have enough storage to support the model are not selected. Similarly, devices that do not have enough memory to store the minimum size data chunks and the results of the data processing are not selected. Data chunks are distributed from the remaining devices based on the available storage.

[0076] A table of devices showing resource availability and allocation of payload chunks [Table 1] In the above table, devices that do not have the least amount of available storage are not selected. The list has been filtered to identify node devices with the amount of RAM that is considered optimal for a given workload.

[0077] 9 is a schematic flow chart of the operation of the application software 140 of the computing device 115. As described above, the computing device 115 receives a payload including a computational model (step 902). The computational model is stored (step 904), for example, in the storage 420. It also receives data chunks to be applied to the computational model (step 906) and stored (step 908), for example, in the storage 425. Depending on the execution, the data chunks may be received in an initial batch, as part of a payload sent to the payload service 410, or via a separate messaging service 406. The application software 140 also determines that additional data chunks are required and issues a request to the management platform 110, for example, via the messaging service 406. The payload processing engine 412 applies the data chunks to the model and stores the results, for example, in the storage 430 (step 910). The results of the computation are then sent back to the management platform 110 (step 912).

[0078] Results may be returned intermittently, for example, as processing of a single data set is completed, or results from multiple processed data chunks may be accumulated and stored on the computing device and returned as a batch. Results are returned in result packages that include a computational workload identifier that the management platform uses to associate the results with known computational workloads. For security, identification may be done at the result set level, but anonymized with respect to the data owner. Data chunks and processing results may also be encrypted for more secure transmission.

[0079] Each data chunk can have its own unique identifier, which is referenced in the data processing results returned for that chunk, and the identifier is used to place the set of results from a job in the proper order. In the case of non-sequenced data, the computational workload identifier can still associate the returned or generated results with a known computational workload request.

[0080] The application software 140 in the computing device 115 may also monitor device health to detect conditions that may affect the speed at which an assigned workload completes or the ability to complete completely. Referring to FIG. 10, the health resource monitoring module 440 in the application software 140 periodically polls or monitors system conditions while processing a workload (step 1002). The health monitoring module relies on native health monitoring capabilities built into the computing device's operating system, firmware, and / or hardware, and responds to relevant interrupts triggered by the native capabilities. Conditions that may be monitored or signaled by interrupts include CPU temperature, low memory or available storage, low battery charge. If a health condition is detected (step 1004), the application software 140 sends a health alert message regarding the condition to the management platform 110 (step 1006). In response, the management platform 110 reevaluates the estimated completion times of data chunks assigned to the computing device, taking into account the reduced resources, and determines whether it can maintain the allocation of data chunks to the device without missing target completion deadlines or other performance indicators. Also, if the execution speed, available memory, or another required resource of the computing device 115 becomes insufficient, some or all of the data chunks assigned to the problematic computing device can be reallocated to a different computing device that is already selected to execute a given computing job, or to another computing device not yet used for this job. If the need arises to reallocate a workload, the workload may be treated as the workload of a new job for the purposes of identifying the device to allocate its portions to. The selection of a device to reallocate a workload takes precedence over the selection of a device to allocate to a new job. If a registered computing device is not available, a virtual device can be launched.

[0081] Also, the portion of the workload allocated to a given device may need to be reallocated if the device goes offline or is not performing at an expected level even without receiving health alerts. For example, a computing device may lose network connectivity for an extended period of time, or the device may have software or hardware issues that cause it to perform poorly but not yet report it. If a given computing device does not regain network connectivity or return a result payload within the expected completion deadline, the device is considered unavailable to complete a job request. If data is being delivered to the device via a streaming protocol, the device is considered unavailable to complete a job request after retrying the attempted payload delivery several times.

[0082] Given that many types of edge nodes maintain transient network connections, system 100 continues to keep devices in a disconnected state for a period of time. The tolerable disconnection time may depend on the expected time it takes for a computing device to execute a workload. This estimate is based on the time when one or more data chunks are provided to a given computing device, the size of the data chunks, and the benchmark-derived processing speed of the computing device. Additionally, the offline tolerable time may be included in the data payload delivery methodology that is known prior to placement of the payload.

[0083] FIG. 11 is a schematic flow diagram of a health monitoring function that can be performed periodically or continuously or aperiodically as needed in the management platform. The predicted data processing rate of the computing device 115 is used to determine when results should be received from the computing device (step 1102). If the execution time has not elapsed (step 1104), the next device is optionally checked (step 1106). The availability and status of the device is then checked (step 1108), which may include querying the computing device to return health information and / or checking recently received data (step 1108). If the computing device is unavailable for more than a maximum allowed time (step 1110), the device is deemed to have failed (step 1112). A new device is requested, unprocessed data chunks are assigned to the replacement device, the new device is added to the cluster, and the payload is placed (step 1114). If the device is available, health is checked (step 1116). If health is not compromised (step 1118), the next computing device is checked. If the health is not OK, the impact is considered, including whether restarting the computational workload is possible (step 1120). If restarting is not possible, or if the health state is such that the allocated data chunk cannot be completed in full or within the required time, some or all of the data chunk is reallocated to a different node (step 1114).

[0084] In some cases, a computing device that is being used to process a workload needs to reacquire resources that were allocated to the workload to be used for other purposes. For example, if the computing device is a laptop computer, a user of the laptop computer needs to run a resource-intensive program. In one embodiment, this occurs with an end-user application.

[0085] Users of external devices receive monetary or other compensation when they register the management platform 110, install application software 140 that enables their device to be used as a computing device 115, and use the device to perform jobs according to the application software 140. The measure of compensation may be based on the amount of data a given computing device processes.

[0086] A user of the computing device 115 to which the workload is assigned may need to use the computing device for other purposes at the time the workload is scheduled to be executed. This reduces the resources available for the execution of the workload, slowing down the data processing rate of the workload. The application software 140 used as part of the system 100 in the computing device 115 monitors system activity to detect when such resource reduction occurs and whether the resource reduction is greater than a predetermined threshold. A measure of the impact of a task may be based on whether the available CPU cycles are reduced by a predetermined amount. The impact may also be due to a reduction in memory available to store the computation results of the workload, for example, insufficient memory available to store the results of processing a relevant number of data chunks, such as one chunk, the total number of chunks stored on the computing device for processing, or the total number of chunks assigned to the computing device if not all chunks have been delivered. It may also be considered whether the resource reduction is temporary or expected to be temporary.

[0087] If the resource reduction exceeds a predetermined threshold for a predetermined period of time, it is treated as akin to a health alert and a message is sent from the computing device 115 to the management platform 110 indicating the resource change (e.g., substantial CPU slowdown, memory shortage, etc.). In response, the management platform 110 evaluates the change and determines whether it is of sufficient impact to warrant a request to reallocate data chunks from the signaling device to a different computing device, and if so, takes appropriate action in a manner similar to the handling of a health alert.

[0088] The application software 140 on the computing device 115 also monitors other activity on the computing device 115, such as disconnection activity with the system 100, to determine whether an actual or potential resource reduction is due to a user action, such as the launch of a new application or task, rather than an activity required by the operating system. If the reduction is due to a user action, and in particular if such reduction affects the operation of an assigned computing job beyond a threshold indicating that the reduction is too great, a message is output on the display of the computing device 115 to inform the user that the operation must use resources allocated to active workloads to continue. In one embodiment, the application software 140 operates to temporarily suspend or reduce the priority of the new resource-consuming task until the user affirmatively indicates that the task should continue. If the user continues the resource-consuming task, a monetary or other penalty is assessed against any compensation the user would have received to make the particular computing device available.

[0089] It should be noted that when the computing device 115 completes processing of the workload, the computation results are sent back to the management platform 110. The application software 140 sends the results directly to the management platform 110 via the network 105. However, a direct network connection may not be available. Referring to FIG. 1, the device 115a is connected to the network 105 via a data link 106, which may be, for example, a cellular data Internet connection or a Wi-Fi connection to a node that has Internet access. Sometimes, the network connection 106 may not be available to the device 115a. When a connection to the network 105 is not available and a computing device such as the device 115a has completed a workload and has computation data to send back to the management platform 110, the device 115a establishes a connection with another peripheral device such as the computing device 115b using a direct Wi-Fi connection, or a connection using Bluetooth, Bluetooth Low Energy (BLE), or an alternative data connection such as an NFC protocol. The completed computation results are transferred from computing device 115a to computing device 115b, which then transfers the results from device 115a to management platform 110.

[0090] To aid in establishing such a connection, application software 140 on the computing device may cause the device to broadcast its presence to be discovered and recognized by other computing devices within communication range. For example, computing device 115b, as a BLE follower device, periodically sends a BLE advertisement message indicating its availability for a connection. The advertisement message may also include payload data such as a local name indicating that computing device 115b is running application software 140. To establish a data connection between two devices, the follower device sends an advertisement message indicating that it is available for a connection. The central device scans for advertisement packets. If it detects an advertisement packet from a suitable follower device, the leader device sends a connection request packet to the follower device and establishes a data connection if the follower responds appropriately.

[0091] NFC data connection is fairly fast but only provides a data link over a short distance, such as about 10 cm. The option of connecting via an NFC link allows an individual to place one computing device 115 where a transfer of computation result data or computation job is required in close proximity to a second computing device, and the application software 140 in the device performs the transfer. A scenario where this is useful is when a job is being performed on a mobile device that has gone into a low battery level state. A low battery computing device is placed adjacent to a second computing device that is sufficiently powered to transfer data and / or computation jobs to the second computing device.

[0092] 12A-12C are flow charts illustrating an alarm monitoring process that results in the transfer of data and / or assigned computing jobs from one computing device to another. The health resource monitor 440 in the device 115a monitors status and / or device generated alarms and notifies the supervisor service 405 in response. It also monitors alarms from the payload processing engine 412 (step 1202). If an alarm is received, the response depends on the alarm type (step 1204). In case of a workload completion alarm, a check is made for the availability of network data transmission (step 1206). If the data network is available, the result data is sent back (step 1207). If not, it starts the process of transferring the result data to the peer device (step 1208). The system may request that the network data remain unavailable for more than a threshold time before starting the offloading process.

[0093] The software on device 115a periodically checks for the presence of an edge computing device running appropriate software, e.g., device 115b having application software, to enable data transfer (step 1210). It checks for one or more available networks, such as Wi-Fi, Bluetooth, NFC, etc. It may also perform a normal network status check, and if it re-establishes a network connection, resumes the normal process of sending back computing data (steps 1212, 1214).

[0094] If a suitable device is found, a connection to the device is established (step 1216). The status of the connected device 115b is queried, which can be done similarly to the device query from the management platform 110, to determine whether the device 115b has sufficient storage to receive the computational data (steps 1218, 1220). The application software 140 on the queried device determines the available storage that has not yet been reserved or allocated for use, for example in connection with its computational job, and returns that data, or the queried device returns more general data regarding available and allocated storage, and the amount of storage determined by the initiating device.

[0095] If sufficient storage is not available, a connection to another device is attempted. If sufficient storage is available, a payload with the computational data is prepared and sent with the data to identify the contents of the payload and the device from which it originated (step 1222). A message is also queued for delivery to the management platform 110 to notify the management platform 110 of the data handoff (step 1224). Each computing device has its own ID. As part of the data handoff, the computing device 115a transfers its ID to the other connected computing device 115b. This information is included as part of the data subsequently transferred by the device 115b to the management platform 110. This allows the management platform 110 to keep track of the transfers that have taken place. If the device 115b has network connectivity issues, at some point it may initiate its own process to transfer the data. The device 115b may attempt to connect to the device 115a before the data from the device 115a is sent back to the management platform (or forwarded to another device 115). When a computing device is in a state where it is returning data to the device from which it originally obtained the data, the target device is queried and indicates the network state and data returned to it only if it has a network connection that can return the data to the management platform 110, thereby avoiding repeated passing of data between devices that do not have the appropriate network connection to return the data.

[0096] If the alert is a device health alert or a resource alert, the process is similar (step 1204). If a network data connection is available (step 1250), a message is sent to the management platform 110 indicating this issue (step 1258). Assuming the network connection is active, the management platform 110 manages the reallocation of computing jobs.

[0097] If a network connection to the management platform 110 is not available, it begins a process to initiate a transfer of the computational job workload to another computing device (steps 1250, 1252). Software on device 115a periodically checks for the presence of an edge computing device running appropriate software, such as device 115b having application software, to enable data transfer (step 1254). It also performs regular network status checks and notifies the management platform of an alert if it re-establishes a network connection (steps 1256, 1258).

[0098] If a suitable device 115b is found, a connection is established to the device (step 1260). The state of the connected device 115b is queried, which can be done similarly to the device query from the management platform 110, to determine if the device 115b has sufficient storage to receive the computational data (steps 1262, 1264). If sufficient storage is not available, a connection to another device is attempted. If the device 115b has sufficient storage to receive the workload payload, it is sent from device 115a to device 115b with the appropriate device and job ID information, and a message is queued to the management platform to inform the transfer when the network connection is re-established (steps 1266, 1268).

[0099] If device 115b is not running a job, it will send the model and data from device 115a to device 115b, which will process the model and data as if it had received the job directly from the management platform 110. Device 115b will also communicate a message to the management platform 110 to inform it of the transferred job. In general, if device 115b is already running a computation job, it will not perform a job transfer. However, if device 115b is running a job with the same computation model as device 115a, but only with different data chunks, then only the transferred data chunks (or only the data chunk IDs, in case of streaming data chunks on request) will be transferred. Device 115a can initially keep device 115b in reserve and look for another computation device that does not currently have a job assigned to it. Increasing the number of data chunks that device 115b must process may affect the time it takes device 115b to complete the job. However, since it is presumed that device 115a is unable to process the job in a meaningful way, the transfer allows processing to continue and management platform 110 then transfers the data chunks of device 115a from device 115b to another suitable computing device.

[0100] After the handoff, the computing device 115a periodically checks whether the network connection has been restored (steps 1226, 1228). If the connection is restored, it contacts the management platform 110 to check the status of the handed off job / job result (step 1230). If the management platform 110 indicates that the job has been completed on another device (step 1232), the device 115a contacts the device 115b where the handoff occurred to check the status of the handoff (step 1234). If the connection between devices 115a / 115b is no longer active, a reconnection can be attempted. If device 115b is out of range or unreachable, the status check is aborted.

[0101] In both data result handoff and job processing handoff scenarios, while computing device 115a is not connected to the network, it may happen that management platform 110 determines computing device 115a is offline and has already transferred the workload to another computing device. If management platform 110 indicates that it has already received the expected results from device 115a, it notifies device 115b to delete the transferred data / cancel the transferred job. If management platform 110 has not received the job results, it signals device 115b to request it to send back the transferred result data, which device 115a sends back to management platform 110. Similarly, computing device 115a may request to send back all or part of the transferred workload (steps 1236-1242) if, for example, computing device 115b is currently processing a data chunk and there are additional chunks remaining to be processed. Transferring back the computation result data or workload may be done similarly to the initial handoff process.

[0102] Figure 13 is a schematic diagram showing communication and data flow between various elements of the system 100 in a cloud computing implementation. The Admin engine 205 is shown in combination with a web server 230. Similarly, various other modules and engines in the cloud platform may be combined in different ways to provide the same overall functionality. The Admin engine 205, connected to a payload manager 215, is used to register and manage computing devices and to publish service metadata about them. The web server stores payload data in the payload container storage 213. The Admin engine / web server 230 also accesses other cloud data storages to store and retrieve other data.

[0103] The payload manager 215 accesses payload data from the payload container storage 213 and connects to the payload services 410 to provide computational model payloads to the computing devices. The connection to the payload services 410 is also used to check payload processing status and handle other payload infrastructure tasks. The connection from the management engine to the messaging service 230 is used to communicate with the computing devices 115, push / pull data for services, and send data chunks to the computing devices, which are retrieved by the messaging service 220 from the payload container storage 213 or other storage.

[0104] Administrators 120' connect to the system via web server interface 230 to access the functionality of admin engine 205, such as creating and managing user accounts and computing devices. Authorized administrators can also access the system via messaging service 220 using a messaging app. Connecting to the payload manager (directly or via a front-end interface such as web server 230 or messaging service 220) is used to manage placement and register computing nodes with patterns or policies to place the system on the appropriate edge node computing devices. Regular users can access the system 100 via web server 230. User links are used to create and manage user accounts, publish and manage edge services and computing payloads, and register / manage their own computing nodes.

[0105] This specification discloses and describes various aspects, embodiments, and examples of the present invention. Alterations, additions, and modifications may be made by one of ordinary skill in the art without departing from the spirit and scope of the invention, as defined in the claims.

Claims

1. A method for distributing a data processing job across a plurality of computing devices in a heterogeneous distributed computing environment, each of the plurality of computing devices having an individual device type and an individual network address and available device resources, the available device resources including processing resources and storage resources, the method comprising: receiving, at a central server, a job payload including an executable computing model and job data to be processed by the computing model, the job data being divisible into a plurality of data chunks, each of the data chunks being processed individually by the computing model; using a benchmark computing platform to determine a benchmark job computing speed by executing the received computing model on a first portion of the received job data; identifying a plurality of computing devices having available device resources sufficient to execute the computing model in at least one data chunk and being remote from the benchmark computing platform; based on the benchmark job computing speed, determining an individual estimated job execution speed for each of the plurality of computing devices; selecting a cluster of computing devices from the plurality of computing devices for executing the computing model in the plurality of data chunks; assigning each of the data chunks to an individual computing device in the cluster, each device in the cluster being assigned at least one data chunk, the individual computing device having an estimated time shorter than a predetermined value for completing processing of the assigned data chunk by the computing model based on the individual estimated job execution speed; transmitting the calculation model and the data chunk assigned to the individual computing device to the individual computing device in the cluster; receiving a calculation result data set from the individual computing device; comprising wherein each calculation result data set is associated with an individual data chunk.

2. The step of transmitting the calculation model and the data chunk assigned to the individual computing device to the individual computing device in the cluster comprises: transmitting the calculation model to the individual computing device using a first software engine; transmitting the data chunk to the individual computing device using a second software engine different from the first software engine; The method according to claim 1, characterized by comprising.

3. The first software engine includes a calculation payload manager, which is configured to communicate with a corresponding payload service in an identified remote device and transfer a payload having a calculation job to the identified remote device, the payload service is operable to extract the calculation job from the payload and execute the calculation job on the selected data, The method according to claim 2, characterized in that the second software engine includes a messaging engine configured to transfer a data message to a corresponding messaging engine in a remote device.

4. The step of determining the benchmark job calculation speed comprises: The method according to claim 1, comprising the step of transmitting the received calculation model and the first part of the received job to a computer system having known operating characteristics.

5. The benchmark platform includes a plurality of benchmark computers, each of the plurality of benchmark computers having a different device type from each other, each of the plurality of benchmark computers being transmitted with the received calculation model and the first part of the received job, and having an individual benchmark job calculation speed, The method further includes the step of identifying, for an individual computing device, a corresponding individual benchmark computer having operating characteristics similar to those of the individual computing device as compared to other benchmark computers among the plurality of benchmark computers according to a similarity vector. The determined estimated job execution speed of the individual computing device is based on the benchmark job calculation speed of the corresponding benchmark computer, the method according to claim 1.

6. Determining a deadline for receiving a calculation result from the first computing device based on the individual estimated job execution speed for the first computing device in the cluster; In response to a determination that the deadline has elapsed without a result being received from the first computing device, reassigning the data chunk assigned to the first computing device from which the calculation result was not received to another computing device selected from the plurality of computing devices; The method according to claim 1, further comprising.

7. The method according to claim 6, wherein the another computing device is selected from a plurality of computing devices not present in the cluster.

8. The first computing device is a battery-powered device, The method further includes, in response to receiving a low battery warning from the first computing device, reassigning a plurality of data chunks that are assigned to the first computing device and for which computation results have not yet been received, to different computing devices, the method according to claim 1, characterized in that.

9. A system for distributing a data processing job in a heterogeneous distributed computing environment, including a computer management platform including a processor, the processor communicates with the memory of the computer, is connected to a network, and the computer management platform is capable of communicating with a plurality of computing devices via the network, the memory stores data identifying an individual network address, device type, and available device resources for an individual computing device, the available device resources include processing resources and storage resources, the memory stores computer instructions, when the computer instructions are executed, receiving a job payload including an executable computing model and job data to be processed by the computing model, the job data being divisible into a plurality of data chunks each to be processed individually by the computing model, determining a benchmark job computation speed based on execution of the received computing model in a first portion of the received job data by a benchmark computing platform remote from the computing device, identifying a plurality of computing devices having available device resources sufficient for execution of the computing model in at least one data chunk, determining an individual estimated job execution speed for each of the plurality of computing devices based on the benchmark job computation speed, Selecting a cluster of computing devices from the plurality of computing devices to execute the computing model in the plurality of data chunks; Allocating each of the data chunks to an individual computing device in the cluster, wherein at least one data chunk is allocated to each device in the cluster, and the individual computing device has an estimated time to complete processing of the allocated data chunk to be executed by the computing model based on the individual estimated job execution speed, and the estimated time is shorter than a predetermined value; allocating each of the data chunks; Transmitting the computing model and the data chunk allocated to the individual computing device to the individual computing device in the cluster; Receiving a computation result dataset from the individual computing device; Configuring the processor to execute; A system, wherein each computation result dataset is associated with an individual data chunk.

10. The computer instructions further include: Using a first software engine to transmit the computing model to the individual computing device; Using a second software engine different from the first software engine to transmit the data chunk to the individual computing device; The system according to claim 9, wherein the processor is configured to execute.

11. The first software engine has a computation payload manager; The computation payload manager is configured to communicate with a corresponding payload service in an identified remote device and transfer a payload having a computation job to the identified remote device; The payload service is operable to extract the calculation job from the payload and execute the calculation job in the selected data. The second software engine has a messaging engine, and the messaging engine is configured to transfer a data message to a corresponding messaging engine in a remote device. The system according to claim 10, characterized in that.

12. The benchmark platform includes a plurality of benchmark computers, each having a different device type. The computer instructions further include For each of the plurality of benchmark computers, to determine an individual benchmark job calculation speed for an individual benchmark computer, sending the received calculation model and the first part of the received job to each of the plurality of benchmark computers; Identifying, for the individual computing device, a corresponding individual benchmark computer having operating characteristics similar to the operating characteristics of the individual computing device compared to other plurality of benchmark computers among the plurality of benchmark computers according to the similarity vector; Configuring the processor to execute The determined estimated job execution speed of the individual computing device is based on the benchmark job calculation speed of the corresponding benchmark computer. The system according to claim 9, characterized in that.

13. The computer instructions further include Determining a deadline for receiving a calculation result from the first computing device based on the individual estimated job execution speed for the first computing device in the cluster; In response to a determination that the deadline has passed without a result being received from the first computing device, reassign the data chunk that is assigned to the first computing device and for which a computation result has not yet been received to another computing device selected from the plurality of computing devices. Configure the processor to execute, the system according to claim 9, wherein the system is characterized by this.

14. The system according to claim 13, wherein the another computing device is selected from a plurality of computing devices that do not exist in the cluster.

15. The first computing device is a battery-powered device. The computer instructions configure the processor to reassign a plurality of data chunks that are assigned to the first computing device and for which computation results have not yet been received to different computing devices in response to receiving a low battery warning from the first computing device, the system according to claim 9, wherein the system is characterized by this.

16. For each of the computing devices, the individual estimated job execution speed is determined based on the benchmark job calculation speed, taking into account a scale factor that reflects the difference in computing power between each of the computing devices and the benchmark device, the method according to claim 1.

17. The method according to claim 16, further comprising the step of determining the scale factor of each of the computing devices at the initial registration of each of the computing devices.

18. The step of determining the scale factor of each of the computing devices is Executing a benchmark test unrelated to the computational model on a benchmark computer in a benchmark system to provide a first performance result; Executing the benchmark test on each of the computing devices to provide a second performance result; Determining the scale factor based on a comparison of the first performance result and the second performance result; The method according to claim 17, comprising:

19. Receiving, from a first computing device in the cluster, a health message indicating a reduction in resources available on the first computing device; Re-evaluating an estimated completion time of a data chunk assigned to the first computing device in consideration of the reduction in resources; Re-assigning at least one data chunk assigned to the first computing device to a different computing device if the estimated completion time exceeds a target deadline; The method according to claim 1, further comprising:

20. The method according to claim 19, wherein the different computing device is selected from the plurality of computing devices.

21. The system according to claim 9, wherein the computer instructions configure the processor to determine an individual estimated job execution speed for each of the computing devices based on the benchmark job calculation speed and considering a scale factor reflecting a difference in computing power between each of the computing devices and a benchmark device.

22. The system according to claim 21, wherein the computer instructions configure the processor to determine the scale factor for each of the computing devices at an initial registration of each of the computing devices.

23. The computer instructions: Execute a benchmark test unrelated to the computing model on a benchmark computer in a benchmark system to provide a first performance result; Execute the benchmark test on each of the computing devices to provide a second performance result; The system according to claim 22, wherein the processor is configured to determine the scale factor based on a comparison between the first performance result and the second performance result. **Claim 24**: The computer instructions, in response to receiving a health message indicating a reduction in resources available at the first computing device from the first computing device in the cluster, re-evaluate an estimated completion time of a data chunk assigned to the first computing device taking into account the reduction in resources, and if the estimated completion time exceeds a target deadline, re-assign at least one data chunk assigned to the first computing device to a different computing device, wherein the processor is configured to perform the above, and the system according to claim 9 is characterized in that. **Claim 25**: The system according to claim 24, wherein the different computing device is selected from the plurality of computing devices. **Claim 26**: The method according to claim 19, further comprising the step of periodically requesting a response message indicating the health status of an individual computing device from the individual computing device in the cluster. **Claim 27**: The system according to claim 24, wherein the computer instructions further configure the processor to periodically request a response message indicating the health status of an individual computing device from the individual computing device in the cluster.