Intelligent management of machine learning inferences in edge cloud systems

By running local machine learning models on edge devices and deciding whether to transmit queries to cloud computing systems, the problem of limited computing resources of edge devices is solved, and efficient machine learning inference and device control is achieved.

CN120163236APending Publication Date: 2025-06-17ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411848800.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-12-16
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Deploying machine learning models on edge devices has problems with limited computing resources and various constraints, especially models that require more computing resources such as deep neural networks are difficult to run on these devices.

Method used

Decide whether to transfer queries to cloud computing systems by running local machine learning models on edge devices and associated with confidence score data. Based on the evaluation results, local prediction data or cloud prediction data are selected as the prediction results to control the edge device.

Benefits of technology

Effectively manage machine learning inference to ensure efficient utilization of computing resources on edge devices, and realize intelligent management and control under limited resource conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163236A_ABST
    Figure CN120163236A_ABST
Patent Text Reader

Abstract

The invention relates to intelligent management of machine learning inferences in an edge cloud system. A computer-implemented system and method associates an edge device with a local machine learning model that generates local prediction data and confidence score data in response to sensor data. Query threshold data is received from a cloud computing system. The evaluation results are evaluated using the confidence score data and the query threshold data. The evaluation result indicates whether a query with sensor data is generated for transmission to the cloud computing system. When the evaluation result indicates that the query is not generated and transmitted, the local prediction data is assigned as a prediction result. The cloud prediction data is assigned as a prediction result when the evaluation result indicates that the query is being generated and transmitted. In response to the query, cloud prediction data is received from the cloud computing system. And controlling the edge device by using the prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to intelligent management of machine learning inference of various machine learning models across edge-cloud systems. Background Art

[0002] Recently, the demand for machine learning (ML) services on edge devices has increased significantly. However, deploying ML models on edge devices (e.g., cleaning robots, smartwatches, etc.) is challenging because edge devices have limited computing resources (e.g., processing resources, memory resources, etc.) and various constraints (e.g., size constraints, power constraints, weight constraints, etc.). There are also many ML models, such as deep neural networks (DNNs), which require far more computing resources than some smaller edge devices can provide. Summary of the Invention

[0003] The following is an overview of certain embodiments described in detail below. The presented aspects are merely provided to give the reader a brief overview of these particular embodiments, and the description of these aspects is not intended to limit the scope of the present disclosure. In fact, the present disclosure may include multiple aspects that may not be explicitly set forth below.

[0004] According to at least one aspect, a computer-implemented method involves controlling an edge device. The method includes receiving sensor data. The method includes generating local prediction data using the sensor data via a local machine learning model. The local prediction data is associated with confidence score data indicating the likelihood of the local prediction data. The method includes receiving query threshold data from a cloud computing system. The method includes generating an evaluation result indicating whether to transmit a query with the sensor data to the cloud computing system. The evaluation result is evaluated using the confidence score data and the query threshold data. The method includes assigning the local prediction data as the prediction result when the evaluation result indicates that the query is not transmitted to the cloud computing system. The method includes assigning cloud prediction data as the prediction result when the evaluation result indicates that the query is being transmitted to the cloud computing system. The cloud prediction data is received from the cloud computing system in response to the query. The method includes using the prediction result to control the edge device.

[0005] According to at least one aspect, a system includes one or more processors and one or more memories. The one or more memories are in data communication with the one or more processors. The one or more memories include computer-readable data stored thereon, which, when executed by the one or more processors, causes the one or more processors to execute a method for controlling an edge device. The method includes receiving sensor data. The method includes generating local prediction data using the sensor data via a local machine learning model. The local prediction data is associated with confidence score data indicating the likelihood of the local prediction data. The method includes receiving query threshold data from a cloud computing system. The method includes generating an evaluation result that indicates whether to transmit a query with the sensor data to the cloud computing system. The evaluation result is evaluated using the confidence score data and the query threshold data. The method includes assigning the local prediction data as the prediction result when the evaluation result indicates that the query is not transmitted to the cloud computing system. The method includes assigning cloud prediction data as the prediction result when the evaluation result indicates that the query is being transmitted to the cloud computing system. The cloud prediction data is received from the cloud computing system in response to the query. The method includes using the prediction result to control the edge device.

[0006] According to at least one aspect, one or more non-transitory computer-readable media have computer-readable data including instructions stored thereon, which, when executed by one or more processors, causes the one or more processors to execute a method for controlling an edge device. The method includes receiving sensor data. The method includes generating local prediction data using the sensor data via a local machine learning model. The local prediction data is associated with confidence score data indicating the likelihood of the local prediction data. The method includes receiving query threshold data from a cloud computing system. The method includes generating an evaluation result that indicates whether to transmit a query with the sensor data to the cloud computing system. The evaluation result is evaluated using the confidence score data and the query threshold data. The method includes assigning the local prediction data as the prediction result when the evaluation result indicates that the query is not transmitted to the cloud computing system. The method includes assigning cloud prediction data as the prediction result when the evaluation result indicates that the query is being transmitted to the cloud computing system. The cloud prediction data is received from the cloud computing system in response to the query. The method includes using the prediction result to control the edge device.

[0007] These and other features, aspects, and advantages of the present invention will be discussed in the following detailed description with reference to the drawings, throughout which like characters represent like or identical parts. Additionally, the drawings are not necessarily to scale, as some features may be enlarged or reduced to show details of particular components. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1It is a flowchart showing an exemplary aspect of an edge cloud system according to an exemplary embodiment of the present disclosure.

[0009] Figure 2 It is a block diagram showing an exemplary aspect of an edge device according to an exemplary embodiment of the present disclosure.

[0010] Figure 3 It is a block diagram showing an exemplary aspect of a cloud computing system according to an exemplary embodiment of the present disclosure.

[0011] Figure 4 It is a flowchart showing an exemplary aspect of an edge device according to an exemplary embodiment of the present disclosure.

[0012] Figure 5 It is a flowchart showing an exemplary aspect of a cloud computing system according to an exemplary embodiment of the present disclosure. Detailed Description

[0013] From the foregoing description, it will be understood that the embodiments described herein and their many advantages have been shown and described by way of example, and it is apparent that various changes may be made in the form, construction, and arrangement of the components without departing from the disclosed subject matter or sacrificing one or more of its advantages. In fact, the described forms of these embodiments are merely illustrative. These embodiments are susceptible to various modifications and alternative forms, and the following claims are intended to cover and include such changes and are not limited to the specific forms disclosed, but rather cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure.

[0014] Figure 1 It is a flowchart showing aspects of a system 100 that includes at least an edge cloud system according to an exemplary embodiment. The system 100 is configured to intelligently manage machine learning inferences generated by various machine learning models on different computer systems (e.g., a cloud computing system of an edge device and an edge cloud system, a local server on an edge device and a computer network system, etc.). The system 100 includes at least one local computing system (e.g., an edge device 200) and at least one remote computing system (e.g., a cloud computing system 300). More specifically, in Figure 1 the example shown, the system 100 includes a plurality of edge devices and a cloud computing system 300. The cloud computing system 300 communicates data with each edge device 200 via a communication technology 10, which can be wired, wireless, or a combination thereof.

[0015] In Figure 1In [the context], each edge device 200 is operably connected to a cloud computing system 300 via a communication technology. Each edge device 200 performs data processing functions at the "edge" of the network. In this regard, each edge device 200 is a functional technical device, which is also configured to act as at least an entry point to the network. Each edge device 200 is configured to interface with one or more users. As a non-limiting example, for instance, the edge device 200 may include a mobile robot (e.g., a robotic vacuum cleaner, etc.), a smart watch, an Internet of Things (IoT) device, or any similar edge technology.

[0016] Reference Figure 1 , as a non-limiting example, each edge device 200 is a cleaning robot (e.g., a robotic vacuum cleaner). Each edge device 200 includes one or more sensors, which are configured to capture sensor data based on, for example, objects in its environment. In this regard, the edge device 200 generates sensor data at least based on the environment. Thus, there may be variations in the timing and amount of sensor data obtained by a particular edge device 200. Each edge device 200 also includes a plurality of components related to its application, as well as a machine learning (ML) model 210 and a query controller 212.

[0017] The ML model 210 is located on the edge device 200. The ML model 210 may be referred to as a local ML model for local adoption on the edge device 200. The ML model 210 may be a lightweight model. The ML model 210 may include a convolutional neural network (CNN) and / or any artificial neural network, which is configured to perform a predetermined task (e.g., classification, etc.) for the edge device 200. In this regard, the ML model 210 is configured to generate at least prediction data (or local prediction data 102) based on input data. The ML model 210 is also configured to generate confidence score data, which provides likelihood data or probability data for the corresponding prediction data. As a non-limiting example, for instance, in Figure 1 , each ML model 210 includes a You Only Look Once (YOLO) model, such as YOLOv5s (small), which is configured to perform the predetermined task of real-time object detection.

[0018] The query controller 212 is located on or associated with the edge device 200. The query controller 212 is configured to receive local prediction data from the ML model 210. Additionally, the query controller 212 is configured to receive confidence score data corresponding to each local prediction data. The query controller 212 is configured to generate an evaluation result regarding the local prediction data. The query controller 212 is configured to determine whether a query including the input data should be sent to the cloud computing system 300 based on the evaluation result. In this regard, the query controller 212 is configured to determine whether to offload the input data to the cloud computing system 300 to obtain a more accurate machine learning inference from the cloud computing system 300 compared to the local machine learning inference obtained from the ML model 210. The query controller 212 includes software technology, hardware technology, or a combination of hardware and software technologies.

[0019] As described above, the query controller 212 is configured to perform an adaptive cloud query based on the evaluation result. The query controller 212 is configured to generate an evaluation result based on a set of components. In one example embodiment, the query controller 212 is configured to generate an evaluation result based on three components. For example, the first component includes the prediction confidence of the ML model 210. In this regard, for example, the ML model 210 is configured to provide confidence score data associated with the local prediction data. The confidence score data indicates the likelihood or probability regarding the corresponding local prediction data. The second component includes query threshold data, which is a confidence level compared with the confidence score data generated from the ML model 210 to determine whether to send a query to the cloud computing system 300. The query threshold data is determined and set by the cloud computing system 300 based on the activity level or busy level of the cloud computing system 300. The third component includes the average round-trip network latency between the edge device 200 and the cloud computing system 300. This third component can be obtained by the edge device 200 by subtracting the cloud processing time from the round-trip time. The cloud processing time refers to the amount of time the cloud computing system 300 takes to process a query. Thus, based on at least these three components, the query controller 212 is configured to determine whether to generate a query to the cloud computing system 300.

[0020] The query controller 212 is configured to generate a query that includes a portion of the input data, a version of the input data, or all of the input data. As an example, the query controller 212 is configured to generate a query that includes the input data (e.g., sensor data such as a digital image). In this regard, the query controller 212 forwards the same input data that is processed by the ML model 210 to the cloud computing system 300. As another example, the query controller 212 is configured to transmit a query that includes intermediate data of the ML model 210 (instead of sending the entire input data). As a non-limiting example, the intermediate data is extracted or output from a specific layer of the ML model 210 (e.g., the second CNN layer). When generating a query that includes some form of input data of the ML model 210, the query controller 212 is configured to send the query to the cloud computing system 300 for processing, such that the cloud computing system 300 can provide more accurate machine learning inferences to the edge device 200 by using one of its ML models 314.

[0021] The cloud computing system 300 includes multiple ML models 314, which are hosted in multiple cluster nodes 316. A group of cluster nodes 316 can be grouped together in a cluster 318. Each cluster can include a set of "n" nodes, where "n" represents an integer greater than 1. The ML models 314 are configured to generate prediction data (which can be referred to as cloud prediction data 104) based on the input data, which is transmitted to the ML models 314 as a query from the edge device 200. Compared with the ML model 210 of the edge device 200, the ML models 314 of the cloud computing system 300 are larger models, such that the ML models 314 provide higher accuracy than the ML model 210. Compared with the ML model 210, the ML models 314 are higher-performance models. The number of parameters of the ML models 314 can be greater than the number of parameters of the ML model 210. The amount of resources (e.g., memory resources, processing resources, etc.) used by the ML models 314 may be greater than the amount of resources used by the ML model 210. The ML models 314 are configured to perform the same or similar tasks as the ML model 210. As a non-limiting example, for instance, in Figure 1 each ML model 314 includes a YOLO model, such as YOLOv5x (x-large), which is configured to perform a predetermined task of real-time object detection.

[0022] The cloud computing system 300 further includes a load balancer 310 configured to receive all queries from the edge device 200 before they are transmitted to the ML model 314. The number of queries transmitted to the cloud computing system 300 varies over time in terms of intensity, load, and / or arrival pattern. The load balancer 310 acts as an intermediary between the edge device 200 and the ML model 314 to manage the distribution of queries to the ML model 314. The load balancer 310 includes software technology, hardware technology, or a combination of software technology and hardware technology. Additionally, the load balancer 310 is configured to generate load status data for a given time period. For example, for a given time period, the load status data may include the number of connected edge devices 200 and / or the number of queries received from the edge device 200.

[0023] The cloud computing system 300 can be queried by one or more edge devices 200 simultaneously. The cloud computing system 300 can be queried by different numbers of edge devices 200 at different times. The cloud computing system 300 is configured to intelligently calculate query threshold data for a given time period. The query threshold data is determined by the activity level or busyness level of the cloud computing system 300. The cloud computing system 300 can generate a query response including the query threshold data. The cloud computing system 300 is configured to transmit the query threshold data and / or the query response to each edge device 200.

[0024] As Figure 1 shown, the cloud computing system 300 further includes a reinforcement learning (RL) agent 312. The RL agent 312 is configured to collect load status data from the load balancer 310 and collect cluster status data from multiple cluster nodes 316. The RL agent 312 is configured to use the load status data and the cluster status data to generate system status data. The RL agent 312 is configured to perform multiple actions on one or more cluster nodes 316 based on the system status data, such that the cloud computer system 300 is intelligently managed.

[0025] Figure 2 is a block diagram of an example of the edge device 200 according to an example embodiment. The edge device 200 includes at least a processing system 204 having at least one processing device. For example, the processing system 204 may include an electronic processor, a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a microprocessor, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), any processing technology, or any number and combination thereof. The processing system 204 is operable to provide the functions described herein.

[0026] The edge device 200 includes a memory system 206 operably connected to a processing system 204. In this regard, the processing system 204 communicates data with the memory system 206. In an example embodiment, the memory system 206 includes at least one non-transitory computer-readable storage medium configured to store and provide access to various data such that at least the processing system 204 can perform the operations and functions disclosed herein. In an example embodiment, the memory system 206 includes a single memory device or multiple memory devices. The memory system 206 may include electrical, electronic, magnetic, optical, semiconductor, electromagnetic, or any suitable storage technology operable with the edge device 200. For example, in one example embodiment, the memory system 206 may include random access memory (RAM), read-only memory (ROM), flash memory, disk drives, memory cards, optical storage devices, magnetic storage devices, memory modules, any suitable type of memory device, or any number and combination thereof.

[0027] The memory system 206 at least includes an edge program 208, an ML model 210, a query controller 212, and other related data 214, which are stored on the memory system 206 and include computer-readable data with instructions that, when executed by the processing system 204, are configured to perform the functions disclosed herein. The computer-readable data may include instructions, code, routines, various related data, any software technology, or any number and combination thereof. The edge program 208 is configured to perform various functions of the edge device 200. For example, the edge program 208 is configured to manage machine learning inferences and / or control the edge device 200 based on machine learning inferences. The ML model 210 includes at least one machine learning system (e.g., artificial neural network, deep neural network, etc.) configured to perform tasks (e.g., classification, etc.) of the edge device 200. Compared with the ML model 314, the ML model 210 is a smaller model. In this regard, the ML model 210 may have fewer parameters than the ML model 314. The ML model 210 uses fewer resources (e.g., memory resources, processing resources, etc.) than the ML model 314. In this regard, for example, the ML model 210 is configured to generate local prediction data based on input data. The query controller 212 is also configured to evaluate the local prediction data of the ML model 210 and determine whether a query should be sent to the cloud computing system 300 based on its evaluation. At the same time, the other related data 214 provides various data (e.g., operating system, etc.), which enables the system 100 to perform the functions discussed herein.

[0028] The edge device 200 is configured to include at least one sensor system 202. The sensor system 202 includes one or more sensors. For example, the sensor system 202 includes an image sensor, a camera, a radar sensor, a light detection and ranging (LIDAR) sensor, a thermal sensor, an ultrasonic sensor, an infrared sensor, a motion sensor, an audio sensor (e.g., a microphone), any suitable sensor, or any number and combination thereof. The sensor system 202 is operable to communicate with one or more other components of the edge device 200 (e.g., the processing system 204 and the memory system 206). For example, the sensor system 202 may provide sensor data, which is then used by the processing system 204 to generate digital image data based on the sensor data. In this regard, the processing system 204 is configured to obtain sensor data as digital image data directly or indirectly from one or more sensors of the sensor system 202. The sensor system 202 is local, remote, or a combination thereof (e.g., partially local and partially remote). Upon receiving the sensor data, the processing system 204 is configured to process the sensor data (e.g., image data) in conjunction with the edge program 208, the ML model 210, the query controller 212, other relevant data 214, or any number and combination thereof.

[0029] In addition, the edge device 200 may include at least one other component. For example, as Figure 2 shown, the memory system 206 is also configured to store other relevant data 214, which relates to the operation of the edge device 200 associated with one or more components (e.g., at least one sensor system 202, at least one I / O device 216, and other functional modules 218). In addition, the edge device 200 includes one or more I / O devices 216 (e.g., a display device, a microphone, a speaker, etc.). The edge device 200 also includes other functional modules 218, such as any suitable hardware, software, or combination thereof that aids or contributes to the functionality of the edge device 200. For example, the other functional modules 218 include communication technologies (e.g., wired communication technologies, wireless communication technologies, or a combination thereof), which enable the components of the edge device 200 to communicate with each other as described herein. The other functional modules 218 may also include actuators 220. As a non-limiting example, for instance, when the edge device 200 is a robotic vacuum cleaner, then one or more actuators 220 may be involved in driving, steering, stopping, and / or controlling the movement of the robotic vacuum cleaner. In this regard, using the edge program 208, the edge device 200 is configured to manage machine learning inferences, which are typically generated locally on the edge device 200 and / or obtained remotely via the cloud computing system 300.

[0030] Figure 3FIG. is a block diagram of an example of a cloud computing system 300 according to an example embodiment. The cloud computing system 300 includes at least one processing system 302 having at least one processing device. For example, the processing system 302 may include an electronic processor, CPU, GPU, TPU, microprocessor, FPGA, ASIC, any processing technology, or any number and combination thereof. Each cluster node 316 may be associated with one or more processors of the processing system 302. The processing system 302 is operable to provide the functions described herein.

[0031] The cloud computing system 300 includes a memory system 306 operably connected to the processing system 302. In this regard, the processing system 302 communicates data with the memory system 206. In an example embodiment, the memory system 306 includes at least one non-transitory computer-readable storage medium configured to store various data and provide access to various data such that at least the processing system 302 can perform the operations and functions disclosed herein. The memory system 306 is typically very large in size. In this regard, the memory system 306 is significantly larger than the memory system 206 of the edge device 200. In an example embodiment, the memory system 306 includes a single memory device or multiple memory devices. The memory system 306 may include electrical, electronic, magnetic, optical, semiconductor, electromagnetic, or any suitable storage technology operable with the cloud computing system 300. For example, in an example embodiment, the memory system 306 may include random access memory (RAM), read only memory (ROM), GPU high bandwidth memory (HBM), flash memory, disk drives, memory cards, optical storage devices, magnetic storage devices, memory modules, any suitable type of memory device, or any number and combination thereof.

[0032] The memory system 306 includes at least a cloud application 308, a load balancer 310, an RL agent 312, one or more cluster nodes 316 having an ML model 314 (“cloud ML model”), and other related data 320, which is stored thereon and includes computer-readable data with instructions that, when executed by the processing system 204, are configured to perform the functions disclosed herein. More specifically, the cloud application 308 is configured to operate and control the cloud computing system 300. The computer-readable data may include instructions, code, routines, various related data, any software technology, or any number and combination thereof. In an example embodiment, the ML model 314 includes at least one machine learning model that is a larger and higher-performance model than the ML model 210 while being configured to perform at least the same tasks as the ML model 210. In this regard, for example, the ML model 210 may be a lightweight version of the ML model 314. Additionally, each cluster node 316 hosts a set of ML models 314. The load balancer 310 is also configured to receive and manage queries from the edge devices 200. The load balancer 310 is configured to generate load status data regarding the current load (e.g., the number of queries) of the cloud computing system 300. The RL agent 312 is configured to generate system status data based on the load status data from the load balancer 310 and the cluster status data from the cluster nodes 316. The RL agent 312 is configured to take one or more corresponding actions based on the system status data. Meanwhile, the other related data 320 provides various data (e.g., operating systems, etc.), which enables the cloud computing system 300 to perform the functions discussed herein.

[0033] In addition, the cloud computing system 300 may include at least one other component. For example, as Figure 3 shown, the memory system 306 is also configured to store other related data 320, which relates to the operation of the cloud computing system 300 associated with one or more components of the cloud computing system 300 and / or the edge devices 200 of the network. Additionally, the cloud computing system 300 is configured to include one or more input / output (I / O) devices 322 (e.g., display devices, keyboard devices, speaker devices, etc.), which are associated with the cloud computing system 300. The cloud computing system 300 also includes other functional modules 304, such as any suitable hardware, software, or combination thereof that aids or contributes to the functions of the cloud computing system 300. For example, the other functional modules 304 include communication technologies (e.g., wired communication technologies, wireless communication technologies, or a combination thereof), which enable the components of the cloud computing system 300 to communicate with each other and / or with each edge device 200, as described herein.

[0034] Figure 4is a flowchart of an example of process 400 of edge device 200 according to an example embodiment. Process 400 includes multiple steps. In this regard, process 400 may include more or fewer steps than Figure 4 shown, as long as the same or similar functions are provided. Process 400 is performed by at least one or more processors of edge device 200. In this regard, although system 100 may include multiple edge devices 200, process 400 is explained with respect to one edge device 200 as an illustrative example.

[0035] In step 402, according to one example, edge device 200 is configured to receive input data. In this regard, for example, edge device 200 is in an operating state and waiting to receive input data. The input data may include sensor data from one or more sensors of sensor system 202. The input data includes user input from one or more I / O devices 216 of edge device 200. For example, the input data may include sensor data or sensor fusion data (e.g., one or more digital images and / or digital videos).

[0036] In step 404, according to one example, edge device 200 determines whether input data for ML model 210 has been received. Edge device 200 may also determine whether the input data is valid and / or suitable input for ML model 210. For example, edge device 200 is configured to receive input data that may include sensor data from sensor system 202, user input from I / O device 216, any suitable data, or any number and combination thereof. When edge device 200 receives input data in step 404, process 400 proceeds to step 406. Alternatively, when edge device 200 does not receive input data in step 404, then process 400 proceeds to step 402.

[0037] In step 406, according to one example, edge device 200 locally performs inference using the input data via ML model 210. More specifically, ML model 210 generates output data (e.g., local prediction data 102 and confidence score data) based on the input data. After locally performing inference on edge device 200, then process 400 proceeds to step 408.

[0038] In step 408, according to one example, edge device 200 determines whether to query cloud computing system 300. For example, as discussed with respect to Figure 1 the query controller 212 is configured to generate an evaluation result indicating whether to offload input data (e.g., sensor data) to cloud computing system 300. The query controller 212 generates the evaluation result based on at least three components (e.g., confidence score data, query threshold data, and network latency data).

[0039] The edge device 200 is configured to generate evaluation data by evaluating a non - negative monotonically increasing function that involves at least confidence score data regarding query threshold data and network latency. For example, f(Conf, L network ) can be used to represent a non - negative monotonically increasing function that receives Conf and L network as input data. More specifically, as an example, for instance, f(Conf, L network ) = Conf + w·L network , where Conf refers to the confidence score data, w refers to the weighting factor, L network refers to the latency of the network, and Thres refers to the query threshold data. In this example, w can be selected based on the expected network latency of the average edge device 200 from the system 100. Specifically, w refers to the application - specific relative sensitivity of the offloading decision with respect to the confidence score compared to the network latency. The query threshold data can be determined by the cloud computing system 300 based on the activity level and / or offloading trend controlled by the cloud computing system 300.

[0040] Conf + w·L network <Thres [1]

[0041] If Equation 1 is satisfied and the inequality is true (i.e., the left - hand expression is less than the query threshold data), then the edge device 200 generates an evaluation result indicating offloading / querying the cloud computing system 300. When Equation 1 is satisfied, the process 400 proceeds to step 410. Alternatively, if Equation 1 is not satisfied and the inequality is false (i.e., the left - hand expression is greater than or equal to the query threshold data), then the edge device 200 generates an evaluation result indicating not offloading and not querying the cloud computing system 300. When Equation 1 is not satisfied, the process 400 proceeds to step 412.

[0042] In step 410, according to an example, the edge device 200 generates a query that includes the input data or some form of the input data. The edge device 200 sends the query as an asynchronous request to the cloud computing system 300. At this point, the cloud computing system 300 receives the query from the edge device 200. More specifically, the load balancer 310 receives the query and transmits the query (e.g., sensor data) to the ML model 314. The cloud computing system 300 generates prediction data (or cloud prediction data 104) based on the input data via the ML model 314.

[0043] In step 412, according to an example, the edge device 200 processes the outputs from the following respectively: (i) the ML model 210 or (ii) the ML model 210 and the ML model 314. More specifically, in the first case, when the evaluation result indicates that the query should not be sent to the cloud computing system 300, the edge device 200 can process the output from the ML model 210 (e.g., the local prediction data 102). In this first case, the edge device 200 is configured to assign the local prediction data 102 as the prediction result 106.

[0044] Alternatively, in the second case, the edge device 200 processes the output from the ML model 210 and then determines to offload the input data to the cloud computing system 300 based on the evaluation result. The edge device 200 then processes the output from the cloud computing system 300 (e.g., the cloud prediction data 104 and the query threshold data). In this second case, the edge device 200 is configured to assign the cloud prediction data 104 as the prediction result 106. That is, in this second case, the edge device 200 does not assign the local prediction data 102 as the prediction result 106.

[0045] The edge device 200 is configured to provide the prediction result 106 as the output data for a given machine learning task in response to the input data. As a non-limiting example, for instance, if the edge device 200 is a robotic vacuum cleaner, the edge device 200 is configured to use the prediction result 106 to control one or more actuators of the robotic vacuum cleaner, and the prediction result 106 is selected as the local prediction data 102 or the cloud prediction data 104. In this case, the robotic vacuum cleaner can receive a digital image from a camera sensor on the robotic vacuum cleaner as the input data. The ML model 210 and the ML model 314 can be configured to perform a classification task to identify objects in the digital image in order to control the operation of the robotic vacuum cleaner based on these identified objects.

[0046] As Figure 4As discussed, the edge device 200 receives input data. Before the query controller 212 decides whether to offload the input data to the cloud computing system 300, the input data is locally processed by the ML model 210 of the edge device 200. If the query controller 212 decides to offload the input data as a query to the cloud computing system 300, the edge device 200 transmits the query including the input data to the cloud computing system 300 as an asynchronous request. In this regard, with this asynchronous request, the edge device 200 does not have to wait to receive a query response from the cloud computing system 300 before processing the next input data. The edge device 200 is configured to immediately and directly wait for and / or obtain the next input data of the ML model 210, even if a query response has not been received. The edge device 200 is configured to receive the latest query threshold data after querying the cloud computing system 300. Since the cloud computing system 300 will include the query threshold data in the query response, this query threshold data will be used to update the query controller 212. The query controller 212 uses the updated query threshold data to determine whether to query the cloud computing system 300 next time.

[0047] As discussed, the process 400 is advantageous in implementing an adaptive query technique to the cloud computing system 300, which maintains a relatively simple but effective query controller 212 on the edge device 200 while maintaining a more complex and computationally demanding RL agent 312 on the cloud computing system 300. In this regard, the RL agent 312 is configured to generate system state data and perform multiple actions based on the system state data. For example, the RL agent 312 is configured to allocate or deallocate resources of the cloud computing system 300 during load scaling. The RL agent 312 is configured to monitor and / or control latency costs. The RL agent 312 is configured to monitor and / or control the costs associated with operating the cloud computing system 300. The RL agent 312 is configured to calculate query threshold data and transmit the query threshold data to each edge device 200. In addition, the RL agent 312 is configured to perform at least one action at one or more fixed time intervals. For example, the RL agent 312 is configured to perform at least one action in at least one set of predetermined actions, as shown in Table 1.

[0048]

[0049] Compared with the previous action time, the system 100 may also have different system dynamics at the current action time (e.g., more edge devices 200 may have started serving, more ML models 314 may be adopted, more cluster nodes 316 may be allocated). The system 100 may exhibit or include multiple system states. Each system state is represented by system state data. In this regard, the system state may be represented by one or more of the following characteristics, as shown in Table 2.

[0050]

[0051]

[0052] Since the number of actions is limited (e.g., the five actions in Table 1), and the number of system states is high-dimensional and has continuous features, the RL agent 312 is trained via at least one deep RL algorithm. For example, the RL agent 312 includes a standard Deep Q-Network (DQN) that uses a DNN model to approximate Q-values and selects the action that returns the best Q-value (i.e., the cumulative long-term reward). More specifically, Table 3 describes multiple aspects of the reward function for the above example, where the RL agent 312 uses DQN.

[0053]

[0054] Although the above example involves the RL agent 312 using standard DQN, the RL agent 312 can involve other RL algorithms. As an example, the RL agent 312 includes a Soft Actor-Critic, which involves an off-policy actor-critic deep RL algorithm based on the maximum entropy reinforcement learning framework. In this off-policy actor-critic deep RL framework, the RL agent 312 or the stochastic actor aims to maximize the expected reward while also maximizing the entropy. As another example, the RL agent 312 uses a Double DQN algorithm. In this regard, the RL agent 312 can include RL algorithms that provide the functions and objectives described in the present disclosure.

[0055] Figure 5 is a flowchart showing an example of a process 500 of the RL agent 312 of the cloud computing system 300 according to an example embodiment. In this regard, Figure 5 shows a process 500 having multiple steps executed by one or more processors of the cloud computing system 300 via the RL agent 312. The process 500 can include more or fewer steps than Figure 5 shown, as long as such steps provide at least the same or similar functions as described herein.

[0056] In step 502, the RL agent 312 interacts with the environment of the cloud computing system 300. The RL agent 312 is configured to evaluate the current system state of the cloud computing system 300 during the current time period. More specifically, at fixed intervals, the RL agent 312 uses the load status data from the load balancer 310 and the cluster status data of the cluster nodes 316 to determine and generate the current system state data. After determining the current system state of the environment and generating the current system state data, the process 500 proceeds to step 504.

[0057] In step 504, the RL agent 312 selects and implements at least one RL policy that is applicable based on the current system state data obtained in step 502. Once the RL agent 312 has the system state data, the trained RL agent 312 uses one or more RL policies to determine the best action to take in the current time period. For example, the RL agent 312 can change the amount of cloud computing resources (e.g., GPUs, CPUs, TPUs, etc.) in the cluster nodes 316, update the query threshold data, take another predetermined action, or any combination thereof.

[0058] In addition, in the example discussed above, the system 100 is modeled as a Markov decision process (MDP), as shown in Equation 2. More specifically, in Equation 2, S represents the complete system state space (including all edge devices and cloud servers). In Equation 2, A represents a set of actions (discrete) that can be performed by the RL agent 312 in the cloud computing system 300. At the same time, in Equation 2, R represents the reward for taking a specific action under specific system state data. The system 100 maps the state-action pair to an immediate reward. P is the transition probability kernel that defines the probability measure of the next system state and reward.

[0059] M sys =(S, A, R, P) [2]

[0060] The goal is to find a policy as shown in Equation 3, where the expectation is taken with respect to the transition probability kernel and γ ∈ (0, 1) represents the future discounted reward. The time step is denoted as i, ranging from the current time t to the infinite future.

[0061]

[0062] In addition, as Figure 1 shown, the RL agent 312 is configured to manage cloud resources and control the offloading trend of the edge devices 200. The RL agent 312 is configured to select and execute an action from a set of predetermined actions. For example, in this example, the actions include doing nothing. The actions include adding or generating a certain amount of processing containers (e.g., GPUs, CPUs, etc.) in the cloud computing system 300. The actions include removing a certain amount of processing containers (e.g., GPUs, CPUs, etc.) from the cloud computing system 300. The actions include increasing the query threshold data or increasing the offloading level of the edge devices 200. The actions include decreasing the query threshold data or decreasing the offloading level of the edge devices 200. After selecting an action, the RL agent 312 is configured to execute the action at time t, which can be denoted as a in Equation 4 tIn addition, in Equation 4, M represents a number that can be set by the user, for example. As a non-limiting example, for instance, if M = 10, then a t = 1 means that the RL agent 312 is configured to increase the number of GPU containers by 10% at time t. In addition, L represents a number that can be set according to the applications and / or configurations of the system 100. As a non-limiting example, for instance, if L = 0.05, then a t = 3 means that the RL agent 312 is configured to increase the edge offloading threshold by 0.05, and then a t = 4 means that the RL agent 312 is configured to decrease the edge offloading threshold by 0.05.

[0063]

[0064] In addition, the state space is for the entire system 100, which includes the central cloud computing system 300 and all edge devices 200. As described above, the system state data can be determined by one or more of the features discussed in Table 2. More specifically, in the example embodiment, the system state data s t at time t can be represented by a seven-state tuple vector, as expressed in Equation 5.

[0065] s t = [N edge , ΔN edge , N ctoua , N req , L req , C cloua , Thres] [5]

[0066] In addition, the goal of the RL agent 312 on the cloud side of the system 100 is to maximize the number of predictions executed by the cloud computing system 300 without exceeding the cost budget, while ensuring the latency goal L target . As an example, the system 100 and / or the RL agent 312 define the immediate reward as R(s t , a t ), as shown in Equation 6. N req and C cloud only depend on the current state, while Cost(a t ) returns the cost of adding / removing cloud resources. In addition, in Equation 6, k represents a given application-specific number that weights the relative sensitivity of the immediate reward to the number of requests or queries completed by the cloud computing system 300 versus the cloud cost Cost(a t ). More specifically, in Equation 6, k ≥ 0.

[0067]

[0068] As described in this disclosure, embodiments include many advantages and provide many benefits. For example, system 100 is cost-effective and robust to scaling of edge devices 200 because system 100 is configured to quickly adjust its resource allocation on cloud computing system 300. Cloud computing system 300 dynamically adapts to the needs of edge devices 200 and controls the query threshold data of each edge device 200 accordingly. Additionally, system 100 is configured to prevent at least two major drawbacks associated with a fixed amount of computing resources. For example, system 100 does not incur unnecessary cloud resource costs during low demand periods from edge devices 200. When transitioning from low demand to high demand, system 100 also prevents overshooting latency of cloud computing system 300.

[0069] Moreover, system 100 is advantageous in providing edge devices 200, whereby each edge device 200 is configured to provide prediction results of machine learning tasks at a higher prediction accuracy and faster rate via its well-managed communication with cloud computing system 300. A company providing edge devices 200 as products can also manage and control its cloud operation costs while providing the benefits of cloud resources to its customers.

[0070] Furthermore, system 100 is configured to intelligently manage cloud operation costs and its scaling when the number of edge devices 200 connected to cloud computing system 300 changes (e.g., the number of edge devices 200 increases sharply within a relatively short amount of time). Cloud operation cost is a key metric that can directly impact revenue when providing high-accuracy inference of cloud machine learning models to edge devices 200. Additionally, scalability is also a beneficial feature because many users can use their edge devices 200 simultaneously (i.e., peak times). Advantageously, system 100 is configured to manage cloud costs and latency even when the number of edge devices 200 changes.

[0071] Furthermore, the foregoing description is intended to be illustrative and not restrictive, and is provided in the context of a particular application and its requirements. Those skilled in the art will understand from the foregoing description that the present invention can be implemented in a variety of forms, and that various embodiments can be implemented individually or in combination. Thus, although embodiments of the present invention have been described in connection with specific examples of the present invention, the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of the described embodiments, and the true scope of the embodiments and / or methods of the present invention is not limited to the embodiments shown and described, as various modifications will become apparent to those skilled in the art after studying the drawings, the specification, and the appended claims. Additionally or alternatively, components and functions can be separated or combined in a different manner than in the various described embodiments, and different terms can be used to describe them. These and other variations, modifications, additions, and improvements may fall within the scope of the present disclosure as defined in the subsequent claims.

Claims

1. A computer-implemented method for controlling an edge device, the computer-implemented method comprising: Receive sensor data; generating, via a local machine learning model, local prediction data using the sensor data, the local prediction data being associated with confidence score data indicating a likelihood of the local prediction data; receiving query threshold data from a cloud computing system; generating an evaluation result indicating whether to transmit the query with the sensor data to the cloud computing system, the evaluation result being evaluated using the confidence score data and the query threshold data; When the evaluation result indicates that the query is not transmitted to the cloud computing system, assigning the local prediction data as the prediction result; When the evaluation result indicates that the query is being transmitted to the cloud computing system, assigning cloud prediction data as a prediction result, the cloud prediction data being received from the cloud computing system in response to the query; as well as Use the prediction results to control edge devices.

2. The computer-implemented method of claim 1 , wherein: The evaluation result is generated by evaluating an inequality, wherein the inequality is f(Conf,L network )<Thres, in Conf is a number representing the confidence score data, L network is a number greater than zero and represents the average round-trip network delay between the edge device and the cloud computing system, f(Conf,L network ) is based on Conf and L network is a non-negative monotonically increasing function of the input data, and Thres indicates query threshold data.

3. The computer-implemented method of claim 2, wherein: When the inequality is satisfied and evaluated to be true, the evaluation result indicates querying the cloud computing system; as well as When the inequality is not satisfied and is evaluated to be false, the evaluation result indicates that the cloud computing system is not queried.

4. The computer-implemented method of claim 2, wherein: f(Conf,L network )=Conf+w·L network , Where w is a number representing a weighting factor indicating the relative network sensitivity.

5. The computer-implemented method of claim 1 , wherein: The cloud prediction data is generated by a cloud machine learning model that is remote from the edge device and is part of a cloud computing system; and Cloud machine learning models are larger and more accurate than local machine learning models on edge devices.

6. The computer-implemented method of claim 1 , wherein: When the evaluation result indicates that the query is not transmitted to the cloud computing system, the sensor data is not transmitted to the cloud computing system.

7. The computer-implemented method of claim 1 , further comprising: Controlling the actuator based on the prediction results, Wherein the actuator is controlled via an edge device.

8. A system comprising: one or more processors; as well as One or more memories in data communication with the one or more processors, the one or more memories including computer readable data stored thereon, which when executed by the one or more processors causes the one or more processors to perform a method for controlling an edge device, the method comprising: Receive sensor data; generating, via a local machine learning model, local prediction data using the sensor data, the local prediction data being associated with confidence score data indicating a likelihood of the local prediction data; receiving query threshold data from a cloud computing system; generating an evaluation result indicating whether to transmit the query with the sensor data to the cloud computing system, the evaluation result being evaluated using the confidence score data and the query threshold data; When the evaluation result indicates that the query is not transmitted to the cloud computing system, assigning the local prediction data as the prediction result; When the evaluation result indicates that the query is being transmitted to the cloud computing system, assigning cloud prediction data as a prediction result, the cloud prediction data being received from the cloud computing system in response to the query; and Use the prediction results to control edge devices.

9. The system according to claim 8, wherein: The evaluation result is generated by evaluating an inequality, wherein the inequality is f(Conf,L network )<Thres, in Conf is a number representing the confidence score data, L network is a number greater than zero and represents the average round-trip network delay between the edge device and the cloud computing system, f(Conf,L network ) is based on Conf and L network is a non-negative monotonically increasing function of the input data, and Thres indicates query threshold data.

10. The system of claim 9, wherein: When the inequality is satisfied and evaluated to be true, the evaluation result indicates querying the cloud computing system; and When the inequality is not satisfied and is evaluated to be false, the evaluation result indicates that the cloud computing system is not queried.

11. The system of claim 9, wherein: f(Conf,L network )=Conf+w·L network , Where w is a number representing a weighting factor indicating the relative network sensitivity.

12. The system of claim 8, wherein: The cloud prediction data is generated by a cloud machine learning model that is remote from the edge device and is part of a cloud computing system; and Cloud machine learning models are larger and more accurate than local machine learning models on edge devices.

13. The system according to claim 8, wherein: When the evaluation result indicates that the query is not transmitted to the cloud computing system, the sensor data is not transmitted to the cloud computing system.

14. The system of claim 8, further comprising: Actuator, Wherein the actuator is controlled via an edge device using the prediction results.

15. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform a method for controlling an edge device, the method comprising: Receive sensor data; generating, via a local machine learning model, local prediction data using the sensor data, the local prediction data being associated with confidence score data indicating a likelihood of the local prediction data; receiving query threshold data from a cloud computing system; generating an evaluation result indicating whether to transmit the query with the sensor data to the cloud computing system, the evaluation result being evaluated using the confidence score data and the query threshold data; When the evaluation result indicates that the query is not transmitted to the cloud computing system, assigning the local prediction data as the prediction result; When the evaluation result indicates that the query is being transmitted to the cloud computing system, assigning cloud prediction data as a prediction result, the cloud prediction data being received from the cloud computing system in response to the query; as well as Use the prediction results to control edge devices.

16. The one or more non-transitory computer-readable media of claim 15, wherein: The evaluation result is generated by evaluating an inequality, wherein the inequality is f(Conf,L network )<Thres, in Conf is a number representing the confidence score data, L network is a number greater than zero and represents the average round-trip network delay between the edge device and the cloud computing system, f(Conf,L network ) is based on Conf and L network is a non-negative monotonically increasing function of the input data, and Thres indicates query threshold data.

17. The one or more non-transitory computer-readable media of claim 16, wherein: When the inequality is satisfied and evaluated to be true, the evaluation result indicates querying the cloud computing system; as well as When the inequality is not satisfied and is evaluated to be false, the evaluation result indicates that the cloud computing system is not queried.

18. The one or more non-transitory computer-readable media of claim 16, wherein: f(Conf,L network )=Conf+w·L network , Where w is a number representing a weighting factor indicating the relative network sensitivity.

19. The one or more non-transitory computer-readable media of claim 15, wherein: The cloud prediction data is generated by a cloud machine learning model that is remote from the edge device and is part of a cloud computing system; and Cloud machine learning models are larger and more accurate than local machine learning models on edge devices.

20. The one or more non-transitory computer-readable media of claim 15, wherein: When the evaluation result indicates that the query is not transmitted to the cloud computing system, the sensor data is not transmitted to the cloud computing system.