Universal inference engine for AI / ML models compatible with cloud, local and edge deployments
Through the dynamic runtime mapping and activation mechanism of the general inference engine, the problem of high resource consumption on exclusive containers is solved, and multiple models are efficiently hosted and managed on edge devices.
Patent Information
- Application Number
- CN202510079239.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-18
- Filing Date
- 2025-01-17
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, AI/ML models running on exclusive docker images and containers lead to high resource consumption and difficult to manage, especially when deploying at the edge, it is difficult to orchestrate thousands of models.
Using a general inference engine, the common inference engine is used to dynamically map and activate the runtime, share the common runtime, reduce the resource requirements of each model, and utilize the MLFlow framework and ONNX/PMML transformation model to dynamically generate the runtime to support the inference of multiple models.
Reduces computing resources and hosting costs, provides the ability to run multiple models on edge devices, and improves resource utilization efficiency.
Smart Images

Figure CN120338092A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an inference engine. Specifically, the present disclosure relates to a general inference engine for different types of models. Background Art
[0002] Generally, artificial intelligence / machine learning (AI / ML) models run on exclusive docker images and / or containers that embed the runtime environment required to run a specific model. Some industrial use cases require the simultaneous execution of thousands of AI / ML models. In such cases, maintaining and serving models using a single-container approach would be very difficult and expensive for hosting and maintenance. Additionally, in the case of edge deployment, due to resource and management model limitations, it is practically impossible to orchestrate thousands of containers. Summary of the Invention
[0003] A first aspect of the present disclosure provides a method for generating an inference for target data using a general inference engine, the method comprising: receiving details of a model to be inferred and target data from a client device; determining whether a runtime associated with the model is available in the general inference engine; based on determining that the runtime associated with the model is available in the general inference engine, determining whether the runtime has already been executed in a virtual environment of the general inference engine; based on determining that the runtime has already been executed in the virtual environment of the general inference engine, loading the model in the runtime; providing the target data as an input to the model; receiving an inference based on the target data as an output from the model; and providing the inference to the client device.
[0004] According to an implementation of the first aspect, determining whether a runtime associated with a model is available in the general inference engine includes: retrieving identification information associated with the model; using the identification information to determine whether the model received from the client requires an Open Neural Network Exchange (ONNX) or Predictive Model Markup Language (PMML) runtime; and based on determining that the model received from the client requires an ONNX or PMML runtime, retrieving a preset runtime as the runtime.
[0005] According to an implementation of the first aspect, based on determining that the model received from the client does not require an ONNX or PMML runtime, the method further includes: using the identification information to determine whether the model exists in a model tracking database for the model; based on determining that the model exists in the model tracking database, retrieving the stored runtime associated with the model from an artifact repository as the runtime.
[0006] According to an implementation of the first aspect, based on determining that the runtime associated with the model is not available in the general inference engine, the method further includes: retrieving identification information associated with the model, where the identification information includes a list of libraries and dependencies required for the execution of the model; dynamically generating a runtime based on the list of libraries and dependencies; storing the runtime in the model runtime repository; executing the runtime in a new virtual environment of the general inference engine; and loading the model in the new virtual environment.
[0007] According to an implementation of the first aspect, based on determining that the runtime has not been executed in the virtual environment of the general inference engine, the method further includes: executing the runtime in a new virtual environment of the general inference engine; and loading the model in the new virtual environment.
[0008] According to an implementation of the first aspect, the method further includes: determining whether other models are being executed in the virtual environment; and deactivating the virtual environment based on determining that no other models are being executed in the virtual environment.
[0009] According to the first aspect, based on determining that other models are being executed in the virtual environment, the method further includes: deactivating the virtual environment after the other models have completed execution.
[0010] According to an implementation of the first aspect, determining whether the runtime associated with the model is available in the general inference engine includes: retrieving identification information associated with the model; using the identification information to determine whether the model received from the client requires a MATLAB or R runtime; and based on determining that the model received from the client requires a MATLAB or R runtime, retrieving a preset runtime as the runtime.
[0011] An implementation of the second aspect of the present disclosure provides a system for generating inferences for target data using a general inference engine. The system includes: a controller configured to: receive details of the model to be inferred and the target data from a client device; determine whether the runtime associated with the model is available in the general inference engine; based on determining that the runtime associated with the model is available in the general inference engine, determine whether the runtime has been executed in the virtual environment of the general inference engine; based on determining that the runtime has been executed in the virtual environment of the general inference engine, load the model in the runtime; provide the target data as an input to the model; receive the inference based on the target data as an output from the model; and provide the inference to the client device.
[0012] According to an implementation of the second aspect, determine whether the runtime associated with the model is available in the general inference engine, and cause the controller to: retrieve the identification information associated with the model; use the identification information to determine whether the model received from the client requires an Open Neural Network Exchange (ONNX) or Predictive Model Markup Language (PMML) runtime; and based on determining that the model received from the client requires an ONNX or PMML runtime, retrieve a preset runtime as the runtime.
[0013] According to the implementation of the second aspect, based on determining that the model received from the client does not require an ONNX or PMML runtime, the controller is further configured to: use the identification information to determine whether the model exists in a model tracking database for the model; based on determining that the model exists in the model tracking database, retrieve the stored runtime associated with the model from the artifact repository as the runtime.
[0014] According to the implementation of the second aspect, based on determining that the runtime associated with the model is not available in the general inference engine, the controller is further configured to: retrieve the identification information associated with the model, where the identification information includes a list of libraries and dependencies required for model execution; dynamically generate a runtime based on the list of libraries and dependencies; store the runtime in a model runtime repository; execute the runtime in a new virtual environment of the general inference engine; and load the model in the new virtual environment.
[0015] According to the implementation of the second aspect, based on determining that the runtime has not been executed in the virtual environment of the general inference engine, the controller is further configured to: execute the runtime in a new virtual environment of the general inference engine; and load the model in the new virtual environment.
[0016] According to the implementation of the second aspect, the controller is further configured to: determine whether other models are being executed in the virtual environment; and based on determining that no other models are being executed in the virtual environment, deactivate the virtual environment.
[0017] According to the implementation of the second aspect, based on determining that other models are being executed in the virtual environment, the method further includes: after the other models have completed execution, deactivate the virtual environment.
[0018] According to the implementation of the second aspect, determine whether the runtime associated with the model is available in the general inference engine, and cause the controller to: retrieve the identification information associated with the model; use the identification information to determine whether the model received from the client requires a MATLAB or R runtime; and based on determining that the model received from the client requires a MATLAB or R runtime, retrieve a preset runtime as the runtime. Description of the Drawings
[0019] The subject matter of the present disclosure will now be described in more detail based on exemplary figures. All features described and / or shown herein can be used alone or in different combinations. By reading the following detailed description with reference to the figures, the features and advantages of various embodiments will become apparent, and the figures illustrate the following:
[0020] Figure 1 A simplified block diagram depicting a system for registering and deploying a model, according to one or more examples of the present disclosure;
[0021] Figure 2 A simplified block diagram depicting a system for determining inferences associated with a model, according to one or more examples of the present disclosure;
[0022] Figure 3 A simplified block diagram of a general inference engine, according to one or more examples of the present disclosure;
[0023] Figures 4 to 6 An exemplary process for invoking various models using general inference, according to one or more examples of the present disclosure;
[0024] Figure 7 An exemplary process for invoking a general inference engine on an edge device, according to one or more examples of the present disclosure;
[0025] Figure 8 A block diagram of an exemplary controller associated with a general inference engine, according to one or more examples of the present disclosure; and
[0026] Figure 9 An exemplary process for invoking a model using a general inference engine, according to one or more examples of the present disclosure. Detailed Description
[0027] Examples of the presented application will now be described more fully hereinafter with reference to the figures, in which some but not all examples of the application are shown. In fact, the present application may be embodied in different forms and should not be construed as limited to the examples set forth herein; rather, these examples are provided so that the application will meet applicable legal requirements. Wherever possible, any term expressed in the singular herein is intended to also include the plural and vice versa, unless otherwise expressly stated. Additionally, as used herein, the term "a" and / or "an" shall mean "one or more," even if the phrase "one or more" is also used herein. Further, when something is said to be "based on" something else herein, it may also be based on one or more other things. In other words, unless otherwise expressly indicated, "based on" as used herein means "at least in part based on" or "at least partially based on."
[0028] Traditionally, artificial intelligence / machine learning (AI / ML) models run on exclusive docker images and / or containers that embed the runtime environment required to run a specific model. This disclosure provides a method for executing an AI / ML model that eliminates the need to create a separate docker image for each AI / ML model by introducing "general inference engine" logic. The method is based on the dynamic mapping, activation, and deactivation of the model runtime during inference.
[0029] For example, if there are 1,000 different AI / ML models (non-Open Neural Network Exchange (ONNX) type), 1,000 separate docker images and / or containers need to be orchestrated to execute each different model. Even assuming that each docker image / container consumes 10 millicores of CPU and 200 megabytes (Mi) of memory in the idle state, just running the different containers / docker images in the idle state may require a total infrastructure of 10,000 millicores of CPU and 200,000 Mi of memory.
[0030] According to an embodiment of the present disclosure, the general inference engine described in the present disclosure uses a 10-replica setup to achieve the hosting and concurrent inference of 1,000 models, thereby reducing the resources required in operation. For example, the general inference engine mainly works based on the concept of model runtime (virtual environment) specific to the model, where the runtime associated with each model is stored when the model is registered. In cases where the runtimes of different models are common, the common runtime can be shared across multiple models to reduce resource consumption during execution.
[0031] The general inference engine is configured to dynamically activate the runtime, call the model within the runtime, and deactivate the runtime, thereby eliminating the need to run a single container (or docker image) for each Al / ML model. The method utilizes multiple frameworks, such as the Machine Learning Flow (MLFlow) framework, ONNX, Predictive Model Markup Language (PMML), wherever applicable to convert the model into a lightweight, cross-platform compatible model, and utilizes the runtime generated during operation to call the model. This saves computing resources, the cost of hosting a large number of models, and provides the ability to run multiple models on edge devices (such as tablet computers and mobile phones).
[0032] Figure 1 A simplified block diagram is shown depicting a system for registering and deploying models according to one or more examples of the present disclosure. Figure 1The system 100 shown includes a model input 102, a repository 104, and MLFlow 106. The model input 102 includes model details related to a new AI / ML model received from a model architecture graphical user interface. In some embodiments, the system 100 is used to register a new AI / ML model before a general inference engine can utilize the model to perform inference. In Figure 1 the system, the model input 102 includes model metadata 102a, a model file 102b, and a PIP list 102c. In some cases, the PIP list 102c can be a package management system written in a programming language (e.g., Python) and is used to install and manage packages. For example, the PIP list 102c can include a list of libraries required to execute the new AI / ML model. The model metadata 102a includes metadata associated with the new AI / ML model, the model file 102b includes program files and / or executable files associated with the new AI / ML model, and the PIP list 102c includes a list of libraries required to execute the new AI / ML model. The model input 102 including the model metadata 102a, the model file 102b, and the PIP list 102c is stored in the artifact repository 106a.
[0033] The artifact repository 106a is part of the MLFlow framework 106. In addition to the artifact repository 106a, the MLFlow framework 106 also includes a tracking database 106b. Identification information about the new AI / ML model, including name, version, and deployment tracking, can be stored in the tracking database 106b. The tracking database includes identification information for each model registered using the system 100, which can be accessed by the general inference engine. In some embodiments, the information associated with the new model stored in the tracking database 106b is extracted from the model input 102 stored in the artifact repository 106a. In some cases, when a request is received from a client, the system can search the tracking database for the requested model using the identification information including name, version, and deployment tracking.
[0034] System 100 also includes a repository 104, and the repository 104 includes a model runtime repository 104b. The model runtime repository 104b stores runtimes associated with the received new AI / ML model. In some embodiments, a PIP list 102c of the model input 102 is used to create a runtime associated with the new AI / ML model. For example, the artifact repository 106a may provide the PIP list 102c associated with the new AI / ML model to the model runtime repository 104b. In such cases, the PIP list 102c may be provided by the artifact repository 106a to the runtime repository 104b as a PIP artifact feed 104a. In some cases, the PIP artifact feed 104a may be part of the artifact repository 106a. Using the information in the PIP artifact feed 104a, a runtime is created for the new AI / ML model. For example, the PIP artifact feed 104a may include a list of libraries required to execute the new AI / ML model. The libraries and the associated dependency list are installed in the runtime associated with the new AI / ML model. The created runtime can be used to deploy the AI / ML model when needed. The runtime is stored in the model runtime repository 104b for later use. The information associated with the new AI / ML model stored in the artifact repository 106a is updated to include a reference to the new runtime created for the new AI / ML model stored in the model runtime repository 104b. For example, the model metadata 102a associated with the new AI / ML model stored in the artifact repository 106a is updated with the model information. In some cases, the model information includes the identification information of the runtime associated with the new AI / ML model stored in the model runtime and deployment status.
[0035] Figure 2 FIG. shows a simplified block diagram depicting a system for determining inferences associated with a model in accordance with one or more examples of the present disclosure. Figure 2 System 200 depicts a client 202, a general inference engine 204, MLFlow 106, and a model runtime repository 104b. In some embodiments, the model runtime store 104b is mapped to individual containers (or docker images / pods) generated by the general inference engine 204. Each of the containers in the general inference engine 204 continuously synchronizes with the model runtime repository 104b and extracts the local virtual runtime within the container. In some examples, the general inference engine 204 may check for updates in the model runtime memory 104b to look for new runtimes that can be saved in the model runtime repository 104. In the case where there is a new runtime in the model runtime repository 104b, the new runtime can be copied to the general inference engine 204. A built-in scheduler may be embedded within the general inference engine 204 to synchronize at regular intervals.
[0036] The client can request a specific model to be deployed to perform inference on a target dataset. The client's request can be transmitted from the client device 202 to the general inference engine 204. The request from the client device can include details of the specific model to be inferred and the target data on which the inference should be performed. The general inference engine 204 can invoke the specific model requested by the client device 202 from MLFlow 106. In some embodiments, the general inference engine 204 can be provided with information identifying the specific model. The general inference engine can use the identification information of the specific model to search for the specific model in the MLFlow framework 106. For example, the general inference engine 204 can use the identification information of the specific model to search for the model in the tracking database 106b. When determining that the specific model exists in the tracking database 106b, the general inference engine 204 can retrieve information related to the specific model in the artifact repository 106a. In some embodiments, the information related to the specific model retrieved from the artifact repository 106a includes model metadata 102a, model files 102b, and a PIP list 102c. The general inference engine 204 can also retrieve the runtime associated with the specific model from the model runtime repository 104b. The model information retrieved from the artifact repository 106a and the runtime retrieved from the model runtime repository 104b are provided to the general inference engine 204. In some embodiments, the general inference engine 204 can determine whether an instance of the runtime of the specific model provided by the model runtime repository 104b is already executing. For example, the general inference engine 204 can use the deployment status of the runtime stored together with the runtime in the model runtime repository 104b to determine whether an instance of the runtime of the specific model is being executed. In the case where the general inference engine 204 determines that an instance of the runtime has already been executed, the general inference engine 204 can continue to load the specific model in the already executed instance of the runtime. In this way, the general inference engine 204 does not have to load a new runtime for each model, and the general inference engine 204 can save processing resources.
[0037] In some embodiments, when the general inference engine 204 determines that there are no executed instances of a specific model at runtime, the runtime provided by the model runtime repository 104b can be implemented in a container and a specific model can be loaded in the runtime based on information associated with the model retrieved from the artifact repository 106b. Once the specific model is loaded in the runtime, the target data provided by the client device 202 to the general inference engine 204 can be provided as input to the specific model to generate an inference. Once the inference is generated, the general inference engine 204 provides the inference to the client device 202. After the inference is generated, the general inference engine 204 can deactivate the runtime. In some embodiments, the general inference engine 204 can deactivate the runtime only if no other model is being executed in the runtime.
[0038] Figure 3 A simplified block diagram of a general inference engine is shown in accordance with one or more examples of the present disclosure. Figure 3 System 300 of Figure 2 The detailed layout of the general inference engine 204 is described. The general inference engine 204 includes some preset runtimes for specific types of AI / ML models. For example, 306 represents the runtime stored in the general inference engine 204 associated with an ONNX model. The ONNX model can be of different types, such as Sklearn, TensorFlow, Keras, Pytorch, or others. When the client device 202 requests an ONNX type model to perform inference on target data, the general inference engine 204 can dynamically generate a runtime to load the model and generate an inference. In some embodiments, the general inference engine can retrieve a preset ONNX runtime for loading the model. For example, a preset runtime can be dynamically generated using the preset information stored in the general inference engine 204 associated with the ONNX runtime. Similarly, 308 represents the runtime stored in the general inference engine 204 associated with a PMML model. The PMML model can be of different types, such as ARIMA / ARIMAx, SARIMA / SARIMAx, OLS, exponential smoothing, or others. When the client device 202 requests a PMML type model to perform inference on target data, the general inference engine 204 can dynamically generate a runtime to load the model and generate an inference. In some embodiments, the general inference engine can retrieve a preset PMML runtime for loading the model. For example, a preset runtime can be dynamically generated using the preset information stored in the general inference engine 204 associated with the PMML runtime.
[0039] In some cases, there may be other model types that can be supported in the general inference engine 204.
[0040] On the other hand, in the case where the client device 202 requests to use a model that does not have a preset runtime stored in the general inference engine 204, the general inference engine 204 generates a virtual environment runtime for the model. The process of generating the virtual environment runtime is described in 310. At 310a, the general inference engine 204 receives from the client device 202 a request to use a model that is not part of the general inference engine 204. At 310b, the general inference engine 204 activates the virtual environment of the model. When activating the virtual environment, the general inference engine 204 can first search the trace database 106b for the model as described regarding Figure 2 above. In the case where the model has been previously requested, it can be registered in the trace database 106b, and the runtime associated with the model can exist in the model runtime repository 104b. In the case where the model exists in the trace database, the runtime associated with the model can be retrieved from the model runtime repository 104b and executed in the virtual environment of the general inference engine 204. The model can be loaded at runtime. In some embodiments, the artifact mount 302 of the general inference engine 204 can include a compressed file containing the runtime that exists in the model runtime repository 104b.
[0041] When the runtime is executed in the virtual environment, the general inference engine 204 can extract the files necessary for the runtime from the artifact mount 302. The artifact mount 302 can be configured to synchronize with the virtual environment of the general inference engine 204 at periodic intervals (e.g., every 10 seconds, 30 seconds, 60 seconds, etc.) based on user and / or system requirements. Once the runtime is executed, the general inference engine 204 can retrieve the model to be used for inference from MLFlow 106.
[0042] In the case where the trace database does not include the model, Figure 1The process described in registers the model in the tracking database 106b. Once the model is registered, information related to the model can be stored in the artifact repository 106a. The runtime associated with the model can be generated using the information from the artifact repository 106a and stored in the model runtime repository 104b. As disclosed above, the runtime associated with the model can include all the libraries and dependencies required for the model preloaded in the runtime to execute the model faster. Then, the runtime can be generated by the general inference engine 204 and the model can be loaded in the runtime. At 310c, inferences can be generated using the model loaded in the runtime. Inferences can be generated by providing the target data received from the client device 202 as the input to the model. The generated inferences can be provided to the client device 202. At 310d, after generating the inferences, the general inference engine 204 can deactivate the virtual environment that executes the runtime using the model. In some embodiments, deactivating the virtual environment involves determining whether any other inference requests are being executed in the virtual environment. In the case where other inference requests are being executed in the virtual environment, the virtual environment can be deactivated after all the inferences requests being executed are completed. In the case where no other inference requests are being executed in the virtual environment, the virtual environment can be deactivated.
[0043] In some embodiments, the client device 202 can request the MATLAB or R model to perform inferences on the target data. The process of generating inferences using the R or MATLAB model is described in process 312. At 312a, the general inference engine 204 receives a request to use the R or MATLAB model from the client device 202. At 312b, the general inference engine 204 can send a service request to the container associated with the model. At 312c, inferences associated with the model are received from the pod associated with the model, and at 312d, the inferences associated with the model are reported to the client device 202.
[0044] Figures 4 to 7 An exemplary process for invoking various models using a general inference engine according to one or more examples of the present disclosure is shown.
[0045] Figure 4 An exemplary process for invoking an ONNX or PMML runtime using a general inference engine according to one or more examples of the present disclosure is shown. Process 400 can be executed by Figure 3 the general inference engine 204. However, it should be recognized that any of the following boxes can be executed in any suitable order, and process 400 can be executed in any environment and by any suitable computing device and / or controller.
[0046] At 402, a model inference request is received from the client device 202 at the general inference engine 204. In some embodiments, the model inference request may include target data and a model on which the target data should operate. For example, the model inference request may include an ONNX or PMML model to generate an inference for the target data.
[0047] At 404, a model inference is determined within an existing runtime. For example, as disclosed with respect to Figure 3 an ONNX or PMML model, the runtime associated with the ONNX or PMML model has been stored in the general inference engine 204. In such cases, preset information stored in the general inference engine 204 associated with the ONNX or PMML model can be used to dynamically generate a runtime. Once the runtime is deployed, the PMML or ONNX model is loaded into the runtime for execution. The general inference engine 204 executes the PMML or ONNX model using the target data received from the client device 202 in the runtime and generates an inference.
[0048] At 406, the general inference engine records inference data for monitoring based on a log flag. For example, the log flag can be used to enable data logging for model monitoring purposes. In the case where the log flag is set to false, data is not recorded in the system for monitoring. A log flag may be required to distinguish model evaluation / test data from actual production data that needs to be monitored.
[0049] At 408, the general inference engine 204 provides the determined inference to the client device 202.
[0050] Figure 5 An exemplary process for invoking a virtual environment runtime using a general inference engine is shown in accordance with one or more examples of the present disclosure. Process 500 may be performed by Figure 3 the general inference engine 204. However, it should be recognized that any of the following blocks may be performed in any suitable order, and process 500 may be performed in any environment and by any suitable computing device and / or controller.
[0051] At 502, a model inference request is received from the client device 202 at the general inference engine 204. In some embodiments, the model inference request may include target data and a model on which the target data should operate. For example, the model inference request may include a model that does not have a preset runtime stored in the general inference engine 204 to generate an inference for the target data.
[0052] At 504, the general inference engine extracts model runtime details. In some embodiments, extracting model runtime details may include determining from the model whether a runtime already exists. To determine whether a runtime already exists for the model, the general inference engine 204 may determine whether the model is registered in the tracking database 106b. In the case where the model is registered in the tracking database 106b, the runtime associated with the model may exist in the model runtime repository 104b. In the case where the model exists in the tracking database, the runtime associated with the model may be retrieved from the model runtime repository 104b. In the case where the tracking database does not include the model, the Figure 1 process described in may be used to register the model in the tracking database 106b. Once the model is registered, information related to the model may be stored in the artifact repository 106a. The runtime associated with the model may use the information from the artifact repository 106a to be generated and stored in the model runtime repository 104b.
[0053] At 506, the general inference engine 204 activates the runtime in the relevant subprocess. For example, once the correct runtime model is retrieved, the general inference engine 204 may execute the runtime in a virtual environment.
[0054] At 508, the general inference engine 204 executes a scoring script and invokes the model in the subprocess. For example, the model may be loaded into the runtime for execution. Inference may be generated by providing the target data received from the client device 202 as input to the runtime model. In some embodiments, the scoring script is a script for preprocessing the input data and sending it to the model for prediction. Once the prediction is complete, the scoring script may also be responsible for postprocessing the prediction data and returning the final output to the user. In some other embodiments, the subprocess may be considered a thread / kernel that is dynamically instantiated when a model inference request is received from the client device 202. The virtual environment may be dynamically activated, the model may be inferred, and then the thread / kernel may be ended after the inference is complete.
[0055] At 510, the general inference engine 204 extracts the model prediction back to the main process.
[0056] At 512, the general inference engine 204 records inference data for monitoring based on log tags.
[0057] At 514, the general inference engine 204 provides the determined inference to the client device 202.
[0058] Figure 6 An exemplary process for using a general inference engine to invoke an R or MATLAB runtime according to one or more examples of the present disclosure is shown. Process 600 may be performed by Figure 3It is executed by the general inference engine 204. However, it should be recognized that any of the following boxes can be executed in any suitable order, and process 600 can be executed in any environment and by any suitable computing device and / or controller.
[0059] At 602, a model inference request is received from the client device 202 at the general inference engine 204. In some embodiments, the model inference request may include target data and the model on which the target data should run. For example, the model inference request may include an R or MATLAB model to generate inferences for the target data.
[0060] At 604, the general inference engine 204 sends a service call to the corresponding R / MATLAB model. In some embodiments, the R and MATLAB models may be hosted as separate container images within a cluster. For example, one container for each R / MATLAB model may be stored in the general inference engine 204. The general inference engine 204 can identify the appropriate model container based on the request, call the model prediction from the model container, and return the final prediction to the end user.
[0061] At 606, the general inference engine 204 extracts the response and returns it to the main process.
[0062] At 608, the general inference engine 204 records inference data for monitoring based on log tags.
[0063] At 610, the general inference engine 204 provides the determined inference to the client device 202. For example, the general inference engine 204 can act to coordinate the process of sending the input to the appropriate container and returning the response to the user.
[0064] Figure 7 An exemplary process for invoking a general inference engine on an edge device according to one or more examples of the present disclosure is shown. Process 700 may be executed by Figure 3 the general inference engine 204. However, it should be recognized that any of the following boxes can be executed in any suitable order, and process 700 can be executed in any environment and by any suitable computing device and / or controller.
[0065] In some embodiments, the general inference engine 204 is compatible with running on edge devices such as mobile phones, tablet computers, or other devices. When running on an edge device, the general inference engine 204 can be configured to invoke inferences for models pushed from the cloud to the edge. In these examples, on an edge device with more models, the resources required to invoke multiple models hosting and running different types of runtimes can be utilized.
[0066] At 702, a model that can be used to perform inference is stored in the cloud.
[0067] At 704, the general inference engine 204 on the edge device can request a specific model from the model stored in the cloud.
[0068] At 706, the requested model can be moved from the cloud to the edge device. In addition to the requested model, the runtime associated with the requested model can also be sent to the edge device for execution.
[0069] For edge export, the runtime and the model repository are packaged and shipped to the relevant edge device, where the general inference engine 204 can run by default. On the edge device, the model and the runtime are extracted, and the execution of the model is performed.
[0070] Figure 8 is a block diagram of an exemplary system or device 800 within a system 200 such as the general inference engine 204. The system 800 includes a processor 804, such as a central processing unit (CPU) and / or logic, which executes computer-executable instructions for performing the functions, processes, and / or methods described herein. In some examples, the computer-executable instructions are stored locally and accessed from a non-transitory computer-readable medium such as a storage device 810, which can be a hard disk drive or a flash drive. The read-only memory (ROM) 806 includes computer-executable instructions for initializing the processor 804, and the random access memory (RAM) 808 is the main memory for loading and processing the instructions executed by the processor 804. The network interface 812 can be connected to a wired network or a cellular network and a local area network or a wide area network. The system 800 can also include a bus 802 that connects the processor 804, the ROM 806, the RAM 808, the storage device 810, and / or the network interface 812. Components within the system 800 can communicate with each other using the bus 802. The components within the system 800 are merely exemplary and may not include every component within the controller 104. Additionally, and / or alternatively, the system 800 can also include components that may not be included in every instance of the system 100. For example, in some examples, the general inference engine 204 may not include the network interface 812.
[0071] Figure 9 Illustrates an exemplary process for invoking a model using a general inference engine according to one or more examples of the present disclosure. The process 900 can be performed by Figure 3 the general inference engine 204. However, it should be recognized that any of the following blocks can be performed in any suitable order, and the process 900 can be performed in any environment and by any suitable computing device and / or controller.
[0072] At 902, the general inference engine 204 receives a model and target data from the client device 202.
[0073] At 904, the general inference engine 204 determines whether a runtime associated with the model is available in the general inference engine 204. For example, determining whether a runtime associated with the model is available in the general inference engine 204 includes determining whether the model requires an ONNX runtime, or a PMML runtime, or a MATLAB runtime. The general inference engine 204 includes preset runtimes corresponding to each of these runtimes. In the case where the model requires an ONNX runtime, a PMML runtime, or a MATLAB runtime, the general inference engine 204 can retrieve the preset runtimes corresponding to these runtimes. In some embodiments, the general inference engine 204 can also search a tracking database to determine whether the model requested to be used by the user has been registered with the general inference engine 204. In the case where the model is registered, the runtime associated with the model can be stored in the model runtime repository 104b. In this case, the runtime can be retrieved from the model runtime repository using the identification information retrieved from the model.
[0074] In response to determining that the runtime associated with the model is available in the general inference engine 204, process 900 moves to 906 to determine whether the runtime is being executed in a virtual environment.
[0075] In response to determining that the runtime associated with the model is not available in the general inference engine 204, process 900 moves to 908 to retrieve the identification information associated with the model.
[0076] At 906, the general inference engine 204 determines whether the runtime is being executed in a virtual environment. In response to determining that the runtime is being executed in a virtual environment, process 900 moves to 914 to load the model in the executing runtime. In response to determining that the runtime is not being executed, process 900 moves to 910, where the general inference engine 204 starts a virtual environment to execute the runtime.
[0077] At 908, the general inference engine 204 retrieves the identification information associated with the model. For example, the identification information associated with the model can include a list of libraries and dependencies required to execute the model.
[0078] At 912, the general inference engine 204 dynamically generates a runtime in the virtual environment based on the information retrieved from the model. For example, dynamically generating the runtime can include installing the libraries and dependency functions required for the model in the virtual environment. The dynamically generated runtime can also be stored in the model runtime repository 104b for future reference.
[0079] At 914, the general inference engine 204 loads the runtime model executed in the virtual environment.
[0080] At 916, the general engine 204 provides the target data received from the client device 202 as input to the model. The model analyzes the target data to generate inferences.
[0081] At 918, the determined inferences are provided to the client device 202.
[0082] Although the subject matter of the present disclosure has been described and illustrated in detail in the accompanying drawings and the foregoing description, such description and illustration should be considered illustrative or exemplary and not restrictive. Any statement characterizing the invention made herein is also considered illustrative or exemplary and not restrictive, since the invention is defined by the claims. It should be understood that those of ordinary skill in the art can make changes and modifications within the scope of the appended claims, and the appended claims may include any combination of features from the different embodiments described above.
[0083] The terms used in the claims should be construed to have the broadest reasonable interpretation consistent with the foregoing description. For example, the use of the articles "a" or "the" when introducing an element should not be construed as excluding a plurality of elements. Similarly, the phrase "or" should be construed as inclusive, such that the phrase "A or B" does not exclude "A and B", unless it is clear from the context or the foregoing description that only one of A and B is meant. In addition, the phrase "at least one of A, B, and C" should be construed as one or more of the group of elements consisting of A, B, and C, and should not be construed as requiring at least one of each of the listed elements A, B, and C, regardless of whether A, B, and C are related as a category or otherwise. In addition, the phrases "A, B, and / or C" or "at least one of A, B, or C" should be construed to include any single entity from the listed elements, such as A, any subset from the listed elements, such as A and B, or the entire list of elements A, B, and C.
Claims
1. A method for generating inferences for target data using a general inference engine, the method comprising: Receiving details of a model to be inferred and target data from a client device; Determining whether a runtime associated with the model is available in the general inference engine; Based on determining that the runtime associated with the model is available in the general inference engine, determining whether the runtime has been executed in a virtual environment of the general inference engine; Based on determining that the runtime has been executed in a virtual environment of the general inference engine, loading the model in the runtime; Providing the target data as an input to the model; Receiving an inference based on the target data as an output from the model; And Providing the inference to the client device.
2. The method according to claim 1, wherein determining whether the runtime associated with the model is available in the general inference engine comprises: Retrieving identification information associated with the model; Using the identification information to determine whether the model received from the client requires an Open Neural Network Exchange (ONNX) or a Predictive Model Markup Language (PMML) runtime; And Based on determining that the model received from the client requires the ONNX or the PMML runtime, retrieving a preset runtime as the runtime.
3. The method according to claim 3, wherein based on determining that the model received from the client does not require the ONNX or the PMML runtime, the method further comprises: Using the identification information to determine whether the model exists in a model tracking database for the model; Based on determining that the model exists in the model tracking database, retrieving a stored runtime associated with the model from an artifact repository as the runtime.
4. The method according to claim 4, wherein based on determining that the runtime associated with the model is not available in the general inference engine, the method further comprises: Retrieving identification information associated with the model, wherein the identification information includes a list of libraries and dependencies required for execution of the model; Dynamically generating the runtime based on the list of libraries and dependencies; Storing the runtime in a model runtime repository; Executing the runtime in a new virtual environment of the general inference engine; And Loading the model in the new virtual environment.
5. The method according to claim 1, wherein based on determining that the runtime has not been executed in a virtual environment of the general inference engine, the method further comprises: Executing the runtime in a new virtual environment of the general inference engine; And Loading the model in the new virtual environment.
6. The method according to claim 1, further comprising: Determining whether other models are being executed in the virtual environment; And Based on determining that no other models are being executed in the virtual environment, deactivating the virtual environment.
7. The method according to claim 6, wherein based on determining that other models are being executed in the virtual environment, the method further comprises: After the other model has finished execution, deactivate the virtual environment.
8. The method according to claim 1, wherein determining whether the runtime associated with the model is available in the general inference engine includes: Retrieve the identification information associated with the model; Use the identification information to determine whether the model received from the client requires a MATLAB or R runtime; And Based on determining that the model received from the client requires a MATLAB or R runtime, retrieve a preset runtime as the runtime.
9. A system for generating inferences for target data using a general inference engine, the system comprising: A controller configured to: Receive details of a model to be inferred and target data from a client device; Determine whether the runtime associated with the model is available in the general inference engine; Based on determining that the runtime associated with the model is available in the general inference engine, determine whether the runtime has been executed in the virtual environment of the general inference engine; Based on determining that the runtime has been executed in the virtual environment of the general inference engine, load the model in the runtime; Provide the target data as input to the model; Receive the inference based on the target data as output from the model; And Provide the inference to the client device.
10. The system according to claim 9, wherein determining whether the runtime associated with the model is available in the general inference engine causes the controller to: Retrieve the identification information associated with the model; Use the identification information to determine whether the model received from the client requires an Open Neural Network Exchange (ONNX) or Predictive Model Markup Language (PMML) runtime; And Based on determining that the model received from the client requires the ONNX or the PMML runtime, retrieve a preset runtime as the runtime.
11. The system according to claim 10, wherein based on determining that the model received from the client does not require the ONNX or the PMML runtime, the controller is further configured to: Use the identification information to determine whether the model exists in a model tracking database for the model; Based on determining that the model exists in the model tracking database, retrieve the stored runtime associated with the model from an artifact repository as the runtime.
12. The system according to claim 11, wherein based on determining that the runtime associated with the model is not available in the general inference engine, the controller is further configured to: Retrieve the identification information associated with the model, wherein the identification information includes a list of libraries and dependencies required for the execution of the model; Dynamically generate the runtime based on the list of libraries and dependencies; Store the runtime in a model runtime repository; Execute the runtime in a new virtual environment of the general inference engine; And Load the model in the new virtual environment.
13. The system according to claim 9, wherein based on determining that the runtime has not been executed in the virtual environment of the general inference engine, the controller is further configured to: Execute the runtime in a new virtual environment of the general inference engine; and Load the model in the new virtual environment.
14. The system according to claim 9, wherein the controller is further configured to: Determine whether other models are being executed in the virtual environment; and Deactivate the virtual environment based on determining that no other models are being executed in the virtual environment.
15. The system according to claim 14, wherein based on determining that other models are being executed in the virtual environment, the system is further configured to: Deactivate the virtual environment after the other models have completed execution.
16. The system according to claim 9, wherein determining whether the runtime associated with the model is available in the general inference engine causes the controller to: Retrieve the identification information associated with the model; Use the identification information to determine whether the model received from the client requires MATLAB or R runtime; and And Based on determining that the model received from the client requires MATLAB or R runtime, retrieve a preset runtime as the runtime.