Proxy System and Method for a Multiprocessing Architecture
The proxy computing system addresses the challenge of resource management in multi-client/multi-processor architectures by dynamically assigning processing units based on load states, enhancing system performance and efficiency.
Patent Information
- Application Number
- JP2024568125
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-17
- Filing Date
- 2023-05-17
- Publication Date
- 2025-06-17
AI Technical Summary
Existing systems face challenges in efficiently managing and allocating system resources for executing inference requests from client computing systems across multiple processing devices, particularly in multi-client/multi-processor device architectures.
A proxy computing system is implemented to receive neural network models from client computing systems, access system resource availability on multiple processing devices, select suitable devices, load models, and route inference requests and results accordingly, based on load states and resource availability.
This solution enables efficient resource management and allocation, dynamically assigning processing units to inference requests based on load states, thereby reducing system throughput slowdowns and improving overall system performance.
Smart Images

Figure 2025518513000001_ABST
Abstract
Description
Technical Field
[0001] This application claims the benefit of priority of U.S. Provisional Patent Application Serial No. 63 / 343,014, filed on May 17, 2022, entitled "SYSTEMS AND METHODS FOR MANAGING MULTIPLE MACHINE-LEARNING-SPECIFIC PROCESSORS", the disclosure of which is incorporated herein by reference in its entirety.
[0002] The present disclosure relates to systems and methods for mapping at least one client computing system and associated inference requests to one or more processing devices included in a plurality of processing devices.
Background Art
[0003] Recent advancements in artificial intelligence / machine learning technologies, as well as in processing technologies, have led to an increase in system architectures configured such that one or more inference requests from one or more client computing systems are executed by a plurality of processing devices. In the case of multi-client / multi-processor device mapping / architecture, it can be a challenge to monitor, manage, and allocate system resources on both the client computing system side and the processing device side.
Summary of the Invention
[0004] Aspects of the present invention are directed to systems and methods for implementing a proxy computing system for a multiprocessing architecture. One method includes a proxy computing system receiving a neural network model from a client computing system. The proxy computing system can access system resource availability on a plurality of processing devices and select a subset of available processing devices based on the system resource availability. The proxy computing system can load the neural network model into each of the processing devices in the subset.
[0005] In one aspect, the proxy computing system receives an inference request from a client computing system. In response, the proxy computing system accesses the load state of each of the processing devices in the subset and selects a target processing device from the subset based on the load state. The proxy computing system can send the inference request to the target processing device.
[0006] Other aspects include an apparatus for implementing a workflow associated with the above method.
[0007] Non-limiting and non-exhaustive examples of the present disclosure are described with reference to the following figures, in which like reference numerals refer to like parts throughout the various figures unless otherwise specified. BRIEF DESCRIPTION OF THE DRAWINGS
[0008]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 3C
Figure 3D
Figure 3E
Figure 4A
Figure 4B
Figure 4C
Figure 4D
Figure 4E
Figure 5A
Figure 5B
Figure 5C
Figure 6
Figure 7
DETAILED DESCRIPTION OF THE INVENTION
[0009] In the following description, reference is made to the accompanying drawings that form a part hereof, and in which are shown by way of illustration specific exemplary embodiments in which the disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the concepts disclosed herein, and it is to be understood that various modifications to the disclosed embodiments may be made and other embodiments may be utilized without departing from the scope of the disclosure. Accordingly, the following detailed description is not to be taken in a limiting sense.
[0010] Throughout this specification, references to "one embodiment", "an embodiment", "one example", or "an example" mean that a particular feature, structure, or characteristic described in connection with the embodiment or example is included in at least one embodiment of the disclosure. Thus, the appearances of the phrases "in one embodiment", "in an embodiment", "one example", or "an example" in various places throughout this specification are not necessarily all referring to the same embodiment or example. Furthermore, the particular features, structures, databases, or characteristics may be combined in any suitable combination and / or sub-combination in one or more embodiments or examples. Additionally, it should be recognized that the figures provided with this specification are for illustrative purposes for those skilled in the art and that the figures are not necessarily drawn to scale.
[0011] Embodiments in accordance with the present disclosure may be embodied as an apparatus, a method, or a computer program product. Accordingly, the present disclosure may take the form of an embodiment that is entirely hardware, an embodiment that is entirely software (including firmware, resident software, microcode, and the like), or an embodiment that combines software aspects and hardware aspects that may all be collectively referred to herein as a "circuit", "module", or "system". Furthermore, embodiments of the present disclosure may take the form of a computer program product embodied in any tangible expression medium having computer-usable program code embodied therein.
[0012] One or more computer-usable or computer-readable media in any combination may be utilized. For example, the computer-readable media can include one or more of a portable computer diskette, a hard disk, a random access memory (RAM) device, a read-only memory (ROM) device, an erasable programmable read-only memory (EPROM) or flash memory device, a portable compact disc read-only memory (CDROM), an optical storage device, a magnetic storage device, and any other storage media now known or later discovered. The computer program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages. Such code can be compiled from source code into computer-readable assembly language or machine code suitable for a device or computer that can execute the code.
[0013] The embodiments can also be implemented in a cloud computing environment. In this description and the following claims, "cloud computing" may be defined as a model for enabling ubiquitous, convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services), which can be rapidly provisioned via virtualization and released with minimal management effort or service provider interaction, and then scaled accordingly. The cloud model may be composed of various characteristics (e.g., on-demand self-service, broad network access, resource pooling, rapid elasticity, and measured service), service models (e.g., software as a service ("SaaS"), platform as a service ("PaaS"), and infrastructure as a service ("IaaS")), and deployment models (e.g., private cloud, community cloud, public cloud, and hybrid cloud).
[0014] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a special purpose hardware-based system that performs the specified function or operation, or by a combination of special purpose hardware and computer instructions. These computer program instructions may also be stored in a computer-readable medium and can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture that includes instruction means for implementing the specified function / operation in one or more blocks of the flowchart and / or block diagram.
[0015] Aspects of the present invention are directed to systems and methods for implementing an interface between one or more client computing systems and a plurality of processing units (i.e., processing devices). In one aspect, a proxy computing system implements such an interface. Such a proxy computing system can enable a client computing device to communicate with one or more processing units. The proxy computing system can also facilitate distributing or mapping one or more inference requests from the client computing system to the processing units. Any inference results generated by the processing units can be routed back to the client computing system that initiated the corresponding inference request.
[0016] FIG. 1 is a block diagram illustrating a proxy computing system interface 100. As shown, the proxy computing system interface 100 includes a proxy computing system 102, client computing systems 104 to 108, a PCIe bus interface 110, a USB interface 112, a system call interface 114, processing units 128, 130, 132, 134, 136, 138, and 140, and a simulated processing unit 142. The proxy computing system 102 further includes a storage cache 118, a device library 122, a statistics library 124, a request manager 116, and a physical layer 126. The storage cache 118 further includes a model library 120.
[0017] In one aspect, each of client computing systems 104 to 108 is a computing system including at least a processor, a memory, and a network interface. Each of client computing systems 104 to 108 can boot an operating system (e.g., Linux®, Windows®, MacOS, Unix®, etc.). Examples of computing systems include desktop computers, laptop computers, mobile computing devices such as tablets and smartphones, and others.
[0018] In one aspect, each of processing units (PUs) 128 to 140 (also described as "processing devices") is a stand-alone computing unit that includes at least a processor, a memory, and a network interface. Examples of processing devices include single-board stand-alone computing systems (e.g., ARM-based computing systems and other types of embedded processing systems). In one aspect, each of PUs 128 to 140 is configured such that one or more neural network models are loaded. These neural network models may be loaded by proxy computing system 102 from model library 120 stored in storage cache 118 onto any combination of PUs 128 to 140. Each of PUs 128 to 140 may be configured to initiate one or more inference requests associated with a particular neural network model running on each PU. These inference requests may be received from any of client computing systems 104 to 108 and routed via proxy computing system 102 to an appropriate PU. The associated PU can initiate the inference request and generate an inference result. The inference result can be routed back by proxy computing system 102 to the client computing system that sent the inference request.
[0019] In one aspect, the proxy computing system 102 interfaces with one or more PUs via interfaces such as a PCIe bus 110 (used to interface the proxy computing system 102 with PUs 128, 130, and 132), a USB interface 112 (used to interface the proxy computing system 102 with PUs 134 and 136), and one or more system calls 114 (used to interface the proxy computing system 102 with PUs 138 and 140). The PCIe bus 110, USB 112, and system call 114 of the interfaces are implemented and generated by the physical layer 126, and via an appropriate communication protocol, enable the proxy computing system 102 to interface with PUs 128 to 140. Other interfaces such as an inter - process communication (IPC) interface (not shown in FIG. 1) may be used to interface the proxy computing system with one or more PUs.
[0020] The device library 122 may be used by the proxy computing system 102 to properly interface with the PUs. The device library 122 can contain information about each device associated with the PUs. For example, if the PU is a computing board, the data associated with the PU stored in the device library 122 can include the type of processor (e.g., ARM processor, GPU), the number of computing cores, system RAM, processing unit memory state, model occupancy, and others for each PU. The proxy computing system 102 may interface with client computing systems 104 to 108 via interfaces such as USB, Ethernet®, Wi - Fi, Bluetooth, ZigBee, or any other connectivity protocol.
[0021] In one aspect, the proxy computing system 102 receives a model load request from any combination from the client computing system 104 to the client computing system 108. For example, the proxy computing system 102 can receive a model load request from the client computing system 104. The model load request may be a request to load a neural network model onto a PU (e.g., PU 128). The proxy computing system 102 can retrieve an appropriate model from the model library stored in the storage cache 118 and load the model onto one or more appropriate PUs via the associated interface.
[0022] In one aspect, the proxy computing system 102 can receive an inference request from any of the client computing systems 104 to 108. The inference request may be a request to initiate an inference operation on a specific neural network model that is running on any of the PUs 128 to 140 on which the neural network model has been previously loaded. (For example, determined by the proxy computing system 102 using the device library 122 and the statistics library 124) Based on the available resources of the PUs 128 to 140, the proxy computing system 102 can select the PU on which to initiate the inference request. The request manager 116 may be configured to route the inference request to the selected PU via the physical layer 126 and the associated communication interface (e.g., the PCIe bus 110, USB 112, or system call 114). The selected PU can initiate the inference request and generate an inference output (or inference result). The inference output / result may be sent to the proxy computing system 102 via the associated communication interface. The proxy computing system 102 can send the inference output to the client computing system that generated the inference request.
[0023] In one aspect, instead of launching an inference request on the PU, the client computing system may desire to launch a simulation of the inference request. Such a scenario may be used when a developer working on the client computing system is in the process of developing or debugging software code associated with the inference request. In this case, the proxy computing system 102 can route the inference request to the simulated processing unit 142. The simulated processing unit 142 can execute the associated inference request and send back the simulated inference result to the client computing system via the proxy computing system 102.
[0024] In one aspect, the simulated processing unit 142 is built using C++ or any other applicable programming language. The simulated processing unit 142 may be used when the hardware is not yet ready (e.g., not yet manufactured). The simulated processing unit 142 may also be used when greater observability that the hardware may not be able to provide is required (e.g., debugging an FPGA-class device). In one aspect, the simulated processing unit 142 includes one or more simulation models that represent the device behavior associated with the simulated device. Such models may have various different performances or may abstract the system (e.g., the PU) depending on the project requirements. For example, a model of a processor may choose not to be as thorough or detailed in the modeling of the memory hierarchy while choosing to mimic arithmetic calculations in fine detail. In this case, the model may model a flat memory hierarchy instead of the L1, L2, L3, DRAM levels.
[0025] Essentially, the proxy computing system 102 functions as a transparent interface (i.e., a proxy) between the client computing systems 104 to 108, the PUs 128 to 140, and the simulated processing unit 142. A client computing system may wish to load a preferred neural network model onto a PU. In this case, the proxy computing system 102 selects one or more available / suitable PUs and loads the neural network model onto that PU. The client computing system can then request to initiate an inference request using the preferred neural network model. The proxy computing system 102 can select a PU with sufficient available computing resources from among the PUs on which the preferred neural network model is running and load the inference request onto that PU. The PU can initiate the inference request and generate an inference result. The inference result is then sent back to the client computing system that sent the inference request via the proxy computing system 102.
[0026] Each of client computing systems 104 to 108 may be configured to launch application software that enables the client computing system to communicate with proxy computing system 102. The application software can include a development environment that enables software developers to develop a program that launches one or more inference requests on any combination of PUs 128 to 140 or on simulated processing unit 142. Proxy computing system 102 can function as an intermediary between client computing systems 104 to 108, PUs 128 to 140, and simulated processing unit 142. Proxy computing system 102 can route model loading (fulfillment of model load requests) and inference requests to any selected combination of PUs 128 to 140 and route inference results, statistical results, and other data back to client computing systems 104 to 108. Examples of proxy computing system 102 can support any number of client computing systems and PUs.
[0027] As an example, proxy computing system 102 may be associated with deploying artificial intelligence algorithms to launch one or more neural network models instantiated / loaded on any combination of PUs 128 to 140. These artificial intelligence algorithms may be associated with machine vision applications such as inferring, object detection, object identification, object tracking, and the like.
[0028] FIG. 2 is a block diagram illustrating a request management workflow 200 by a proxy computing system 102. The request management workflow 200 may be associated with the proxy computing system 102 receiving one or more inference requests from client computing systems 104 to 108. In one aspect, each of client computing systems 104 to 108 launches a client application. For example, client computing system 104 can launch client application 206, client computing system 106 can launch client application 208, and so on until client computing system 108 launches client application 210.
[0029] The inference requests generated by each client application may be generated as a request queue 212. Each of client applications 206 to 210 can generate its own request queue. Request queue 212 may be received by a load balancer 202 that activates a load balancing policy 204. Load balancer 202 may be implemented as a component of proxy computing system 102. Load balancing policy 204 can determine the load status from PU unit 1 206 to unit 2 218, unit N 220. PU unit 1 216 to unit N 220 may be similar to PU1 28 to 140.
[0030] Based on the load state, the load balancing policy 204 can assign and route the request queue in the request queue 212 as the endpoint execution queue 214. In one aspect, the endpoint execution queue 214 is created based on the individual load states (e.g., available resources) on each of units 1 to N 220. The endpoint execution queue 214 can include requests such as inference requests to be launched with an appropriate neural network model instantiated on each of units 1 to N 220.
[0031] In one aspect, when each of units 1 to N 220 completes the execution of its respective endpoint execution queue, each PU generates an inference output (i.e., an inference response) and sends the inference output to the response handler 222. The response handler 222 distributes the inference responses from units 1 to N 220 into the response queue 224. The response queue 224 is constructed by the response handler 222 such that each inference response is routed back to the appropriate client application from client application 206 to client application 210. In one aspect, the response queue is a set of inference responses that will be routed to a particular client application.
[0032] As an example, a request queue including an inference request from client application 206 may be routed to unit 2 218 as the endpoint execution queue. Unit 2 218 can execute the inference request and generate an inference response. The response handler 222 can receive this inference response, add this inference response to the response queue, and route it back to client application 206.
[0033] By dynamically assigning processing units to appropriate inference requests based on the load state, the proxy computing system 102 provides a flexible operating environment that reduces the slowdown (e.g., bottleneck) of system throughput.
[0034] Figures 3A, 3B, 3C, 3D, and 3E are block diagrams illustrating an inference request / response workflow 300.
[0035] Figure 3A illustrates the operation on the client computing system side, where a model 304 may be retrieved from a model parameter database 302 by a client computing system such as client computing system 104. The model 304 may be a neural network model. In one aspect, the model 304 can include an input tensor space 306 (e.g., an image or video frame) and a set of weight tensors 308 and 310. The input tensor stage 306 can be combined with the weight tensors 308 and 310 (e.g., via a weighting operation) to generate an output tensor space 312. The input tensor space 306, the weight tensors 308 and 310, and the output tensor space 312 can be grouped together and included in the definition of the model 304.
[0036] In one aspect, the client computing system can generate a model load request 314 that includes an input tensor space 316 (e.g., input tensor space 306), a set of weight tensors 318 (e.g., weight tensors 308 and 310), and an output tensor space 320 (e.g., output tensor space 310). This model load request 314 may be sent to the proxy computing system as a model load request package 321.
[0037] As shown in FIG. 3B, the proxy computing system 102 receives a model load request package 321 and assigns a model ID to the model load request 314 (process 322). The proxy computing system can also load the associated neural network model onto a selected PU (e.g., PU128) according to the PU computing load and available resources. The proxy computing system 102 can send back a load response 323 to the client computing system. The load response 323 can include a model ID 324 associated with a specific neural network model.
[0038] As shown in FIG. 3C, the client computing system receives the load response 323 and constructs an inference request 332 that includes the model ID 324. The client computing system can also construct an input tensor 330 based on the input received from any combination of the input tensor database 326 and the image sensor 328. In one aspect, the input tensor database 326 includes a sequence of one or more images or video frames. The image sensor 328 can be a camera sensor that generates a sequence of images or video frames. The input tensor 330 can be included in the inference request 332 by the client computing system. The client computing system can send the inference request 332 to the proxy computing system 102 as an inference request package 333.
[0039] As illustrated in FIG. 3D, the proxy computing system 102 receives a model load request package 321. In response, as part of the model ID assignment process 322, the proxy computing system 102 can access the processing unit memory state 334 from a device library 336 (similar to the device library 122). For example, unit 1 216 can have 6 MB of available memory, unit 2 218 can have 3 MB of available memory, and so on until unit N 220 has 5 MB of available memory. Based on the processing unit memory state 334, the PU selection 338 can select one or more PUs (e.g., unit 1 and other PUs) as the PU / unit selection 339 for loading the model thereon. This selection of PUs may be referred to as a subset of PUs. The loaded model may be a neural network model corresponding to the model ID 324.
[0040] Continuing, when the proxy computing system 102 receives the inference request package 333, the PU selection 338 can access one or more endpoint execution queues associated with a subset of the PUs. The endpoint execution queue 340 may be analyzed along with the load state of each PU in the subset. Based on the analysis, the proxy computing system can select a target PU (e.g., unit 1 342) for executing the inference request 332. The inference request 332 included in the inference request package 333 may be routed to unit 1 342. In this case, the model context bank, the input tensor space (corresponding to the input tensor space 316), and the output tensor space (corresponding to the output tensor space 320) may be written to the main memory portion of unit 1 342. Unit 1 342 can process the inference request and output an output tensor 344 as shown in FIG. 3E. The output tensor 344 can include an inference (inference, i.e., "tree") from the neural network model loaded on unit 1 342. This neural network model may correspond to the model ID 324. The output tensor 344 may be included in the inference response 346 and sent back to the client computing system that sent the model load request package 321 and the inference request package 333.
[0041] FIGS. 4A, 4B, 4C, 4D, and 4E are block diagrams illustrating an inference request / response workflow 400.
[0042] FIG. 4A illustrates the operation on the client computing system side, where one or more endpoint execution queues 404 are stored in the client library 402. The endpoint execution queue 404 can include one or more endpoint execution queues associated with one or more PUs. For example, the endpoint execution queue 406 may be associated with unit 1 216. In each endpoint execution queue in the endpoint execution queue 404, "X" represents an inference request similar to the inference request 332.
[0043] In one aspect, the client computing system can issue a request thread 408 based on the queue state associated with the endpoint execution queue 404.
[0044] Based on the queue state associated with each endpoint execution queue in the endpoint execution queue 404, the queue selector 410 can select a specific endpoint execution queue for the inference request as the selected queue 412. The client library 402 can also receive an inference response from the proxy computing system 102 via the response thread 409.
[0045] In one aspect, when an application (e.g., a client application running on the client computing system 104) makes a request, this request is sent to the endpoint execution queue 404. In response, the queue selector 410 can select a queue (i.e., the selected queue 412) via an arbitration mechanism. The request thread 408 can extract the request existing in the selected queue (one of 406 identified by 410). The extracted request can include the extracted thread 413.
[0046] As shown in FIG. 4B, the client computing system generates an inference request 418. In one aspect, the inference request 418 consists of a model ID 324 (generated from a previous model load request operation), a selected queue 412, an extracted thread 413, and an input tensor 416 retrieved from an input tensor database 414. In one aspect, the input tensor database 414 includes a sequence of one or more images or video frames. The client computing system can send the inference request 418 to the proxy computing system 102 as an inference request package 420.
[0047] In one aspect, prior to the client computing system sending any inference requests to the proxy computing system 102, the client computing system can send a model load request to the proxy computing system 102 at least once. This model load request triggers model ID generation at the proxy computing system 102. This workflow is similar to the workflow illustrated in FIGS. 3A and 3B. After model ID generation, both the proxy computing system 102 and the client computing system store this model ID (e.g., model ID 324) in a memory or storage cache 118 and can use the model ID in any subsequent inference requests. Upon restart of any of the systems (i.e., the proxy computing system 102 and / or the client computing systems 104 to 108), each storage cache can retrieve the model ID and be used by all systems without the need to send a model load request again.
[0048] As shown in FIG. 4C, the proxy computing system 102 can retrieve the model 424 from the storage cache 422. The model 424 may be a neural network model consisting of an input tensor space 426, a set of weight tensors 428, and an output tensor space 430. The model 424 is output as model data 432.
[0049] As shown in FIG. 4D, the proxy computing system 102 receives the model data 432 that houses the model 424. The proxy computing system 102 can load the model 424 onto any, all, or one of the PUs, such as unit 1 442 and other PUs. This selection of PUs may be referred to as a subset of PUs. In one aspect, based on the processing unit memory state 436 retrieved from the device library 438 (similar to the device library 122), a subset of PUs is selected via the PU selection 440. For example, unit 1 216 can have 6 MB of available memory, unit 2 218 can have 3 MB of available memory, and so on until unit N 220 has 5 MB of available memory. Based on the processing unit memory state 436, the PU selection 440 can select one or more PUs (e.g., unit 1 and other PUs) as the PU / unit selection 441 for loading the model thereon.
[0050] The proxy computing system 102 can also receive an inference request package 420. The PU selection 440 can access one or more endpoint execution queues associated with a subset of PUs. The endpoint execution queue 434 may be analyzed along with the load state of each PU in the subset. Based on the analysis, the proxy computing system 102 can select a target PU (e.g., unit 1 442) for executing the inference request 418. The inference request 418 included in the inference request package 420 may be routed to unit 1 442. In this case, the model context bank, the input tensor space, and the output tensor space may be written to the main memory portion of unit 1 442. Unit 1 442 can process the inference request and output an output tensor 444 as illustrated in FIG. 4E. The output tensor 444 can include an inference (inference, i.e., "tree") from the neural network model loaded on unit 1 442. The output tensor 444 may be included in the inference response 446 and sent back to the client computing system that sent the model inference request package 420. The inference response 446 may be sent back to the client library 402 via the response thread 409.
[0051] FIGS. 5A, 5B, and 5C are block diagrams illustrating an inference request / response workflow 500.
[0052] FIG. 5A illustrates the operation on the client computing system side, where the client computing system constructs an inference request 508 that includes a model ID 324 generated from a previous model load request. The client computing system can also construct an input tensor 506 based on inputs received from any combination of the input tensor database 502 and the image sensor 504. In one aspect, the input tensor database 502 includes a sequence of one or more images or video frames. The image sensor 504 may be a camera sensor that generates a sequence of images or video frames. The input tensor 506 may be included in the inference request 508 by the client computing system. The client computing system can send the inference request 508 to the proxy computing system 102 as an inference request package 510.
[0053] The client computing system can also generate a statistics request 512, which includes a model ID 514 and an average inference time 516. The statistics request 512 may also be sent by the client computing system to the proxy computing system 102.
[0054] Figure 5B illustrates the operations on the proxy computing system side, where the proxy computing system 102 processes the inference request package 510. In response to receiving the inference request 501, the proxy computing system 102 can access the model occupancy 520 for the selected model ID from a device library 518 (similar to the device library 122). The model occupancy 520 for the selected model ID can provide a list of one or more PUs (e.g., unit 1 216, unit 2 218, etc.) on which the neural network model associated with the model ID 324 is loaded. The data from the model occupancy 520 for the selected model ID may be used by the proxy computing system 102 to select a subset of the PUs on which the neural network model associated with the model ID 324 is loaded.
[0055] In one aspect, the proxy computing system 102 can access a set of endpoint execution queues 522 associated with a subset of the PUs on which the neural network model associated with the model ID 324 is loaded. Based on the load state of each PU indicated by the endpoint execution queue 522, the PU selection 524 can perform a PU / unit selection 526 to select one or more target PUs from the subset to initiate the inference request 508. To initiate the inference request 508, the input tensor space and the output tensor space associated with the inference request 508 may be loaded into the main memory of each PU (e.g., unit 1 528).
[0056] FIG. 5C illustrates the operation on the proxy-computing system side, where one or more PUs initiate an inference request 508 to produce an output tensor 530. The output tensor 530 can include an inference (inference, i.e., "tree") from the neural network model loaded on unit 1 528. The output tensor 530 can be included in an inference response 532 and sent back to the client-computing system that sent the model inference request package 510.
[0057] In one aspect, the proxy-computing system 102 can process a statistics request 512 for model ID 514 to determine an average inference time 516. In one aspect, the model ID 514 is the same as model ID 324. Based on the proxy-computing system monitoring the execution of inference requests by the target processing device based on the neural network model, statistical calculation 534 can process the statistics request 512. The statistical calculation 534 can generate a response (including the average inference time) to the statistics request 512 and send the response to the client application running on the client-computing system.
[0058] FIG. 6 is a workflow diagram illustrating a workflow 600 for generating a model load response. The workflow 600 may be implemented on any combination of the proxy-computing system 102 and the client-computing systems 104 to 108.
[0059] Workflow 600 can include the client computing system accessing the model (602). For example, the client computing system can access model 304 from model parameter database 302. Workflow 600 can include the client computing system sending a model load request associated with the model to proxy computing system 102. For example, the client computing system can send model load request 314 to proxy computing system 102 as model load request package 321.
[0060] Workflow 600 can include the proxy computing system (e.g., proxy computing system 102) receiving the model load request (606). For example, the proxy computing system can receive model load request 314 via model load request package 321. Workflow 600 can include proxy computing system 102 accessing the memory state of each of one or more processing units 620 (608). For example, proxy computing system 102 can access processing unit memory state 334 from device library 336.
[0061] Workflow 600 can include proxy computing system 102 selecting a subset of processing units (610). For example, the processes of PU selection 338 and PU / unit selection 339 can select a subset of processing units (e.g., unit 1 342). Workflow 600 can include proxy computing system 102 loading a model (e.g., a neural network model associated with model ID 324) onto the subset of processing units (612).
[0062] Workflow 600 can include the subset of the processing system loading the model into the associated context bank and main memory (614). For example, unit 1342 can include the model context bank in the main memory that loads the neural network model corresponding to model ID 324.
[0063] Workflow 600 can include the proxy computing system 102 sending the model load response to the client computing system (616) and the client computing system receiving the model load response (618). For example, the proxy computing system 102 can send the load response 323 to the client computing system that originated the model load request.
[0064] Figure 7 is a workflow diagram illustrating workflow 700 for generating an inference response.
[0065] Workflow 700 can include the client computing system accessing the input tensor of the neural network model (702). For example, any of client computing systems 104 to 108 can access the input tensor from the input tensor database 326 or from the image sensor 328.
[0066] Workflow 700 can include the client computing system sending an inference request associated with an input tensor to the proxy computing system 102 (704). For example, the client computing system can send inference request 332 to the proxy computing system 102 as inference request package 333. Workflow 700 can include the proxy computing system receiving the inference request (706) and accessing the load state of each PU in the set of PUs (708). For example, proxy computing system 102 can receive inference request package 333 and access the load state associated with endpoint execution queue 340. The endpoint execution queue can correspond to processing units 710 that may be similar from unit 1 216 / 342 to unit 2 218, unit N 220.
[0067] Workflow 700 can include the proxy computing system 102 selecting a target processing unit (712). For example, PU selection 338 associated with proxy computing system 102 can select a target processing unit (e.g., unit 1 342) for executing inference request 332. Workflow 700 can include the proxy computing system 104 sending the input tensor to the target processing unit (714). For example, proxy computing system 104 can send inference request 332 including input tensor 330 to unit 1 342 for inference.
[0068] Workflow 700 can include performing an inference (716) based on a neural network model and an input tensor. For example, unit 1 342 can perform an inference request 332 based on input tensor 330 and model ID 324. Unit 1 342 can also generate an output tensor (e.g., output tensor 344 with an inference result such as "tree") as a result of performing the inference request. Workflow 700 can include the proxy computing system 102 retrieving the output tensor from the target processing unit (718). For example, the proxy system can retrieve output tensor 344 from unit 1 342.
[0069] Workflow 700 can include the proxy computing system 102 sending an inference response to a client application running on the client computing system (720). For example, the proxy computing system 102 can construct an inference response 346 including output tensor 344 and send the inference response 346 to the client computing system. Workflow 700 can include the client computing system receiving the inference response from the proxy computing system 102 (722).
[0070] In one aspect, the process of inferring involves launching an artificial intelligence (AI) algorithm on an input tensor obtained from an image or a video file. The proxy computing system 102 may be configured to map one or more client computing systems 104 to 108 to one or more PUs 128 to 140. Each of the client computing systems 104 to 108 may be a remote computing system or a local computing system. The client computing systems 104 to 108 can interface with the proxy computing system 102 via a communication protocol such as TCP / IP or other networking protocols.
[0071] In one aspect, the request manager 116 performs a load balancing function among the PUs 128 to 140. The request manager 116 can also account for aspects such as fault tolerance, device monitoring, device failures, and others.
[0072] Generally, any inference operation requires at least one AI / neural network model. AI models are generally large and difficult to handle in terms of computing resources. Faults or crashes in the client computing system and PU network can cause significant delays because these AI models may need to be reloaded into the memory device during system recovery. In one aspect, the proxy computing system 102 maintains a model library 120 in a storage cache 118. In the event of a fault or crash, the proxy computing system 102 can access the appropriate AI / neural network model in the storage cache 118 to quickly bring all systems back online.
[0073] In one aspect, the proxy computing system 102 can enable substantially seamless switching between different neural network models (from the perspective of client computing systems 104 to 108). Each neural network model may be associated with a unique model ID. Based on different inference requests from different client computing systems, different neural network models can be interchangeably loaded onto any combination of PUs 128 to 140.
[0074] In one aspect, an API running on a client computing system (e.g., client computing system 104) can interact directly with, or via the proxy computing system 102, with a PU (e.g., PU 128). In this sense, the proxy computing system 102 can function as a device driver. However, while a typical device driver is limited to a single interface (USB or PCIe), the proxy computing system 102 implements a unified interface (e.g., PCIe 110 and USB 112 are simultaneously connected to multiple PUs) that supports multiple interface protocols with multiple instances of the same protocol at the same time. The end user does not need to care about how such connectivity occurs, and the connectivity process is transparently implemented by the proxy computing system 102.
[0075] The proxy computing system can also perform the following functions. * Cycle stealing during inference for performing other tasks * Provisioning memory prior to the time of an inference request * Resource management / allocation and load balancing * Model management *Device management including fault handling (inference workload) *Thread management
[0076] Although the present disclosure has been described with respect to certain exemplary embodiments, other embodiments, including embodiments that do not provide all of the benefits and features described herein, will be apparent to those of ordinary skill in the art in view of the benefits of the present disclosure, and these are also within the scope of the present disclosure. It should be understood that other embodiments may be utilized without departing from the scope of the present disclosure.
Claims
1. Receiving a neural network model from a client computing system; Evaluating system resource availability on a plurality of processing devices; Selecting a subset of available processing devices based on the system resource availability; Loading the neural network model into each processing device in the subset; Receiving an inference request from the client computing system; Accessing the load state of each processing device in the subset; Selecting a target processing device from the subset based on the load state; Sending the inference request to the target processing device A method comprising.
2. After executing the inference request based on the neural network model, receiving an inference result generated by the target processing device; Sending the inference result to the client computing system The method according to claim 1, further comprising.
3. The method according to claim 2, wherein the inference result is an output tensor.
4. The method according to claim 1, wherein the neural network is a convolutional neural network or a neural network consisting of one or more linear algebra operators.
5. The method according to claim 1, further comprising automatically determining and negotiating the type of a processing device interface associated with the processing device.
6. The method according to claim 5, wherein the processing device interface is any one of a PCIe (Peripheral Component Interconnect-Express) bus interface, a USB (Universal Serial Bus) interface, or an IPC (Inter-Process Communication) interface.
7. The method according to claim 1, wherein the inference request includes an input tensor.
8. The method according to claim 7, wherein the input tensor is an image generated by an image sensor.
9. The method according to claim 1, further comprising selecting the subset based on analyzing the processing unit memory states of the plurality of processing devices.
10. The method according to claim 1, further comprising assigning a model ID to the neural network model.
11. A proxy computing system, A client computing system communicatively coupled to the proxy computing system, And a plurality of processing devices communicatively coupled to the proxy computing system An apparatus comprising: The proxy computing system receives a neural network model from the client computing system, The proxy computing system evaluates system resource availability on the plurality of processing devices, The proxy computing system selects a subset of available processing devices based on the system resource availability, The proxy computing system loads the neural network model into each processing device in the subset, The proxy computing system receives an inference request from the client computing system, The proxy computing system accesses the load state of each processing device in the subset, The proxy computing system selects a target processing device from the subset based on the load state, The proxy computing system sends the inference request to the target processing device, The target processing device executes the inference request based on the neural network model. Device.
12. The target processing device generates an inference result based on the execution, The target processing device sends the inference result to the proxy computing system, The proxy computing system sends the inference result to the client computing system. The device according to claim 11.
13. The device according to claim 12, wherein the inference result is an output tensor.
14. The device according to claim 11, wherein the neural network is a convolutional neural network or a neural network composed of one or more linear algebra operators.
15. The processing device in the plurality of processing devices is communicably coupled to the proxy computing system via a processing device interface, and the proxy computing system automatically determines and negotiates the type of the processing device interface. The apparatus according to claim 11.
16. The apparatus according to claim 15, wherein the processing device interface is any one of a PCIe bus interface, a USB interface, or an IPC interface.
17. The apparatus according to claim 11, wherein the inference request includes an input tensor.
18. The apparatus according to claim 17, wherein the input tensor is an image generated by an image sensor.
19. The apparatus according to claim 11, wherein the subset is selected based on analyzing the processing unit memory state of each of the plurality of processing devices.
20. The apparatus according to claim 11, wherein the proxy computing system assigns a model ID to the neural network model.