Machine learning model updates for machine learning accelerators

The hybrid gateway in peripheral I/O devices separates computational resources into I/O and coherent domains, enabling efficient, low-latency updates of machine learning models by leveraging cache-coherent shared memory multiprocessor paradigms, addressing security and reliability issues in traditional systems.

JP7795358B2Active Publication Date: 2026-01-07XILINX INC
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2021562796
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-04-26
Filing Date
2020-04-10
Publication Date
2026-01-07
Estimated Expiration
2040-04-10

AI Technical Summary

Technical Problem

Peripheral I/O devices in traditional computing systems lack the ability to benefit from the cache-coherent shared memory multiprocessor paradigm, leading to security and reliability issues due to multiple I/O device drivers, and inefficient updating of machine learning models.

Method used

Implementing a hybrid gateway in peripheral I/O devices that separates computational resources into an I/O domain and a coherent domain, allowing the use of a cache-coherent shared memory multiprocessor paradigm for updating machine learning models, with the ML model stored in the coherent domain and the ML engine in the I/O domain, and utilizing a coherent interconnect protocol to extend the coherent domain to the peripheral I/O devices.

Benefits of technology

Enables efficient, low-latency updates of machine learning models in peripheral I/O devices, improving security and reliability by leveraging both I/O and coherent domain paradigms, and allowing parallel execution of tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007795358000001
    Figure 0007795358000001
  • Figure 0007795358000002
    Figure 0007795358000002
  • Figure 0007795358000003
    Figure 0007795358000003
Patent Text Reader

Abstract

Examples herein describe peripheral I / O devices with a hybrid gateway that allows the device to have both an I / O domain and a coherent domain. As a result, computational resources in the coherent domain of the peripheral I / O device can communicate with the host in a manner similar to inter-CPU communication within the host. Dual domains within a peripheral I / O device may be utilized for machine learning (ML) applications. While I / O devices may be used as ML accelerators, these accelerators previously only used the I / O domain. In embodiments herein, computational resources may be divided between the I / O domain and the coherent domain, where the ML engine resides in the I / O domain and the ML model resides in the coherent domain. An advantage of doing so is that the ML model can be coherently updated using a reference ML model stored within the host.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Examples of the present disclosure generally relate to executing machine learning models within peripheral I / O devices that support both the I / O domain and the coherent domain. [Background technology]

[0002] In the traditional I / O model, a host computing system interfaces with a peripheral I / O device when executing an accelerator task or function using a custom I / O device driver specific to the peripheral I / O device. Having multiple I / O devices, or even multiple instances of the same I / O device, means that the host interfaces with multiple I / O device drivers or multiple running copies of the same I / O device driver. This can result in security and reliability issues because the I / O device drivers, which are typically developed by the supplier that provides the peripheral I / O device, must be incorporated along with all the software and hardware within the host computing system.

[0003] On the other hand, the hardware cache-coherent shared memory multiprocessor paradigm leverages a general, instruction set architecture (ISA)-independent model of interfacing execution tasks or functions on multiprocessor CPUs. The general, ISA-independent (e.g., C-code) model of interfacing scales both the number of processing units and the amount of shared memory available to those processing units. Traditionally, peripheral I / O devices have not been able to benefit from the coherent paradigm used by CPUs running on a host computing system. Summary of the Invention

[0004]

[0003] Techniques for executing machine learning models using an I / O domain and a coherent domain within a peripheral device are described. One example is a peripheral I / O device including: a hybrid gateway configured to communicatively couple the peripheral I / O device to a host; I / O logic including a machine learning (ML) engine assigned to the I / O domain; and coherent logic including the ML model assigned to a coherent domain, where the ML model shares the coherent domain with computational resources within the host.

[0005] In some embodiments, the hybrid gateway is configured to use a coherent interconnection protocol to extend the coherent domain of the host to the peripheral I / O devices.

[0006] In some embodiments, the hybrid gateway includes an update agent configured to use a cache-coherent shared memory multiprocessor paradigm to update the ML model in response to changes made to a reference ML model stored in a memory associated with the host.

[0007] In some embodiments, using the cache-coherent shared memory multiprocessor paradigm to update the ML model results in only a portion of the ML model being updated when the reference ML model is updated.

[0008] In some embodiments, the peripheral I / O device includes a network-on-chip (NoC) coupled to the I / O logic and the coherent logic, and at least one of the NoC, inter-programmable logic (PL) messages, and wired signaling is configured to allow parameters associated with a layer in the ML model to be transferred from the coherent logic to the I / O logic.

[0009] In some embodiments, the ML engine is configured to process an ML dataset received from the host using the parameters in the ML model.

[0010] In some embodiments, a peripheral I / O device includes a PL array, wherein a first plurality of PL blocks in the PL array are part of the I / O logic and assigned to the I / O domain, and a second plurality of PL blocks in the PL array are part of the coherent logic and assigned to the coherent domain, and a plurality of memory blocks, wherein a first subset of a plurality of memory blocks are part of the I / O logic and assigned to the I / O domain, and a second subset of the plurality of memory blocks are part of the coherent logic and assigned to the coherent domain, wherein the first subset of the plurality of memory blocks can communicate with the first plurality of PL blocks but cannot directly communicate with the second plurality of PL blocks, and the second subset of the plurality of memory blocks can communicate with the second plurality of PL blocks but cannot directly communicate with the first plurality of PL blocks.

[0011] One example described herein is a computing system including a host and a peripheral I / O device. The host includes a memory that stores a reference ML model and a plurality of CPUs that, together with the memory, form a coherent domain. The I / O device includes I / O logic including an ML engine assigned to an I / O domain and coherent logic including an ML model assigned to the coherent domain with the memory and the CPUs in the host.

[0012] In some embodiments, the peripheral I / O devices are configured to use a coherent interconnect protocol to extend the coherent domain to the peripheral I / O devices.

[0013] In some embodiments, the peripheral I / O device comprises an update agent configured to use a cache-coherent shared memory multiprocessor paradigm to update the ML model in response to changes made to the reference ML model stored in the memory.

[0014] In some embodiments, using the cache-coherent shared memory multiprocessor paradigm to update the ML model results in only a portion of the ML model being updated when the reference ML model is updated.

[0015] In some embodiments, the host includes multiple reference ML models stored in the memory, and the peripheral I / O device includes multiple ML models in the coherent domain corresponding to the multiple reference ML models, and multiple ML engines assigned to the I / O domain, where the multiple ML engines are configured to execute independently of one another.

[0016] In some embodiments, the peripheral I / O device includes a PL array, wherein a first plurality of PL blocks in the PL array are part of the I / O logic and assigned to the I / O domain, and a second plurality of PL blocks in the PL array are part of the coherent logic and assigned to the coherent domain, and a plurality of memory blocks, wherein a first subset of a plurality of memory blocks are part of the I / O logic and assigned to the I / O domain, and a second subset of the plurality of memory blocks are part of the coherent logic and assigned to the coherent domain, wherein the first subset of the plurality of memory blocks can communicate with the first plurality of PL blocks but cannot directly communicate with the second plurality of PL blocks, and the second subset of the plurality of memory blocks can communicate with the second plurality of PL blocks but cannot directly communicate with the first plurality of PL blocks.

[0017] One example described herein is a method that includes updating a portion of a reference ML model in a memory associated with a host; if the memory of the host and the coherent logic of the peripheral I / O device are in the same coherent domain, updating a subset of cached ML models in the coherent logic associated with the peripheral I / O device coupled to the host; retrieving the updated subset of cached ML models from the coherent domain; and, if an ML engine is in I / O logic in the peripheral I / O device assigned to an I / O domain, using the ML engine to process an ML dataset according to parameters in the retrieved subset of cached ML models.

[0018] In some embodiments, retrieving the updated subset of the cached ML models from the coherent domain is performed using at least one of NoC-to-PL messages and wired signaling that communicatively couples the coherent logic assigned to the coherent domain to the I / O logic assigned to the I / O domain.

[0019] Therefore, in a manner that the above-listed features can be understood in detail, a more particular description, briefly summarized above, may be made with reference to example implementations, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only typical example implementations and therefore should not be considered to limit the scope of the implementations. [Brief explanation of the drawings]

[0020] [Figure 1] 1 is a block diagram of a host coupled to a peripheral I / O device having an I / O domain and a coherent domain, according to an example embodiment. [Figure 2]1 is a block diagram of a peripheral I / O device having programmable logic, memory, and a network-on-chip logically partitioned into an I / O domain and a coherent domain, according to an example. [Figure 3] FIG. 1 is a block diagram of a peripheral I / O device having a machine learning model and a machine learning engine, according to an example. [Figure 4] 1 is a flowchart for updating a machine learning model in a coherent domain of an I / O device, according to an example. [Figure 5] 1 is a block diagram of an I / O expansion box containing multiple I / O devices, according to an example. [Figure 6] 1 is a flowchart for updating machine learning models cached on multiple I / O devices, according to an example. [Figure 7] 1 is a flowchart for using a recursive learning algorithm to update a machine learning model, according to an example. [Figure 8] FIG. 1 illustrates a field programmable gate array implementation of a programmable IC, according to an example. DETAILED DESCRIPTION OF THE INVENTION

[0021] Various features are described hereinafter with reference to the figures. It should be noted that the figures may or may not be drawn to scale, and that elements of similar structure or function are represented by like reference numerals throughout the figures. It should be noted that the figures are merely for ease of description of the features. The figures are not intended as a comprehensive description of the specification or as a limitation on the scope of the claims. Additionally, the illustrated examples do not necessarily have all aspects or advantages shown. Aspects or advantages described with a particular example are not necessarily limited to the example described above and may be implemented in any other example even if not so illustrated or explicitly described.

[0022] Examples herein describe peripheral I / O devices with hybrid gateways that allow the device to have both an I / O domain and a coherent domain. That is, the I / O device can benefit from the traditional I / O model in which an I / O device driver manages some of the computational resources within the I / O device, as well as adding other computational resources within the I / O device to the same coherent domain used by a processor (e.g., a central processing unit (CPU)) in a host computing system. As a result, computational resources within the peripheral I / O device's coherent domain can communicate with the host in a manner similar to inter-CPU communication within a host. This means that the computational resources can take advantage of coherent-type features such as direct communication, more efficient memory usage, non-uniform memory access (NUMA) awareness, etc. At the same time, the computational resources within the I / O domain can benefit from the traditional I / O device model (e.g., direct memory access (DMA)), which provides efficiencies when performing large memory transfers between the host and the I / O device.

[0023] The dual domains within peripheral I / O devices can be leveraged for machine learning (ML) applications. While I / O devices are sometimes used as ML accelerators, these accelerators previously only used the I / O domain. In embodiments herein, computational resources can be divided between the I / O domain and the coherent domain, with the ML engine assigned to the I / O domain and the ML model stored in the coherent domain. The advantage of doing so is that the ML model may be coherently updated using a reference ML model stored on the host. That is, some types of ML applications benefit from being able to quickly (e.g., in real time or with low latency) update one or more ML models within the I / O device. Storing the ML model in the coherent domain (instead of the I / O domain) means that a cache-coherent shared-memory multiprocessor paradigm can be used to update the ML model, which is much faster than relying on a traditional I / O domain model (e.g., direct memory access (DMA)). The ML engine, however, can run within the I / O domain of the peripheral I / O device. This is beneficial as ML engines often process large amounts of ML data that is more efficiently transferred between the I / O device and the host using DMA rather than a cache-coherent paradigm.

[0024] FIG. 1 is a block diagram of a host 105 coupled to a peripheral I / O device 135 having an I / O domain and a coherent domain, according to an example. The computing system 100 of FIG. 1 includes a host 105 communicatively coupled to the peripheral I / O device 135 using a PCIe connection 130. The host 105 can represent a single computer (e.g., a server) or multiple physical computing systems interconnected. In either case, the host 105 includes an operating system 110, multiple CPUs 115, and memory 120. The OS 110 can be any OS capable of performing the functions described herein. In one embodiment, the OS 110 (or hypervisor or kernel) defines a cache-coherent shared-memory multiprocessor paradigm for the CPUs 115 and memory 120. In one embodiment, the CPUs 115 and memory 120 are managed by the OS (or managed by the kernel / hypervisor) to form coherent domains that adhere to the cache-coherent shared-memory multiprocessor paradigm. However, as noted above, in the traditional I / O model, peripheral I / O devices 135 (and all of their computational resources 150) are excluded from the coherence domain defined within host 105. Instead, host 105 relies on I / O device drivers 125 stored within host memory 120 to manage computational resources 150 within I / O devices 135. That is, peripheral I / O devices 135 are controlled by and accessible through I / O device drivers 125.

[0025] In embodiments herein, the shared memory multiprocessor paradigm is available to peripheral I / O device 135 with all the performance advantages, software adaptability, and overhead reduction of the above paradigms. Furthermore, adding computational resources within I / O device 135 to the same coherent domain as CPU 115 and memory 120 enables a generic, ISA-agnostic development environment. As shown in FIG. 1 , some of the computational resources 150 of peripheral I / O device 135 are assigned to coherent domain 160, the same coherent domain 160 used by computational resources within host 105—e.g., CPU 115 and memory 120.

[0026] Computational resources 150A and 150B are assigned to I / O domain 145, while computational resources 150C and 150D are logically assigned to coherent domain 160. Thus, I / O device 135 benefits from having computational resources 150 assigned to both domains 145, 160. While I / O domain 145 provides efficiency when performing large memory transfers between host 105 and I / O device 135, coherent domain 160 provides the performance advantages, software adaptability, and overhead reduction discussed above. By logically dividing hardware computational resources 150 (e.g., programmable logic, network-on-chip (NoC), data processing engine, and / or memory) into I / O domain 145 and coherent domain 160, I / O device 135 can benefit from both types of paradigms.

[0027] To enable host 105 to send and receive both I / O data traffic and coherent data traffic, peripheral I / O device 135 includes hybrid gateway 140 that separates data received over PCIe connection 130 into I / O data traffic and coherent data traffic. I / O data traffic is forwarded to computational resources 150A and 150B in I / O domain 145, while coherent data traffic is forwarded to computational resources 150C and 150D in coherent domain 160. In one embodiment, hybrid gateway 140 can process I / O data traffic and coherent data traffic in parallel, such that computational resources 150 in I / O domain 145 can execute in parallel with computational resources 150 in coherent domain 160. That is, host 105 can assign tasks to computational resources 150 in both I / O domain 145 and coherent domain 160, which can execute the tasks in parallel.

[0028] The peripheral I / O devices 135 can be many different types of I / O devices, such as pluggable cards (which plug into an expansion slot in the host 105 or into a separate expansion box), systems-on-chips (SoCs), graphics processing units (GPUs), field-programmable gate arrays (FPGAs), etc. Thus, while many of the embodiments discuss I / O devices 135 that include programmable logic (e.g., programmable logic arrays), the embodiments may also apply to I / O devices 135 that do not have programmable logic but simply contain hardened circuitry (which may be software programmable). Furthermore, while the embodiments herein discuss dividing the computational resources 150 into two domains, in other embodiments, the hybrid gateway 140 can be modified to support additional domains or multiple sub-domains within the I / O domain and the coherent domain 145, 160.

[0029] In one embodiment, the hybrid gateway 140 and the host 105 use a coherent interconnect protocol to extend the coherent domain 160 to the peripheral I / O devices 135. For example, the hybrid gateway 140 can use the Cache Coherent Interconnect for Accelerators (CCIX) to extend the coherent domain 160 within the device 135. CCIX is a high-performance, chip-to-chip interconnect architecture that provides a cache coherent framework for heterogeneous system architectures. CCIX brings kernel-managed semantics to the peripheral devices 135. Cache coherency is automatically maintained at all times between the CPU on the host 105 and various other accelerators in the system, which may be located on any number of peripheral I / O devices.

[0030] However, other coherent interconnect protocols such as QuickPath Interconnect (QPI), Omni-Path, Infinity Fabric, NVLink, or OpenCAPI may be used in addition to CCIX to extend the coherent domain within the host 105 to include computational resources within peripheral I / O devices 135. That is, the hybrid gateway may be customized to support any type of coherent interconnect protocol that facilitates forming a coherent domain that includes computational resources within I / O devices 135.

[0031] 2 is a block diagram of a peripheral I / O device 135 having a programmable logic (PL) array 205, memory blocks 220, and an NoC 230 logically partitioned into I / O and coherent domains 145, 160, according to an example. In this example, PL array 205 is formed from multiple PL blocks 210. These blocks can be individually assigned to I / O domain 145 or coherent domain 160. That is, PL blocks 210A and 210B are assigned to I / O domain 145, while PL blocks 210C and 210D are assigned to coherent domain 160. In one embodiment, the set of PL blocks 210 assigned to the I / O domain is mutually exclusive with the set of PL blocks 210 assigned to the coherent domain, such that there is no overlap between the blocks (e.g., a PL block 210 is not assigned to both the I / O domain and the coherent domain).

[0032] In one embodiment, the assignment of hardware resources to either I / O domain 145 or coherent domain 160 does not affect (or indicate) the physical location of the hardware resources within I / O device 135. For example, PL blocks 210A and 210C may be assigned to different domains even though these blocks are adjacent to each other in PL array 205. Thus, although the physical location of hardware resources within I / O device 135 may be taken into consideration when logically assigning hardware resources to I / O domain 145 and coherent domain 160, this is not necessarily required.

[0033] The I / O devices 135 also include memory controllers 215 that are assigned to the I / O domain 145 and the coherent domain 160. In one embodiment, due to the physical interconnections between the memory controllers 215 and the corresponding memory blocks 220, assigning one of the memory controllers 215 to either the I / O domain 145 or the coherent domain 160 means that all memory blocks 220 connected to the memory controller 215 are also assigned to the same domain. For example, a memory controller 215 may be connected to a fixed set of memory blocks 220 (not connected to any other memory controllers 215). In this manner, a memory block 220 may be assigned to the same domain as the memory controller 215 to which it is connected. However, in other embodiments, it may be possible to assign memory blocks 220 connected to the same memory controller 215 to different domains.

[0034] In one embodiment, the NoC includes interface elements that allow hardware elements (e.g., configurable data processing engines, memory blocks 220, PL blocks 210, etc.) in the I / O device 135 to send and receive data using the NoC 230. In one embodiment, rather than using programmable logic to form the NoC 230, some or all of the components that form the NoC are hardened. In either case, the NoC 230 can be logically divided between the I / O domain 145 and the coherent domain 160. In one embodiment, instead of allocating different portions of the NoC 230 to the two domains, parameters of the NoC are configured to provide different service levels for data traffic corresponding to the I / O domain 145 and the coherent domain 160. That is, data traffic for both domains flowing through the NoC 230 may use the same hardware elements (e.g., switches and communication links) but may be treated differently by the hardware elements. For example, the NoC 230 can provide different Quality of Service (QoS), latency, and bandwidth to the two different domains. Additionally, the NoC 230 can also isolate traffic in the I / O domain 145 from traffic in the coherent domain 160 for security reasons.

[0035] In another embodiment, NoC 230 can prevent computational resources in I / O domain 145 from communicating with computational resources in coherent domain 160. However, in one embodiment, it may be advantageous to allow computational resources assigned to I / O domain 145 to communicate with computational resources assigned to coherent domain 160. Previously, this communication would have occurred between I / O device driver 125 and the OS in host 105. Instead, inter-domain communication can occur internal to I / O device 135 using NoC 230 (if the computational resources are far apart within device 135) or using an inter-fabric connection in PL array 205 (if two PL blocks 210 assigned to two different domains are close together and need to communicate).

[0036] FIG. 3 is a block diagram of a peripheral I / O device 135 having an ML model 345 and an ML engine 335, according to an example. In FIG. 3, a host 105 is coupled to a host-attached memory 305 that stores ML data and results 310 and a reference ML model 315. The ML data and results 310 include data that the host 105 sends to the peripheral I / O device 135 (e.g., an ML accelerator) for processing as well as results that the host 105 returns from the I / O device 135. The reference ML model 315, in turn, defines the layers and parameters of the ML algorithms that the peripheral I / O device 135 uses to process the ML data. The reference ML model 315 can also include multiple ML models, each defining the layers and parameters of multiple ML algorithms to be used to process the ML data, such that the host receives results from across multiple ML algorithms. Embodiments herein are not limited to a particular ML model 315 and can include binary classification, multi-class classification, regression, neural networks (e.g., convolutional neural networks (CNN) or recurrent neural networks (RNN)), etc. The ML model 315 can specify the number of layers, how multiple layers are interconnected, weights for each layer, etc. Additionally, while the host-attached memory 305 is shown as separate from the host 105, in other embodiments, the ML data and results 310 and the ML model 315 are stored in memory internal to the host 105.

[0037] The host 105 can update the reference ML model 315. For example, as more data becomes available, the host 105 can change some of the weights of certain layers of the reference ML model 315, change how layers are interconnected, or add / remove layers in the ML model 315. As discussed below, these updates in the reference ML model 315 may be mirrored in the ML model 345 stored (or cached) in the peripheral I / O device 135.

[0038] Hybrid gateway 140 allows host 105's coherent domain to be expanded to include hardware elements within peripheral I / O device 135. Additionally, hybrid gateway 140 defines an I / O domain that can use a traditional I / O model in which hardware resources assigned to the I / O domain are managed by I / O device drivers. To do so, hybrid gateway 140 includes I / O and DMA engine 320, which forwards I / O domain traffic between host 105 and hardware assigned to the I / O domain within peripheral I / O device 135, and update agent 325, which forwards coherent domain traffic between host 105 and hardware assigned to the coherent domain within peripheral I / O device 135.

[0039] In this example, hybrid gateway 140 (as well as I / O and DMA engine 320 and update agent 325) is connected to NoC 230, which facilitates communication between gateway 140 and I / O logic 330 and coherent logic 340. I / O logic 330 represents hardware elements within peripheral I / O device 135 assigned to the I / O domain, while coherent logic 340 represents hardware elements assigned to the coherent domain. In one embodiment, I / O logic 300 and coherent logic 340 include PL blocks 210 and memory blocks 220 illustrated in FIG. 2 . That is, some portions of PL blocks 210 and memory blocks 220 form I / O logic 330, while other portions form coherent logic 340. However, in another embodiment, I / O logic 300 and coherent logic 340 may not include any PL, but instead include hardened circuitry (which may be software programmable). For example, the peripheral I / O device 135 can be an ASIC or a specialized processor that does not include a PL.

[0040] As shown, ML engine 335 is executed using I / O logic 330, while ML model 345 is stored in coherent logic 340. As such, ML model 345 is in the same coherent domain as host-attached memory 305 and a CPU (not shown) in host 105. In contrast, ML engine 335 is not part of the coherent domain and therefore is not coherently updated when data stored in memory 305 is updated or otherwise changed.

[0041] Additionally, peripheral I / O device 135 is coupled to attached memory 350, which stores ML model 345 (which may be a cached version of ML model 345 stored in coherent logic 340). For example, peripheral I / O device 135 may not store the entire ML model 345 in coherent logic 340. Rather, the entire ML model 345 may be stored in attached memory 350, while a portion of ML model 345 currently being used by ML engine 335 is stored in coherent logic 340. In either case, the memory elements in attached memory 350 that store ML model 345 are part of the same coherent domain as coherent logic 340 and host 105.

[0042] ML dataset 355, in contrast, is stored in a memory element allocated to the I / O domain. For example, ML engine 335 can retrieve data stored in ML dataset 355, process the data according to ML model 345, and then store the processed data back in attached memory 350. Thus, in this manner, ML engine 335 and ML dataset 355 are allocated to hardware elements in the I / O domain, while ML model 345 is allocated to hardware elements in the coherent domain.

[0043] Although FIG. 3 illustrates one ML engine and one ML model, peripheral I / O device 135 can run any number of ML engines and ML models. For example, a first ML model may be good at recognizing object A in a captured image in most instances, except when the image contains both object A and object B. However, a second ML model may not recognize object A in many cases, but may be good at distinguishing between object A and object B. In this way, a system administrator can direct ML engine 335 to run two different ML models (e.g., there are two ML models stored in coherent logic 340). Furthermore, running ML engine 335 and ML model 345 may require only a small amount of the available computing resources in peripheral I / O device 135. In that case, the administrator can run another ML engine with a corresponding ML model in device 135. In other words, I / O logic 330 can run two ML engines, while coherent logic 340 stores two ML models. These ML engine / ML model pairs can run independently of each other.

[0044] Additionally, the allocation of computational resources to the I / O domain and the coherent domain may be dynamic; for example, a system administrator may determine that there are insufficient resources in the I / O domain for the ML engine 335 and reconfigure the peripheral I / O equipment 135 so that the computational resources previously allocated to the coherent domain are now allocated to the I / O domain. For example, PL blocks and memory blocks previously allocated to the coherent logic 340 may be reallocated to the I / O logic 330—e.g., an administrator may want to run two ML engines or may need two ML engines 335 to run two ML models. The I / O equipment 135 can be reconfigured with the new allocation, and the hybrid gateway 140 can simultaneously support operation in the I / O domain and the coherent domain.

[0045] FIG. 4 is a flowchart of a method 400 for updating an ML model in a coherent domain of an I / O device, according to an example. In block 405, the host updates a portion of the reference ML model in its memory. For example, an OS in the host (or a software application in the host) can run a training algorithm to modify or fine-tune the reference model. In one embodiment, the ML model is used to evaluate images to detect a particular object. Once the object is detected by the ML engine, the host can re-run the training algorithm, which results in updates to the ML model. That is, because detecting the object in the image can improve the training data, the host can decide to re-run the training algorithm (or a portion of the training algorithm), which can fine-tune the reference ML model. For example, the host can change weights corresponding to one or more layers in the reference ML model or change the way the layers are interconnected. In another example, the host can add or remove layers in the reference ML model.

[0046] In one embodiment, the host updates only a portion of the reference ML model. For example, the host changes weights corresponding to one or more of the layers, while the remaining layers in the reference ML model remain unchanged. As such, much of the data defining the ML model may remain unchanged after re-running the training algorithm. For example, the reference ML model may have a total of 20 MB of data, but an update may affect only 10% of that data. Under the traditional I / O device paradigm, the host must send the entire ML model (updated and non-updated data) to the peripheral I / O device, regardless of how small the update to the reference ML model is. However, by storing the ML model in a coherent domain of the peripheral I / O device, it is possible to avoid sending the entire reference ML model to the I / O device every time there is an update.

[0047] In block 410, the host updates only a subset of the cached ML models for the peripheral I / O device. In particular, the host sends the updated data in the reference ML model to the peripheral device in block 410. This transfer occurs within a coherent domain and may therefore behave like a transfer between memory elements within the host's CPU-memory complex. This is particularly useful in ML or artificial intelligence (AI) systems that rely on frequent (or low-latency) updates to ML models in ML accelerators (e.g., peripheral I / O devices).

[0048] In another example, placing an ML model in a coherent domain of an I / O device can be useful when the same ML model is distributed across many different peripheral I / O devices. That is, a host may be attached to multiple peripheral I / O devices that all have the same ML model. Thus, rather than having to update the entire reference ML model, the coherent domain can be leveraged to update only the data that has changed in the reference ML model in each of the peripheral I / O devices.

[0049] In block 415, the ML engine retrieves the updated portion of the ML model in the peripheral I / O device from the coherent domain. For example, although the NoC can keep I / O domain traffic and coherent domain traffic separate, the NoC can facilitate communication between hardware elements assigned to the I / O domain and hardware elements assigned to the coherent domain when desired. However, the NoC is only one of several transport mechanisms that can facilitate communication between the coherency domain and the I / O domain. Other examples include direct PL-to-PL messages or wired signaling, and communication via metadata written to a shared memory buffer between the two domains. In this way, the peripheral I / O device can transfer data from the ML model to the ML engine. Doing so enables the ML engine to process the ML dataset according to the ML model.

[0050] In one embodiment, the ML engine can retrieve only a portion of the ML model at any particular time. For example, the ML engine can retrieve parameters (e.g., weights) for one layer and configure I / O logic to execute that layer in the ML model. Once complete, the ML engine can retrieve parameters for the next layer in the ML model, and so on.

[0051] In block 420, I / O logic in the peripheral I / O device processes the ML dataset using an ML engine in the I / O domain according to the parameters of the ML model. The ML engine can use I / O domain techniques such as DMA to receive the ML dataset from the host. The ML dataset may be stored in the peripheral I / O device or in attached memory.

[0052] The ML engine processes the ML data set using the parameters in the ML model and returns the results to the host in block 425. For example, once finished, a DMA engine in the hybrid gateway can initiate a DMA write to transfer the processed data from a peripheral I / O device (or attached memory) to the host using an I / O device driver.

[0053] 5 is a block diagram of an I / O expansion box 500 containing multiple I / O devices 135, according to an example. In FIG. 5, a host 105 communicates with multiple peripheral I / O devices 135, which may be separate ML accelerators (e.g., separate accelerator cards). In one embodiment, the host 105 can assign different tasks to different peripheral I / O devices 135. For example, the host 105 can send different ML data sets to each of the peripheral I / O devices 135 for processing.

[0054] In this embodiment, the same ML model 525 runs on all peripheral I / O devices 135. That is, the reference ML model 315 in the host 105 is provided to each of the I / O devices 135, so that these devices 135 use the same ML model 525. As an example, the host 105 may receive feeds from multiple cameras (e.g., multiple cameras for an autonomous vehicle or multiple cameras in an area of ​​a city). To process the data generated by the cameras in a timely manner, the host 105 may aggregate the data and send different feeds to different peripheral I / O devices 135, so that these devices 135 can evaluate data sets in parallel using the same ML model 525. In this manner, using an I / O expansion box 500 with multiple peripheral I / O devices 135 may be preferred in ML or AI environments when quick response times are important or desired.

[0055] In addition to storing the ML model 525 in the peripheral I / O device 135, the expansion box 500 includes a coherent switch 505 that is separate from the I / O device 135. Nevertheless, the coherent switch 505 is also in the same coherent domain as the hardware resources in the host 105 and the cache 520 in the peripheral I / O device 135. In one embodiment, the cache 510 in the coherent switch 505 is another layer of cache between the cache 520 in the peripheral I / O device 135 and the memory elements that store the reference ML model 315 according to the NUMA arrangement.

[0056] While host 105 may send N copies of reference ML model 315 to each device 135 when a portion of reference ML model 315 is being updated (where N is the total number of peripheral I / O devices 135 in the container), only the updated portion of reference ML model 315 is transferred to cache 510 and cache 520 because caches 520 and 510 are in the same coherent domain. Therefore, the arrangement of FIG. 5 scales better than an embodiment in which ML model 525 is stored in hardware resources allocated to the I / O domain of peripheral I / O device 135.

[0057] FIG. 6 is a flowchart of a method 600 for updating a machine learning model cached on multiple I / O devices, according to an example. In one embodiment, the method 600 is used to update multiple copies of an ML model stored on multiple peripheral I / O devices connected to a host, such as the example illustrated in FIG. 5. In block 605, the host updates a portion of a reference ML model stored in host memory. The reference ML model may be stored in local memory or in attached memory. In either case, the reference ML model is part of a coherent domain shared, for example, by CPUs within the host.

[0058] At block 610, method 600 branches depending on whether a push model or a pull model is used to update the ML model. If a pull model is used, method 600 proceeds to block 615, where the host invalidates a subset of the cached ML models in the switch and peripheral I / O device. That is, in FIG. 5 , host 105 invalidates ML model 515 stored in cache 510 in switch 505 and ML model 525 stored in cache 520 in peripheral I / O device 135. Because ML model 525 is in the same coherence domain as host 105, host 105 does not necessarily invalidate all data in ML model 525 except for only the subset that has changed in response to updating reference ML model 315.

[0059] In block 620, an update agent in the peripheral I / O device retrieves the updated portion of the reference ML model from host memory. In one embodiment, block 620 occurs in response to an ML engine (or any other software or hardware actor in the coherent switch or peripheral I / O device) attempting to access an invalidated subset of the ML model. That is, if the ML engine attempts to retrieve data from an ML model in a cache that was not invalidated, the requested data is provided to the ML engine. However, if the ML engine attempts to retrieve data from an invalidated portion of the cache (this is also referred to as a cache miss), doing so triggers block 620.

[0060] In one embodiment, after determining that the requested data is invalidated on a local cache in the peripheral I / O device (e.g., cache 520), the update agent first attempts to determine whether the requested data is available in a cache in the coherent switch (e.g., cache 510). However, as part of performing block 615, the host invalidates the same subset of the caches in both the coherent switch and the peripheral I / O device 135. Doing so forces the update agent to retrieve the updated data from a reference ML model stored in the host.

[0061] In a pull model, updated data in the reference ML model is retrieved after a cache miss (e.g., when the ML engine requests an invalidated cache entry from the ML model). Thus, the peripheral I / O device can perform block 620 at different times (e.g., on demand) depending on when the ML engine (or any other actor in the device) requests the invalidated portion of the ML model.

[0062] In contrast, if the ML model is updated using a push model, at block 610, method 600 proceeds to block 625 where the host pushes the updated portions to caches in the switch and peripheral I / O devices. In this model, the host controls when the ML model cached in the peripheral I / O devices is updated, rather than the ML model being updated when there is a cache miss. The host can push out the updated data to the peripheral I / O devices in parallel or sequentially. In either case, the host does not need to push out all of the data in the reference ML model, except for only the portions of the reference ML model that have been updated or changed.

[0063] 7 is a flowchart of a method 700 for using a recursive learning algorithm to update a machine learning model, according to an example. In one embodiment, method 700 may be used to update a reference ML model using information obtained from running an ML model in a peripheral I / O device. In block 705, the peripheral I / O device (or host) identifies false positives in the result data generated by the ML engine when running the ML model. For example, the ML model may be designed to recognize a particular object or person in an image, but sometimes provides false positives (e.g., identifies the object or person when the object or person is not actually present in the image).

[0064] At block 710, the host uses a recursive learning algorithm to update the reference ML model within the host. In one embodiment, the recursive learning algorithm updates the training data used to train the reference ML model. In response to a false positive, the host can update the training data and then re-run at least a portion of the training algorithm using the updated training data. Thus, the recursive learning algorithm can update the reference ML model in real time using the result data provided by the ML engine.

[0065] At block 715, the host updates the cached ML models using the coherent domain. For example, the host can update one or more ML models in the peripheral I / O devices using the pull model described in blocks 615 and 620 of method 600 or the push model described in block 625. In this manner, by identifying a false positive in the resulting data generated by one or more of the peripheral I / O devices (e.g., one of the ML accelerators), the host can update the reference ML model. The host can then use the push model or the pull model to update the cached ML models of all peripheral I / O devices connected to the host.

[0066] 8 illustrates an FPGA 800 implementation of I / O peripherals 135, more specifically, with PL array 205 of FIG. 2, which includes a number of different programmable tiles including transceivers 37, CLBs 33, BRAMs 34, input / output blocks (“IOBs”) 36, configuration and clocking logic (“CONFIG / CLOCKS”) 42, DSP blocks 35, dedicated input / output blocks (“IOs”) 41 (e.g., configuration ports and clock ports), and other programmable logic 39 such as digital clock managers, analog-to-digital converters, system monitoring logic, etc. The FPGA may also include a PCIe interface 40, an analog-to-digital converter (ADC) 38, etc.

[0067] In some FPGAs, each programmable tile may include at least one programmable interconnect element (“INT”) 43 having connections to input and output terminals 48 of programmable logic elements within the same tile, as shown by the example included at the top of FIG. 8 . Each programmable interconnect element 43 may also include connections to interconnect segments 49 of adjacent programmable interconnect elements within the same tile or other tiles. Each programmable interconnect element 43 may also include connections to interconnect segments 50 of general-purpose routing resources between logic blocks (not shown). The general-purpose routing resources may include routing channels between logic blocks (not shown) that comprise tracks of interconnect segments (e.g., interconnect segments 50) and switch blocks (not shown) for connecting the interconnect segments. The interconnect segments of the general-purpose routing resources (e.g., interconnect segments 50) may span one or more logic blocks. The programmable interconnect elements 43 employed together with the general-purpose routing resources implement a programmable interconnect structure (“programmable interconnect”) for the illustrated FPGA.

[0068] In an example implementation, CLB 33 may include configurable logic elements ("CLEs") 44, which may be programmed to implement user logic plus a single programmable interconnect element ("INT") 43. BRAM 34 may include BRAM logic elements ("BRLs") 45 in addition to one or more programmable interconnect elements. Typically, the number of interconnect elements included in a tile depends on the height of the tile. In the illustrated example, the BRAM tile has the same height as five CLBs, although other numbers (e.g., four) may also be used. DSP block 35 may include DSP logic elements ("DSPLs") 46 in addition to an appropriate number of programmable interconnect elements. IOB 36 may include, for example, one instance of programmable interconnect element 43 plus two instances of input / output logic ("IOL") elements 47. As will be apparent to one skilled in the art, for example, the actual IO pads connected to IO logic elements 47 are typically not limited to the area of ​​the input / output logic elements 47.

[0069] In the illustrated example, a horizontal area near the center of the die (shown in FIG. 8) is used for configuration, clocks, and other control logic. Vertical columns 51 extending from this horizontal area or column are used to distribute clock and configuration signals across the width of the FPGA.

[0070] Some FPGAs utilizing the architecture illustrated in Figure 8 include additional logic blocks that break up the regular columnar structure that makes up most of the FPGA. The additional logic blocks can be programmable blocks and / or dedicated logic.

[0071] It should be noted that Figure 8 is merely illustrative of an exemplary FPGA architecture. For example, the number of logic blocks in a row, the relative width of the rows, the number and order of rows, the types of logic blocks included in the rows, the relative sizes of the logic blocks, and the interconnect / logic implementation included at the top of Figure 8 are merely exemplary. For example, in an actual FPGA, more than one adjacent row of CLBs is typically included wherever a CLB appears to facilitate efficient implementation of user logic, although the number of adjacent CLB rows will vary with the overall size of the FPGA.

[0072] The foregoing references embodiments presented in this disclosure. However, the scope of the disclosure is not limited to the specifically described embodiments. Instead, any combination of the described features and elements, whether related to another embodiment or not, is contemplated for implementing and performing the contemplated embodiments. Furthermore, while the embodiments disclosed herein may be implemented advantageously over other possible solutions or over the prior art, whether or not a particular advantage is achieved by a given embodiment does not limit the scope of the disclosure. Thus, the foregoing aspects, features, embodiments, and advantages are merely exemplary and are not considered elements or limitations of the appended claims unless expressly recited in the claims.

[0073] As will be appreciated by one skilled in the art, embodiments disclosed herein may be embodied as a system, method, or computer program product. Accordingly, aspects may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be referred to generically herein as a "circuit," "module," or "system." Furthermore, aspects may take the form of a computer program product embodied in one or more computer-readable medium(s) having computer-readable program code embodied therein.

[0074] Any combination of one or more computer-readable media can be utilized. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the context of this document, a computer-readable storage medium is any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0075] A computer-readable signal medium may include, for example, a propagated data signal having computer-readable program code embodied therein, such as in a baseband wave or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium is not a computer-readable storage medium but may be any computer-readable medium that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0076] The program code embodied on the computer readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, fiber optic cable, RF, etc., or any suitable combination of the foregoing.

[0077] Computer program code for performing operations related to aspects of the present disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" programming language or similar programming languages. The program code may run entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or a connection may be made to an external computer (e.g., via the Internet using an Internet Service Provider).

[0078] Aspects of the present disclosure are described below with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments presented in this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, executing via the processor of the computer or other programmable data processing apparatus, create means for performing the functions / acts specified in one or more blocks of the flowchart and / or block diagrams.

[0079] These computer program instructions may also be stored on a computer-readable medium that can direct a computer, other programmable data processing apparatus, or other device to function in a particular manner to produce an article of manufacture, where the instructions stored on the computer-readable medium include instructions that implement the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.

[0080] Computer program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device, generating a computer-implemented process such that the instructions executing on the computer or other programmable apparatus provide a process for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.

[0081] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various examples of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially simultaneously, or multiple blocks may sometimes be executed in the reverse order, depending on the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, may be implemented by a dedicated hardware-based system that performs the specified functions or acts or a combination of dedicated hardware and computer instructions.

[0082] While the foregoing is directed to specific examples, other and further examples may be devised without departing from the basic scope of the above examples, which scope is determined by the appended claims.

Claims

1. a hybrid gateway configured to communicatively connect peripheral I / O devices to a host; I / O logic including a machine learning (ML) engine assigned to an I / O domain; a coherent logic including an ML model assigned to a coherent domain; Equipped with a first set of computational resources within the peripheral I / O device are managed in the I / O domain, and a second set of computational resources within the peripheral I / O device are managed in the coherent domain that coherently updates the ML model; the ML engine is configured to process an ML dataset received from the host using parameters associated with layers in the ML model; Peripheral I / O devices.

2. A hybrid gateway configured to communicate with a host and a peripheral I / O device; I / O logic including a machine learning (ML) engine assigned to an I / O domain; a coherent logic including an ML model assigned to a coherent domain; Equipped with a first set of computational resources within the peripheral I / O device are managed in the I / O domain, and a second set of computational resources within the peripheral I / O device are managed in the coherent domain that coherently updates the ML model; The hybrid gateway is configured to use a coherent interconnection protocol to extend the coherent domain of the host to the peripheral I / O device.

3. The hybrid gateway, an update agent configured to use a cache-coherent shared memory multiprocessor paradigm to update a reference ML model according to changes made to the reference ML model stored in a memory associated with the host; The peripheral I / O device of claim 2 , comprising:

4. 4. The peripheral I / O device of claim 3, wherein using the cache-coherent shared memory multiprocessor paradigm to update the ML model results in only a portion of the ML model being updated when the reference ML model is updated.

5. A hybrid gateway configured to communicate with a host and a peripheral I / O device; I / O logic including a machine learning (ML) engine assigned to an I / O domain; a coherent logic including an ML model assigned to a coherent domain; a network-on-chip (NoC) coupled to the I / O logic and the coherent logic, wherein at least one of the NoC to programmable logic (PL) messages and wired signaling is configured to allow parameters associated with a layer in the ML model to be transferred from the coherent logic to the I / O logic; Equipped with a first set of computing resources within the peripheral I / O device managed in the I / O domain, and a second set of computing resources within the peripheral I / O device managed in the coherent domain that coherently updates the ML model; Peripheral I / O devices.

6. A hybrid gateway configured to communicate with a host and a peripheral I / O device; I / O logic including a machine learning (ML) engine assigned to an I / O domain; a coherent logic including an ML model assigned to a coherent domain; a PL array, wherein a first plurality of PL blocks in the PL array are part of the I / O logic and assigned to the I / O domain, and a second plurality of PL blocks in the PL array are part of the coherent logic and assigned to the coherent domain; a plurality of memory blocks, wherein a first subset of the plurality of memory blocks is part of the I / O logic and assigned to the I / O domain, and a second subset of the plurality of memory blocks is part of the coherent logic and assigned to the coherent domain, the first subset of the plurality of memory blocks being able to communicate with the first plurality of PL blocks but not directly with the second plurality of PL blocks, and the second subset of the plurality of memory blocks being able to communicate with the second plurality of PL blocks but not directly with the first plurality of PL blocks; Equipped with a first set of computing resources within the peripheral I / O device managed in the I / O domain, and a second set of computing resources within the peripheral I / O device managed in the coherent domain that coherently updates the ML model; Peripheral I / O devices.

7. a memory for storing a reference ML model; a plurality of CPUs forming a coherent domain together with said memory; a host comprising: I / O logic including an ML engine assigned to an I / O domain; coherent logic including an ML model assigned to the coherent domain along with the memory and the plurality of CPUs in the host; a peripheral I / O device comprising: Equipped with a first set of computational resources within the peripheral I / O device managed in the I / O domain, and a second set of computational resources within the peripheral I / O device managed in the coherent domain that coherently updates the ML model; the ML engine is configured to process an ML dataset received from the host using parameters associated with layers in the ML model; Computing system.

8. A memory that stores a reference ML model; a plurality of CPUs forming a coherent domain together with said memory; a host comprising: I / O logic including an ML engine assigned to an I / O domain; coherent logic including an ML model assigned to the coherent domain along with the memory and the plurality of CPUs in the host; a peripheral I / O device comprising: Equipped with a first set of computational resources within the peripheral I / O device managed in the I / O domain, and a second set of computational resources within the peripheral I / O device managed in the coherent domain that coherently updates the ML model; A computing system, wherein the peripheral I / O devices are configured to use a coherent interconnect protocol to extend the coherent domain to the peripheral I / O devices.

9. 9. The computing system of claim 8, wherein the peripheral I / O device comprises an update agent configured to use a cache-coherent shared memory multiprocessor paradigm to update the ML model according to changes made to the reference ML model stored in the memory.

10. 10. The computing system of claim 9, wherein using the cache-coherent shared memory multiprocessor paradigm to update the ML model results in only a portion of the ML model being updated when the reference ML model is updated.

11. The host includes a number of reference ML models stored in the memory, and the peripheral I / O device includes: a number of ML models in the coherent domain corresponding to the number of reference ML models; a plurality of ML engines assigned to the I / O domain, the plurality of ML engines being configured to execute independently of one another; The computing system of claim 10, comprising:

12. A memory for storing a reference ML model; a plurality of CPUs forming a coherent domain together with said memory; a host comprising: I / O logic including an ML engine assigned to an I / O domain; coherent logic including an ML model assigned to the coherent domain along with the memory and the plurality of CPUs in the host; a PL array, wherein a first plurality of PL blocks in the PL array are part of the I / O logic and assigned to the I / O domain, and a second plurality of PL blocks in the PL array are part of the coherent logic and assigned to the coherent domain; a plurality of memory blocks, wherein a first subset of the plurality of memory blocks is part of the I / O logic and assigned to the I / O domain, and a second subset of the plurality of memory blocks is part of the coherent logic and assigned to the coherent domain, the first subset of the plurality of memory blocks being able to communicate with the first plurality of PL blocks but not directly with the second plurality of PL blocks, and the second subset of the plurality of memory blocks being able to communicate with the second plurality of PL blocks but not directly with the first plurality of PL blocks; a peripheral I / O device comprising: Equipped with a first set of computing resources within the peripheral I / O device managed in the I / O domain, and a second set of computing resources within the peripheral I / O device managed in the coherent domain that coherently updates the ML model; Computing system.

13. updating a small portion of the reference ML model in a memory associated with the host; updating a subset of cached ML models in coherent logic associated with a peripheral I / O device attached to the host, wherein the memory of the host and the coherent logic of the peripheral I / O device update a subset of cached ML models that are in the same coherent domain; retrieving the updated subset of the cached ML models from the coherent domain; using an ML engine to process an ML data set according to parameters in the retrieved subset of the cached ML models, the ML engine executing in I / O logic within the peripheral I / O device assigned to an I / O domain; A method comprising:

14. 14. The method of claim 13, wherein retrieving the updated subset of the cached ML models from the coherent domain is performed using at least one of NoC-to-PL messages and wired signaling that communicatively couples the coherent logic assigned to the coherent domain to the I / O logic assigned to the I / O domain.

Citation Information

Patent Citations

  • Coarse grain coherency

    EP3385847A1

  • Integrated multimedia system

    JP2006179028A

  • Communications traffic processing architecture and method

    JP2016510524A

  • System and method for providing forward progress and avoiding starvation and livelock in a multiprocessor computer system

    US20040093455A1

  • Access management technique with operation translation capability

    US20100228945A1