Artificial Intelligence Hub for Local Network Inference Operations

DE102026107936A1Undetermined Publication Date: 2026-08-27NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102026107936
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-27
Filing Date
2026-02-26
Publication Date
2026-08-27

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Several examples disclose systems and methods relating to inference operations in local area networks (LANs). A system can receive a request for an inference task from a LAN device. The system can select a first inference device from a plurality of inference devices connected to the LAN, based at least on the inference task and one or more processing capabilities of the plurality of inference devices. In response to the request, the system can provide the LAN device with a specification of the first inference device.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND Machine learning inference tasks are typically performed on computing devices equipped with specialized hardware to improve computational efficiency. Efficiently executing machine learning inference tasks in local network environments can be challenging. SUMMARY The invention is defined by the claims. To illustrate the invention, aspects and embodiments that may or may not fall within the scope of the claims are described here. This disclosure relates to techniques for implementing an artificial intelligence node within a local network environment. Conventional approaches to performing AI inference operations involve sending data to remote servers or cloud-based platforms for processing, resulting in significant latency due to network transmission delays and bandwidth limitations. Furthermore, such approaches can violate data privacy restrictions, as sensitive information may be transmitted over potentially insecure networks or processed by systems lacking adequate cybersecurity frameworks. While some solutions implement local processing operations, conventional approaches to implementing AI models in a local network environment do not fully utilize local computing resources.For example, PCs equipped with graphics processing units (GPUs) often remain unused or operate at low utilization. Similarly, Internet of Things (IoT) devices may have specialized hardware capable of performing complex calculations, but this hardware is not used effectively. The techniques described here provide an AI node that can operate within a local network environment and dynamically allocate computing resources for inference operations based on real-time availability and suitability for processing. These approaches can be implemented to minimize latency by processing data as close as possible to its source (e.g., within the local network), thereby reducing reliance on remote servers or cloud platforms, while maintaining high performance through the dynamic allocation of inference tasks. The AI ​​node can access deployed deep learning models hosted as microservices on various devices within the local network.In some implementations, inference tasks can be scheduled at least based on optimal computing locations, taking into account factors such as power mode and network availability. Dynamic AI task allocation improves overall system performance by utilizing unused resources compared to traditional approaches for locally executing AI models. At least one aspect relates to one or more processors. The one or more processors may comprise one or more circuits. The one or more circuits may receive a request for an inference task from a device on a local network. The one or more circuits may select a first inference device from a plurality of inference devices connected to the local network, based at least on the inference task and / or one or more processing capabilities of the plurality of inference devices. The one or more circuits may provide the device on the local network with an indication of the first inference device in response to the request. In some embodiments, the one or more circuits can store a list of identifiers of the plurality of inference devices in conjunction with the one or more processing capabilities. In some embodiments, the one or more circuits can select the first inference device according to an order of the list of identifiers. In some embodiments, the one or more circuits can update the list of identifiers of the multiple inference devices based on at least one or more responses to a multicast message transmitted over the local network. In some embodiments, the one or more circuits can provide a network address of the first inference device as part of the specification. In some embodiments, the one or more circuits can receive data from the first inference device indicating at least one machine learning model stored on the first inference device. In some embodiments, the one or more circuits can update the one or more processing capabilities of the plurality of inference devices, at least based on the data indicating the at least one machine learning model. In some embodiments, the one or more circuits can determine that the first inference device has sufficient computational resources for the inference task. In some embodiments, the one or more circuits can select the first inference device after determining that the first inference device has sufficient computational resources for the inference task. In some embodiments, the one or more circuits can send a message corresponding to the inference task to a second inference device from a plurality of inference devices. In some embodiments, the one or more circuits can detect that the second inference device has failed to provide a response to the message within a time limit. In some embodiments, the one or more circuits can select the first inference device after detecting that the second inference device has failed to provide the response within the time limit. In some embodiments, the one or more circuits can receive data for the inference task from the local network device.In some embodiments, the one or more circuits can provide the data for the inference task to the first inference device in response to the selection of the first inference device. In some embodiments, the one or more circuits can receive output information from the first inference device, generated by the first inference device according to the inference task. In some implementations, the one or more circuits can create a schedule for a multitude of inference tasks received from a multitude of devices over the local network. At least one aspect concerns a system. The system can comprise a multitude of inference devices that communicate with each other via a local network. Each of the numerous inference devices can store a respective machine learning model. The system can include a control device that communicates with the local network. The control device can, through communication with the local network, create a list of the numerous inference devices connected to the local network. The control device can receive a specification of an inference task from a client device. The control device can select an inference device from the multitude of inference devices, based at least on the inference task and the respective machine learning model stored on each of the multitude of inference devices.The control device can generate a response to the request, based at least on the selected inference device. In some embodiments, the control device can provide a local network address of the selected inference device in response to the request. In some embodiments, the control device can transmit data for the inference task to the selected inference device. In some embodiments, the control device can monitor the execution of the inference task on the selected inference device. In some embodiments, the control device can provide the client device with an output of the inference task generated by the selected inference device. In some embodiments, the control device can be contained within the plurality of inference devices.In some embodiments, the selected inference device can be a first inference device, and the control device can determine, based on a second network communication, that a second inference device is not available from the plurality of inference devices for performing the inference task. In some embodiments, the control device can select the first inference device at least based on the unavailability of the second inference device. At least one aspect relates to a procedure. The procedure may include receiving a request for an inference operation from a device on a local network. The procedure may include selecting a first inference device from a plurality of inference devices connected to the local network based on at least one of the inference operations or one or more processing capabilities of the plurality of inference devices. The procedure may include assigning at least one part of the inference operation to the first inference device. In some embodiments, the method may include storing a list of identifiers of the multiple inference devices in conjunction with one or more processing capabilities. In some embodiments, the method may include selecting the first inference device according to an order of the list of identifiers. In some embodiments, the method may include updating the list of identifiers of the multiple inference devices based on at least one or more responses to a multicast message transmitted over the local network. In some embodiments, the method may include providing an indication of the first inference device to the device on the local network in response to the request, the indication including a network address of the first inference device as part of the indication. The processors, systems, and / or methods described herein may be implemented by or include at least one of the following elements: a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine, a system for performing simulation operations, a system for performing digital twin operations, a system for performing light transport simulations, a system for performing collaborative content creation for 3D assets, a system for performing deep learning operations, a system for performing generative AI operations using a small language model, a system for performing generative AI operations using a large language model, a system for performing generative AI operations using a small language model, a system for performing generative AI operations using an image language model.a system for performing generative AI operations using a multimodal language model, a system implemented using an edge device, a system implemented using a robot, a system for performing conversational AI operations, a system for generating synthetic data, a system comprising one or more virtual machines (VMs), a system implemented at least partially in a data center, or a system implemented at least partially using cloud computing resources. Further features of the disclosure are characterized by the independent and dependent claims. Any feature of one aspect of the disclosure can be applied in any suitable combination to other aspects of the disclosure. In particular, aspects of the method can be applied to apparatus or system aspects and vice versa. Furthermore, features implemented in hardware can be implemented in software and vice versa. Any reference to software and hardware features in this disclosure shall be interpreted accordingly. Each system or device feature described herein can also be provided as a process feature, and vice versa. Functionally described system and / or device aspects (including means-plus functional features) can alternatively be expressed in terms of their corresponding structure, such as a suitably programmed processor and associated memory. It should also be noted that certain combinations of the various features described and defined in any aspect of the disclosure can be implemented and / or provided and / or used independently of one another. The disclosure also provides computer programs and computer program products comprising software code designed to perform one of the methods when executed on a data processing device and / or to embody one of the device and system features described herein, including one or all partial steps of a method. The disclosure further provides a computer or computer system (including networked or distributed systems) that has an operating system which supports a computer program to perform one of the methods described herein and / or to embody one of the device or system features described herein. The disclosure also provides a computer-readable medium on which one or more of the aforementioned computer programs are stored. The disclosure also relates to a signal that transmits one or more of the aforementioned computer programs. The disclosure extends to methods and / or devices and / or systems as described herein with reference to the accompanying drawings. Aspects and embodiments of the disclosure will now be described purely by way of example with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS The present systems and methods for implementing an artificial intelligence node for inference operations in local networks are described in detail below with reference to the accompanying drawings, wherein: Fig. 1 is a block diagram of an example system for implementing inference operations in local networks according to some embodiments of the present disclosure; Fig. 2 is an example diagram showing an example process implemented by one or more computer systems described herein to perform inference operations in local networks according to some embodiments of the present disclosure; Fig. 3 is a flowchart of an example method for implementing inference operations in local networks according to some embodiments of the present disclosure; Fig.Figure 4 is a block diagram of an example of a computing device suitable for carrying out some embodiments of the present disclosure; and Figure 5 is a block diagram of an example of a computing center suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION This disclosure relates to systems and methods for implementing an artificial intelligence node within a local network environment. Conventional approaches to performing artificial intelligence inference operations involve sending raw data generated by sensors or devices to remote servers or cloud-based platforms for processing, resulting in significant latency due to network transmission delays and bandwidth limitations. Furthermore, such approaches can violate data privacy restrictions for the data being processed, as sensitive information may be transmitted over potentially insecure networks or processed by systems lacking adequate cybersecurity frameworks. Although some solutions implement local processing, conventional approaches to deploying artificial intelligence models in a local network environment do not fully utilize local computing resources. For example, PCs equipped with graphics processing units (GPUs) often remain idle or operate at low utilization when users are not performing computationally intensive tasks. Similarly, while Internet of Things (IoT) devices possess specialized hardware capable of performing complex calculations, they lack the necessary software frameworks to effectively leverage these capabilities. The techniques described here provide an AI node that can operate within a local network environment and dynamically allocate computing resources for inference operations based on real-time availability and processing suitability. These approaches can be implemented to minimize latency by processing data as close as possible to its source (e.g., within the local network), thereby reducing reliance on remote servers or cloud platforms, while maintaining high performance through the dynamic allocation of inference tasks. The AI ​​node can access deployed deep learning models hosted as microservices on various devices within the local network. These techniques allow inference tasks to be scheduled, at least based on optimal computing locations, taking into account factors such as power mode and network availability. For example, if a PC with an unused GPU is on the network, the AI ​​node can delegate certain computationally intensive tasks to this device, instead of relying solely on less powerful devices or high-latency cloud services. Dynamic AI task allocation improves overall system performance by utilizing unused resources compared to traditional approaches to running AI models locally. Furthermore, implementing a decentralized and flexible resource management strategy enables the AI ​​node to operate seamlessly across various platforms while ensuring a consistent user experience. The use of multicast DNS facilitates the discovery of and access to the AI ​​node within local networks, eliminating the need for complex configuration steps or specialized knowledge on the part of end users. This approach represents a significant improvement over previous solutions that either relied solely on centralized cloud services or failed to allocate resources based on real-time conditions. Referring to Fig. 1, Fig. 1 shows an exemplary computing environment comprising a system for implementing an AI node for inference operations in local networks according to some embodiments of the present disclosure. It is understood that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, instructions, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional units that can be implemented as discrete or distributed components, or in conjunction with other components, and in any suitable combination and location.Various functions described herein as being performed by units can be executed by hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory. System 100 can be used to perform machine learning inference operations (e.g., artificial intelligence) on local networks. System 100 is shown with one or more inference systems 102A-102N (sometimes referred to generally as "inference system(s) 102" or "inference devices 102"), at least one client device 110, and at least one local network 110. In some embodiments, the local network 110 can be connected to one or more cloud systems 122 via at least one external network 120. Each of the inference systems 102 can include at least one machine learning model 104. At least one of the inference systems 102 (shown here as inference system 102A) can include a controller 106. One or more client devices 108 can store and / or run one or more applications 112 to request the execution of inference tasks over the local network 110. Local Area Network 110 (LAN 110) can be any type of network infrastructure that enables communication between devices. In some embodiments, LAN 110 may be limited to a group of devices within a geographical area and may include, but are not limited to, home networks, office / corporate networks, or campus networks. Examples of LAN 110 include, but are not limited to, wireless LANs, Ethernet-based local area networks (LANs), and Bluetooth personal area networks (PANs), or combinations thereof. In some embodiments, LAN 110 may include virtual private networks (VPNs). LAN 110 can support various communication protocols and standards, thus enabling seamless interaction between different types of devices and systems. The local area network 110 can coordinate communication between all connected computer devices, including one or more client devices 108 and one or more inference devices 102. To this end, the local area network 110 can facilitate data transmission, packet routing, and the enforcement of firewalls or other security measures, while ensuring that the devices can efficiently discover and communicate with each other. The local area network 110 can include any number of switches, routers, or other devices that facilitate the transmission or routing of network data (e.g., network packets).In some embodiments, the local network 110 or its devices may use network management protocols and services, such as the Dynamic Host Configuration Protocol (DHCP) for automatic IP address assignment, the Domain Name System (DNS) for name resolution, and Multicast DNS (mDNS) for service discovery, among others. In some embodiments, the local network 110 can be connected to one or more external networks 120, such as wide area networks (WANs), the internet, or other private or public networks. Devices within the local network 110, such as routers or gateways, can coordinate access to the external network 120. For example, a router can manage the routing of network traffic between the local network 110 and the external network 120, ensuring that data packets are correctly forwarded to their intended destinations. The router can also perform network address translation (NAT) to allow devices within the local network 110 to communicate with devices in the external network 120 using private internet protocol (IP) addresses. The external network 120 can provide access to one or more cloud systems 122 and / or external computer systems (e.g., computer systems outside the local network 110). The external network 120 can facilitate the transfer of data between the local network 110 and the cloud systems 122. Devices within the local network 110 can communicate with external computer systems via the external network 120 to retrieve data and / or access functions of the cloud system 122. For example, cloud systems 122 can provide storage space, computing resources, and machine learning models that devices within the local network 110 can access.In some embodiments, routers / switches of the local network 110 can prevent devices of the external network 120 (which may include the cloud system) from accessing certain ports or other network resources of the local network 110 or its devices (e.g., via configuration settings, firewalls, network routing rules, etc.). Cloud System 122 can be any type of cloud computing system or distributed computing environment. In some embodiments, Cloud System 122 can be a collection of remote servers and infrastructure that provide various cloud services. Cloud System 122 can include, among other things, data centers, servers, storage devices, and network equipment. In some embodiments, Cloud System 122 can provide one or more cloud services, including remote computing, storage, and application hosting. Cloud System 122 can be accessed via the external network 120. In some embodiments, Cloud System 122 can provide or otherwise store one or more machine learning models 104 or other types of services for the inference of one or more inference systems 102.For example, the cloud system can host 122 machine learning models 104 in various formats, such as software packages, containers or microservices, and make them available to the inference systems 102 for retrieval and execution, as described in more detail herein. The client device 108 can be any type of computer capable of running applications and communicating over a network. Examples of client devices 108 include smartphones, tablets, PCs, laptops, and smart home devices. The client device 108 can comprise various hardware components, such as a processor, RAM, storage, input / output interfaces, and network interfaces. The client device 108 can communicate over the local network 110 using various communication protocols, including Wi-Fi, Ethernet, or Bluetooth. The client device 108 can run one or more applications 112, which can be stored and run on the device or accessed through a web browser.In implementations where the application 112 is accessed via a web browser, a web-based application can be provided by a controller 106 of an inference system 102A, as described in more detail herein. The application 112 can be any type of software, hardware, or a combination of hardware and software that provides a user interface and / or an application programming interface (API) for performing the various operations described herein. In some embodiments, the application 112 can provide a graphical user interface that allows a user to specify, configure, and / or request the execution of inference tasks by devices (e.g., inference systems 102) of the local network 110. The application 112 can receive input from the user, which may include data, selections thereof, and configurations thereof for inference tasks. In some embodiments, the application 112 can receive configurations / data for inference tasks via one or more API calls, interprocess communication, or other types of input. Inference tasks can encompass any type of processing task involving machine learning models or artificial intelligence operations. Examples of such tasks include image recognition, object recognition, natural language processing, speech recognition, and predictive analytics. Image recognition tasks can involve identifying and classifying objects in images or videos. Object recognition tasks can go beyond simple detection to locate and delineate objects within an image or video frame. Natural language processing tasks can include sentiment analysis, speech translation, text summarization, or other generative text operations involving one or more language models (e.g., large language models (LLMs), small language models (SLMs), etc.).Speech recognition tasks can involve converting spoken words into text, while predictive analytics tasks can forecast future trends based on historical data. As described here, application 112 can receive choices and configurations for various types of inference tasks from other applications running on client device 108 via a user interface or API calls. For example, application 112 can display a menu or a set of options that allow the user to select the specific type(s) of inference task(s) to perform. In some embodiments, other applications running on client device 108 can request inference tasks by making API calls or otherwise communicating with application 112. The other applications / processes on client device 108 can provide or otherwise specify parameters for the inference task, identify input data for the inference task, and / or provide additional configuration settings for the inference task.In some implementations, application 112 can also enable interprocess communication to receive inference task requests from other applications running on client device 108. In an implementation where application 112 allows the configuration of an inference task via a user interface, application 112 can provide one or more configuration options for one or more selectable inference tasks. Examples of configuration operations include identifying / specifying the input data for the inference task, setting task-specific parameters for the inference task, or defining / specifying output formats for the inference task. Once the user or another application has configured the task, application 112 can send a request to execute the inference task to a controller 106 of an inference system 102 on the local network 110. The request can include one or more of the specified parameters of the inference task.In some embodiments, the application 112 can transmit the input data for the inference task to the controller 106 as part of the request or in addition to it. The network address of the inference system 102, on which the controller 106 is running, can be stored in an internal configuration of the application 112 or determined dynamically from multicast or other communication over the local network 110. In some embodiments, a user and / or an application can specify / select the controller 106 and / or the inference system 102 via a corresponding interface of the application 112. In some embodiments, the application 112 can use multicast DNS or other service discovery protocols to identify inference systems 102 with a controller 106. Once the network address of the inference system 102A is known, the application 112 can establish a communication channel with the corresponding controller 106 to transmit inference tasks and receive results, as described in more detail herein.The 112 application can handle various aspects of communication, including but not limited to error checking, retransmission, and status updates, among other things. In some embodiments, the client device 108 can be an IoT device comprising one or more sensors, such as cameras, temperature sensors, humidity sensors, or motion detectors, as well as other types of sensors. Various processes, firmware, or other control commands can cause the sensors of the client device 108 to acquire various sensor data (e.g., images, temperature readings, etc.). In such implementations, the application 112 on the client device 108 can be an embedded application, firmware, or any other type of software component (or a combination of hardware and software) that can run on an IoT device.In some embodiments, the application 112 can access data from the sensors (or from other processes that control / operate the sensors) and send requests to the controller 106 to process the sensor data using one or more machine learning models 104, as described herein. For example, the application 112 can send image data captured by a camera to the controller 106 for object detection or send temperature and humidity data to the controller 106 for environmental analysis / prediction. The client device 108 can communicate with the controller 106 of one or more inference systems 102 over the local network 110 to enable the transmission of sensor data and the execution of inference tasks, as described in more detail herein. The inference systems 102 can be any type of computer capable of performing machine learning operations or parts thereof. Such devices can include, but are not limited to, personal computers, laptops, smartphones, and tablets. In some embodiments, one or more of the inference systems 102 can include specialized hardware, such as graphics processing units (GPUs), tensor processing units (TPUs), and field-programmable gate arrays (FPGAs), among others. In some embodiments, the inference systems 102 can store libraries, drivers, or other types of software, or combinations of hardware and software, optimized for performing specific types of computations, such as matrix multiplications, tensor operations, convolution operations, or other types of mathematical operations commonly used in machine learning tasks. Each of the inference systems 102 can have different processing capabilities. For example, a personal computer with a high-performance GPU can process complex and computationally intensive machine learning tasks, while a laptop, smartphone, or IoT device may be limited to simpler operations due to constraints in processing power and memory. The inference systems 102 can communicate with each other and with other computing devices, such as the client device 108 and / or the external network 120, via the local network 110. For example, one or more of the inference systems 102 can use the local network 110 to exchange data, coordinate the assignment of inference tasks, or perform one of the operations described here.Each inference system 102 can communicate over the local network 110 using any suitable communication protocol, including but not limited to Wi-Fi, Ethernet or Bluetooth, among others. Each inference system 102 can store one or more machine learning models 104 that are used to perform one or more inference tasks or parts thereof. Examples of machine learning models 104 include convolutional neural networks (CNNs) for image and object recognition, transformer-based models such as generative pretrained transformer models (GPTs) or other language models such as recurrent neural networks (RNNs) for natural language processing and speech recognition, diffusion-based models for image and / or video generation, and other types of machine learning models (e.g., regression models, sparse vector machine models, decision tree models, etc.) based on input data to generate predictions, among others.In some embodiments, one or more of the inference systems 102 can store or manage multiple machine learning models, each of which may correspond to a different inference task, a different input data type, or a different prediction / output data. The machine learning models 104 can be stored as software packages, containers, or other types of software services, such as microservices. For example, the machine learning models 104 can be deployed within container components, which may contain runtime libraries and configuration settings for running one or more machine learning models 104. In some implementations, the machine learning models 104 can be retrieved, downloaded, or otherwise accessed via the external network 120 from one or more cloud systems 122 or external devices. The inference system 102 can communicate with the cloud systems 122 or external devices to download or retrieve the appropriate machine learning models 104 and / or other software components / services.In some implementations, the machine learning models 104 can be stored in cloud storage services, container registries or other repositories that can be accessed via the external network 120. In some embodiments, one or more of the inference systems 102 can store and / or execute other software components and / or services, such as data preprocessing services, in addition to one or more machine learning models 104. Data preprocessing services can include instructions to automatically format / convert raw sensor data into a format suitable for one or more machine learning models 104. In some embodiments, data preprocessing can be performed by one or more of the inference systems 102 described herein to prepare input data for an inference task before the execution of one or more machine learning models 104. In some embodiments, and as further described herein, multiple tasks (e.g., preprocessing, inference in machine learning, etc.) can be performed.) are scheduled to be executed by one or more inference systems 102 of the local network 110 in order to perform one or more inference tasks requested by one or more client devices 108. At least one of the inference systems 102 can execute a controller 106. The controller 106 can comprise software, hardware, or a combination of hardware and software. The controller 106 can store / manage a list / data structure that includes each of the inference systems 102 of the local network 110. The controller 106 can schedule, coordinate, and, in some implementations, process one or more inference tasks transmitted by one or more client devices 108 using the inference systems 102 of the local network 110. For example, the controller 106 can receive inference tasks (e.g., via a suitable API call, etc.) from a client device 108, select one or more of the inference systems 102 to execute the inference task, and coordinate the execution of the inference task across the local network 110. In some embodiments, once the controller 106 has selected one or more inference systems 102 to process the inference task, it can transmit a network address (e.g., an API endpoint address, a Uniform Resource Locator (URL), a Uniform Resource Identifier (URI), etc.) of the selected inference system(s) 102 to the requesting client device 108. In such embodiments, the application 112 can use the provided network address to access input data and parameters for the inference task and transmit them to the selected inference system(s) 102. In some embodiments, the controller 106 can coordinate the execution of the inference task by receiving the input data and parameters from the application.In such implementations, the controller 106 can automatically transmit the input data and the parameter(s) for the inference task to the selected inference system(s) 102 via the network address(es). Further details of the processes implemented by the controller 106 are described in conjunction with Fig. 2. With reference to Fig. 2 in conjunction with the components described in connection with Fig. 1, an example diagram is shown depicting an example process 200 implemented by the controller 106 to perform or coordinate the execution of an inference task in a local network according to some embodiments of the present disclosure. As described herein, any inference system 102 can be a computer system capable of performing one or more machine learning or data processing operations (e.g., using machine learning models 104, using other software components, etc.). In step 202 of process 200, the controller 106 can identify one or more inference systems 102 connected to the local network 110 by transmitting discovery requests over the local network 110. The discovery requests can include requests regarding processing capabilities and / or other properties of the inference systems 102 (e.g., information about hardware data, stored / managed machine learning models 104, etc.). In some embodiments, the controller can send one or more mDNS queries to discover and communicate with various inference systems 102 connected to the local network 110. Each inference system 102 can respond to mDNS queries sent by the controller 106 by transmitting reply messages containing appropriate status data (e.g., available, unavailable, etc.) and processing capabilities. In some embodiments, the multicast messages can be transmitted over one or more ports of the local network 110 that are specific to the inference systems 102 (e.g., corresponding to an API endpoint managed by the inference systems 102, etc.). In some embodiments, one or more inference systems 102 can provide an API endpoint through which the controller 106 and / or the client devices 108 can access the inference systems 102. In one example, the API endpoints can provide various functions, such as retrieving information about the inference system 102, providing input data for inference tasks, coordinating the execution of inference tasks, retrieving output data generated by inference tasks, and monitoring the execution of inference tasks, among other functions. In some embodiments, the controller 106 can use these API endpoints to interact with the inference systems 102 and retrieve the processing capabilities for each identified inference system 102. In step 204, the controller 106 can create a list of inference systems 102 that may contain the processing capabilities, status, and network address of each inference system 102 identified via multicast communications transmitted on the local network 110. To this end, the controller 106 can communicate with the inference systems 102 (e.g., via their respective API endpoints / network addresses) to retrieve information regarding the processing capabilities of the inference systems 102. Examples of various processing capabilities include, but are not limited to, hardware specifications (e.g., GPU model, CPU architecture), memory capacity, available storage space, and network bandwidth. The processing capabilities may include information about which machine learning models 104 and / or software components are stored in the inference system 102.In some embodiments, the controller 106 can communicate with the inference systems 102 to coordinate the retrieval of one or more suitable machine learning models 104 from the cloud system 122 or another external computing system. For example, the controller 106 can automatically cause an inference system 102 capable of performing an inference task to automatically download or retrieve a suitable machine learning model 104 from the cloud system 122 or another external computing system. In some embodiments, the controller 106 can periodically update the list of devices by performing step 202 of method 200 to detect any changes in the number or processing capabilities of the inference devices 102 connected to the local network 110. For example, the controller can automatically and dynamically update the list of inference systems 102 when new inference systems 102 connect to or disconnect from the local network 110. In some embodiments, the controller 106 can periodically send multicast requests over the local network 110 to detect changes in processing capabilities, stored / managed machine learning models 104, or state changes in one or more of the inference systems 102 connected to the local network 110. In step 206, the controller 106 can receive one or more inference tasks from one or more client devices 108. As described here, the inference tasks can contain information about parameters for the inference task, such as the type of machine learning model 104 to be used, input data, and any additional configuration settings. Examples of types of information that can be specified as part of the inference task include the format and content of the input data (e.g., image, text, audio), thresholds or constraints for the inference process (e.g., confidence levels, processing time limits), and the format / type of the output data to be generated by the inference task. In some embodiments, the inference tasks can specify parameters / instructions for preprocessing input data (e.g., an instruction to convert input data to a specific format, etc.). In step 208, the controller 106 can analyze the received inference tasks to generate task data. For example, the controller 106 can analyze an API call provided by the application 112 of a client device to generate a data structure containing processing requests and / or parameters for the inference task. This data can be used in subsequent steps to select one or more inference systems 102 to execute the inference task. In some embodiments, analyzing a received inference task request can include validating the request and any parameters it provides. For example, the controller 106 can determine whether the inference task contains valid parameter values ​​and whether the inference task request contains any corrupted or incomplete data.In some embodiments, if the controller 106 determines that the request is invalid, it can send an error message to the requesting client device 108. In some embodiments, the controller 106 can identify specific problems with the request. In step 210, the controller 106 can use the list of inference systems 102 (and their corresponding data) as well as the analyzed data of the inference task to select one or more inference systems 102 to perform the requested inference task. The controller 106 can use any suitable selection procedure to select one or more inference systems 102 to perform the received inference task. In some embodiments, the controller 106 can rank the available inference systems 102 in the list, at least based on their processing capabilities and current utilization (e.g., available processing resources, etc.). The ranking can be determined by evaluating factors such as the hardware type (e.g., GPU model, CPU architecture), memory capacity and available memory, and the processing requirements for the requested inference task.In some embodiments, the controller 106 can order the list of inference systems based on at least some attributes of the inference task, including, but not limited to, the volume, type, or format of the input data; the type of machine learning operation to be performed; the selection of the machine learning model 104 to be executed (if selected / provided in the request); and the requested type / format of the output data to be generated by the execution of the inference task, among other things. In some embodiments, the controller 106 can prioritize inference systems 102 that are currently idle or have a lower utilization to optimize resource utilization and minimize latency. One or more inference systems 102 with the highest priority can be selected as candidate inference systems 102. To determine whether a selected candidate inference system 102 is available to execute the inference task, the controller 106 can communicate with the inference system 102 to check its current status. For example, the controller 106 can send a request to the API endpoint of the candidate inference system 102 to query its availability and responsiveness. The request can be linked to a timer corresponding to an expiration time. If the candidate inference system 102 responds with confirmation that it is available to execute the inference task within the specified timeout period, the controller 106 can select that inference system 102 to execute the inference task.If candidate inference system 102 does not respond within the timeout period or responds with an indication that inference system 102 is unavailable or lacks the processing resources to perform the inference task, the controller 106 can mark it as unavailable and select another candidate inference system 102 with the next highest priority. In some implementations, the controller 106 can select multiple inference systems 102 to perform at least part of the inference task. In some implementations, multiple inference systems 102 can be selected if an inference task involves a data volume that exceeds a threshold and / or if multiple suitable inference systems 102 are available and capable of performing the inference task. In some implementations, the controller 106 can store and manage a list (e.g., a schedule) of active or scheduled tasks that are to be executed or are currently being executed by one or more inference systems 102. The list / schedule can contain details such as the identifier of the inference system 102 assigned to each task, the task status (e.g., pending, in progress, completed), and relevant timestamps or progress indicators. In some embodiments, the controller 106 can update the schedule to display the selected inference systems 102. In some embodiments, if the controller 106 cannot select an inference system 102 for the inference task, it can send an error message back to the requesting client device 108. In step 212, the controller 106 can inform the client device 108 that an inference system 102 has been assigned to the inference task. In some embodiments, the controller 106 can coordinate the execution of the inference task. In such embodiments, the controller 106 can request, retrieve, and / or provide all input data corresponding to the inference task to the selected inference system(s) 102. In some embodiments, the controller 106 does not necessarily coordinate the execution of the inference task. In such embodiments, the controller 106 can transmit the network address(es) of the selected inference system(s) 102 to the requesting client device 108. The network addresses can be URLs or URIs corresponding to the API endpoints of the selected inference system(s).After receiving the network address(es), the application 112 of the client device 108 can transmit the input data for the inference task as well as all associated parameters for the inference task via the network address(es) in one or more requests to the selected inference systems 102 for the execution of the inference task. The selected inference system(s) 102 can automatically perform the requested inference task using the received input data and its stored machine learning models 104 and / or software components / services. The inference system 102 can receive or retrieve the input data from the controller 106 or from the requesting client device 108, as described herein. Once the input data has been received, the inference system 102 can process the data using locally stored machine learning models 104 and / or other software components / services, such as data preprocessing services. For example, the inference system 102 can run a data preprocessing service to preprocess and format the input data to match the input format for the machine learning model 104.The inference system 102 can transmit the input data to the specified machine learning model 104 to generate the requested output data. In some embodiments, several selected inference systems 102 can communicate with each other to perform different parts of the inference task. For example, one inference system 102 can perform data preprocessing operations while another inference system 102 executes the machine learning model 104. The inference systems 102 can communicate, for example, via one or more API calls or other communication protocols. The inference system(s) 102 can deliver the output data to an output location as soon as it is generated or when the task is completed. The output location can include the controller 106, a storage location specified in the request, the requesting client device 108, or any other output location. In some embodiments, the inference systems 102 can store the output data locally until it is requested by a device connected to the local network 110.The inference systems 102 can transmit the output data in real time or in batches; in some embodiments, the configuration for this can be specified in the inference task data. In some embodiments, the controller 106 can monitor the execution progress of one or more inference tasks in step 214. The controller 106 can monitor the execution of an initiated inference task by communicating with the selected inference system(s) 102 to determine the status of the inference system(s) 102 and / or the progress of the inference task. For example, the controller 106 can periodically send status queries to the API endpoint of the inference system 102 to check the current state of the task. The status queries can include, among other things, requests for the current progress, any errors encountered, the estimated time to completion, and other attributes of the inference task. In response to the status query(ies), the inference system 102 can provide one or more response messages containing information about the status of the inference task (e.g.,Percentage / degree of completion, current processing stage, execution performance metrics or logs, etc.). Control 106 can update the list / data structure of active / scheduled tasks to include the new task. Control 106 can add an entry to the list containing the identifier of inference system(s) 102, the task details, and the initial status (e.g., pending, in progress). As inference system(s) 102 makes progress on the inference task, as indicated by the progress messages of inference system(s) 102, control 106 updates the status of the inference task in the list / data structure. For example, control 106 can mark the task as "in progress" as soon as execution begins, or update the task as "completed" as soon as inference system(s) 102 indicate that the execution of the inference task is complete. In some embodiments, an inference system 102 that is performing or scheduled to perform an inference task may report an error, a system failure, or other indication that the execution of the inference task cannot be completed. In some embodiments, in response to receiving an indication that the inference task cannot be completed by the selected inference system(s) 102, the inference controller 106 may automatically perform the operations of step 210 to select another inference system (or systems) 102 to perform the inference task. In some embodiments, the controller 106 may store an indication of the error state or other information relating to the inference systems 102 that were unable to perform the inference task. In step 216, once the controller 106 determines that the execution of the inference task is complete, it can send confirmation of the successful execution of the inference task to the client device 108. In implementations where the controller 106 coordinates the execution of the inference task, the controller 106 can automatically retrieve / receive the output data of the inference task from the inference system(s) 102 and transmit the output data to the client device 108. In implementations where the application 112 coordinates the execution of the inference task, the application 112 can automatically retrieve / receive the output data of the inference task from the inference system(s) 102 and send a message to the controller 106 indicating that the inference task is complete. In some embodiments, the inference task can specify a memory location (e.g.,110 specifies a network drive (a computer system memory) on the local network where the results of the inference task are stored, instead of or in addition to returning the output data to the requesting client device 108. The controller 106 can repeatedly execute each of the steps of the procedure 200, in any order or arrangement, to perform each of the steps described herein. Figure 3 is a flowchart illustrating a method 300 for performing inference operations in local area networks. Method 300 comprises, in block B302, receiving a request for an inference task from a device (e.g., a client device 108) of a local area network (e.g., the local area network 110). The inference task may be a data processing task using one or more services and / or machine learning models (e.g., machine learning model(s) 104, etc.). The inference task may be transmitted from an application (e.g., an application 112) of the device, as described herein. In some embodiments, the device may be or include an IoT device with one or more sensors. In some embodiments, the inference task may be transmitted in response to user input at a user interface.The inference task can specify / identify input data, the operation(s) to be performed, the requested output data, or any other attribute described herein. Method 300, in block B304, comprises the selection of a first inference device from a plurality of inference devices (e.g., the inference systems 102A-102N) connected to the local network, based at least on the inference task and one or more processing capabilities of the plurality of inference devices. For this purpose, one of the operations described in connection with the controller 106 can be performed. For example, the inference devices of the local network can be ranked or prioritized based on their processing capabilities (e.g., available hardware, software, free computing resources, etc.) and the processing requirements (e.g., appropriate machine learning operations, data volume / format, etc.) of the inference task. In some embodiments, one or more inference devices can be queried or requested to provide status information.Inference devices that do not provide status information (e.g., a status signal) within a specified time (e.g., an expiration time) can be disregarded during selection. In some embodiments, the first inference device can be selected as the highest-ranked inference device available to perform the inference task. Method 300, in block B306, includes providing an identifier of the first inference device to the device on the local network. As described here, in some embodiments, a network identifier of the selected inference device can be provided to the device requesting the execution of the inference task. The requesting device can use the network address as an endpoint to send the input data and / or parameters of the inference task for execution. In some implementations, the computing device executing Method 300 (e.g., the controller 106, etc.) can coordinate the execution of the inference task by transmitting the input data and / or parameters of the inference task to the first inference device. The first inference device can receive the input data and perform the operations of the inference task to generate output data.The output data can be provided to the requesting device, stored in a specific memory location, or stored locally on the first inference device for execution. The systems and methods described herein can be used for a wide variety of purposes, including but not limited to the definition of circuit layouts, machine control, machine locomotion, machine propulsion, synthetic data generation, model training, perception, augmented reality, virtual reality, mixed reality, robotics, security and monitoring, simulation and digital twinning, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twinning, data center processing, conversational artificial intelligence (AI), light transport simulation (e.g., ray tracing, path tracing, etc.), collaborative content creation for three-dimensional (3D) assets, cloud computing, generative AI, and / or any other suitable application. The disclosed embodiments can be included in a variety of different systems, such as automotive systems (e.g., a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), systems implemented using a robot, aeronautical systems, medical systems, boat systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing digital twin operations, systems implemented using a device, systems containing one or more virtual machines (VMs), systems for performing operations to generate synthetic data, systems implemented at least partially in a data center, systems for performing conversational AI operations, and systems...that implement one or more language models – such as one or more large language models (LLMs), one or more small language models (SLMs), one or more image language models (VLMs), one or more multimodal language models (MMLMs), systems for performing light transport simulations, systems for performing collaborative content creation for 3D assets, systems that are implemented at least partially using cloud computing resources, and / or other types of systems. EXAMPLE OF A COMPUTER DEVICE Fig. 4 shows a block diagram of an example of a device 400 suitable for implementing some embodiments of the present disclosure. The device 400 may comprise a connection system 402 that directly or indirectly connects the following devices: a memory 404, one or more central processing units (CPUs) 406, one or more graphics processing units (GPUs) 408, a communication interface 410, input / output (I / O) ports 412, input / output components 414, a power supply 416, one or more presentation components 418 (e.g., display(s)), and one or more logic units 420. In at least one embodiment, the devices 400 may comprise one or more virtual machines (VMs), and / or each of the components thereof may comprise virtual components (e.g., virtual hardware components).As non-restrictive examples, one or more of the GPUs 408 can comprise one or more vGPUs, one or more of the CPUs 406 can comprise one or more vCPUs, and / or one or more of the logic units 420 can comprise one or more virtual logic units. Accordingly, a computing device 400 can comprise discrete components (e.g., a complete GPU allocated to the computing device 400), virtual components (e.g., a portion of a GPU allocated to the computing device 400), or a combination thereof. Although the various blocks in Fig. 4 are shown connected via the connection system 402, this is not intended to be restrictive and serves only for clarity. For example, in some embodiments, a presentation component 418, such as a display device, can be considered part of an I / O component 414 (e.g., if the display is a touchscreen). As another example, the CPUs 406 and / or GPUs 408 can include memory (e.g., the memory 404 can represent a storage device in addition to the memory of the GPUs 408, the CPUs 406, and / or other components). In other words, the computing device in Fig. 4 serves only for illustration.No distinction is made between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "game console", "electronic control unit (ECU)", "virtual reality system" and / or other device or system types, as all are considered within the scope of the computing device shown in Fig. 4. The 402 interconnect system can represent one or more connections or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The 402 interconnect system can include one or more bus or connection types, such as an ISA (Industry Standard Architecture) bus, an EISA (Extended Industry Standard Architecture) bus, a VESA (Video Electronics Standards Association) bus, a PCI (Peripheral Component Interconnect) bus, a PCIe (Peripheral Component Interconnect Express) bus, and / or another bus or connection type. In some embodiments, there are direct connections between components. For example, the CPU 406 can be directly connected to the memory 404. Furthermore, the CPU 406 can be directly connected to the GPU 408.If a direct or point-to-point connection exists between components, the 402 connection system can include a PCIe connection to establish the connection. In these examples, the 400 computing device does not need to include a PCI bus. The memory 404 can comprise a variety of computer-readable media. The computer-readable media can be any available media to which the device 400 can access. The computer-readable media can include both volatile and non-volatile media, as well as removable and non-removable media. For example, and without limitation, the computer-readable media can include computer storage media and communication media. Computer storage media can include both volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, the memory 404 can store computer-readable instructions (e.g., those representing a program or programs and / or a program element or elements, such as an operating system). Computer storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, Digital Versatile Discs (DVDs) or other optical storage media, magnetic cartridges, magnetic tapes, magnetic disks or other magnetic storage devices, and any other medium that can be used to store the desired information and that the device 400 can access.For the purposes of this description, computer storage media do not include signals per se. Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transmission mechanism, and encompass any information transmission medium. The term "modulated data signal" can refer to a signal in which one or more of its properties have been set or modified to encode information within the signal. For example, but not limited to, computer storage media can include wired media such as a wired network or a directly wired connection, as well as wireless media such as acoustic, RF, infrared, and other wireless media. Combinations of any of the above media should also fall within the scope of computer-readable media. The CPU(s) 406 can be configured to execute at least some of the computer-readable instructions to control one or more components of the Computing Device 400 to execute one or more of the procedures and / or methods described herein. The CPU(s) 406 can each comprise one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing a plurality of software threads simultaneously. The CPU(s) 406 can comprise any processor type and, depending on the nature of the Device 400 implemented, may include different processor types (e.g., processors with fewer cores for mobile devices and processors with more cores for servers).For example, depending on the type of Device 400, the processor may be an Advanced RISC Machines (ARM) processor implemented using Reduced Instruction Set Computing (RISC), or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or complementary coprocessors, such as mathematical coprocessors, the Device 400 may include one or more CPUs 406. In addition to or as an alternative to the CPU(s) 406, the GPU(s) 408 may be configured to execute at least some of the computer-readable instructions to control one or more components of the device 400 to perform one or more of the methods and / or processes described herein. One or more of the GPU(s) 408 may be an integrated GPU (e.g., with one or more of the CPU(s) 406), and / or one or more of the GPU(s) 408 may be a discrete GPU. In embodiments, one or more of the GPU(s) 408 may be a coprocessor of one or more of the CPU(s) 406. The GPU(s) 408 may be used by the computing device 400 for rendering graphics (e.g., 3D graphics) or for performing general-purpose computations. For example, the GPU(s) 408 may be used for general-purpose GPU computing (GPGPU).The GPU(s) 408 can comprise hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU(s) 408 can generate pixel data for output images in response to rendering commands (e.g., rendering commands from the CPU(s) 406 received via a host interface). The GPU(s) 408 can include graphics memory, such as display memory, for storing pixel data or other suitable data, such as GPGPU data. The display memory can be included as part of the memory 404. The GPU(s) 408 can comprise two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using NVLINK) or connect them via a switch (e.g., using NVSwitch).In combination, each GPU can generate 408 pixel data or GPGPU data for different parts of an output or for different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can have its own dedicated memory or share memory with other GPUs. In addition to or as an alternative to the CPU(s) 406 and / or the GPU(s) 408, the logic unit(s) 420 may be configured to execute at least some of the computer-readable instructions to control one or more components of the device 400 to perform one or more of the methods and / or processes described herein. In embodiments, the CPU(s) 406, the GPU(s) 408, and / or the logic unit(s) 420 may execute any combination of the methods, processes, and / or parts thereof, individually or jointly. One or more of the logic units 420 may be part of and / or integrated into one or more of the CPU(s) 406 and / or the GPU(s) 408, and / or one or more of the logic units 420 may be discrete components or otherwise separate from the CPU(s) 406 and / or the GPU(s) 408.In embodiments, one or more of the logic units 420 can be a coprocessor of one or more of the CPU(s) 406 and / or one or more of the GPU(s) 408. Examples of logic unit(s) 420 include one or more processing core(s) and / or components thereof, such as data processing units (DPUs), tensor cores (TCs), tensor processing units (TPUs), pixel visual cores (PVCs), image processing units (VPUs), graphics processing clusters (GPCs), texture processing clusters (TPCs), streaming multiprocessors (SMs), tree traversal units (TTUs), artificial intelligence accelerators (AIAs), deep learning accelerators (DLAs), arithmetic logic units (ALUs), application-specific integrated circuits (ASICs), floating-point units (FPUs), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements, and / or the like. The communication interface 410 can comprise one or more receivers, transmitters, and / or transceivers that enable the device 400 to communicate with other devices via an electronic communication network, including wired and / or wireless communication. The communication interface 410 can include components and functions to enable communication over a variety of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication via Ethernet or InfiniBand), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet. In one or more embodiments, the logic unit(s) 420 and / or the communication interface 410 can comprise one or more data processing units (DPUs) to directly transmit data received via a network and / or the connection system 402 to (e.g., the communication interface 400).B. a memory of) one or more GPU(s) 408 to transfer. The I / O ports 412 enable the computing device 400 to be logically coupled with other devices, including the I / O components 414, the presentation component(s) 418, and / or other components, some of which may be built into (e.g., integrated with) the device 400. Examples of I / O components 414 include a microphone, mouse, keyboard, joystick, gamepad, game controller, satellite dish, scanner, printer, wireless device, etc. The I / O components 414 can provide a natural user interface (NUI) that processes air gestures, speech, or other physiological inputs generated by a user. In some cases, inputs can be transferred to a suitable network element for further processing.A NUI can implement any combination of speech recognition, pen recognition, facial recognition, biometric recognition, gesture recognition both on-screen and off-screen, air gestures, head and eye tracking, and touch recognition (as further described below) in conjunction with a display of the device 400. The computer device 400 can include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof, for gesture detection and recognition. Additionally, the device 400 can include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit (IMU)) that enable motion detection. In some embodiments, the output of the accelerometers or gyroscopes can be used by the device 400 to display immersive augmented reality or virtual reality. The power supply 416 can be a hardwired power supply, a battery power supply, or a combination thereof. The power supply 416 can supply power to the device 400 to enable the operation of the device 400's components. The presentation component(s) 418 may include a display (e.g., a monitor, a touchscreen, a television screen, a head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component(s) 418 may receive data from other components (e.g., the GPU(s) 408, the CPU(s) 406, DPUs, etc.) and output the data (e.g., as an image, video, sound, etc.). EXAMPLE OF A DATA CENTER Figure 5 illustrates an example of a data center 500 that can be used in at least one embodiment of the present disclosure. The data center 500 can comprise a data center infrastructure layer 510, a framework layer 520, a software layer 530, and / or an application layer 540. As shown in Fig. 5, the data center infrastructure layer 510 can include a resource orchestrator 512, clustered compute resources 514 and node compute resources (“node CRs”) 516(1)-516(N), where “N” is any positive integer. In at least one embodiment, the node compute resources 516(1)-516(N) can include, among other things, any number of central processing units (CPUs) or other processors (including DPUs, accelerators, field-programmable gate arrays (FPGAs), graphics processing units or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or hard disk drives), network input / output devices (NW I / O), network switches, virtual machines (VMs), power supply modules and / or cooling modules, etc. In some embodiments, one or more node CRs from the group of node CRs can be included.s 516(1)-516(N) correspond to a server that has one or more of the computing resources mentioned above. Furthermore, in some embodiments, the node CRs 516(1)-516(N) may include one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or one or more of the node CRs 516(1)-516(N) may correspond to a virtual machine (VM). In at least one embodiment, grouped compute resources 514 can comprise separate groupings of node CRs 516 housed in one or more racks (not shown), or many racks housed in data centers at different geographic locations (also not shown). Separate groupings of node CRs 516 within grouped compute resources 514 can comprise grouped compute, network, storage, or memory resources that can be configured or allocated to support one or more workloads. In at least one embodiment, multiple node CRs 516, including CPUs, GPUs, DPUs, and / or other processors, can be grouped within one or more racks to provide compute resources to support one or more workloads.The one or more racks can also include any number of power supply modules, cooling modules and / or network switches in any combination. The resource orchestrator 512 can configure or otherwise control one or more node CRs 516(1)-516(N) and / or grouped computing resources 514. In at least one embodiment, the resource orchestrator 512 can include a software design infrastructure (SDI) management unit for the data center 500. The resource orchestrator 512 can comprise hardware, software, or a combination thereof. In at least one embodiment, as shown in Fig. 5, the framework layer 520 can include a job scheduler 528, a configuration manager 534, a resource manager 536, and / or a distributed file system 538. The framework layer 520 can include a framework to support the software 532 of the software layer 530 and / or one or more applications 542 of the application layer 540. The software 532 or the application(s) 542 can each include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 520 can be, but is not limited to, a type of free and open-source web application framework, such as Apache Spark™ (hereinafter “Spark”), which can utilize a distributed file system 538 for processing large amounts of data (e.g., “Big Data”).In at least one embodiment, the job scheduler 528 can include a Spark driver to facilitate the scheduling of workloads supported by different layers of the data center 500. The configuration manager 534 can be capable of configuring various layers, such as the software layer 530 and the framework layer 520, including Spark and the distributed file system 538, to support large-scale data processing. The resource manager 536 can be capable of managing clustered or grouped compute resources allocated or assigned to support the distributed file system 538 and the job scheduler 528. In at least one embodiment, clustered or grouped compute resources can include grouped compute resources 514 on the data center infrastructure layer 510.The Resource Manager 536 can coordinate with the Resource Orchestrator 512 to manage these allocated or assigned computing resources. In at least one embodiment, the software 530 contained in the software layer 532 may comprise software used by at least parts of the node CRs 516(1)-516(N), grouped computing resources 514, and / or the distributed file system 538 of the framework layer 520. One or more types of software may include, among others, internet web search software, email virus scanning software, database software, and video streaming software. In at least one embodiment, the applications 542 contained in the application layer 540 can comprise one or more types of applications used by at least parts of the node CRs 516(1)-516(N), the grouped compute resources 514, and / or the distributed file system 538 of the framework layer 520. One or more types of applications can include, among others, any number of genomics applications, cognitive computing applications, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments. In at least one embodiment, the configuration manager 534, the resource manager 536, or the resource coordinator 512 can implement any number and type of self-modifying actions based on any set and type of data acquired in any technically feasible way. Self-modifying actions can relieve the data center operator 500 of potentially making erroneous configuration decisions and help avoid underutilization and / or underperforming areas of a data center. The Data Center 500 may include tools, services, software, or other resources to train one or more machine learning models or to predict or derive information using one or more machine learning models according to one or more of the embodiments described herein. For example, one or more machine learning models may be trained by calculating weight parameters according to a neural network architecture using software and / or computing resources described above in relation to the Data Center 500.In at least one embodiment, trained or provided machine learning models corresponding to one or more neural networks can be used to derive or predict information using the resources described above in relation to the Computing Center 500, by using weight parameters calculated by one or more training techniques, such as, but not limited to, those described herein. In at least one embodiment, the data center can use 500 CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or equivalent virtual computing resources) to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above can be configured as a service to allow users to train or infer information, such as image recognition, speech recognition, or other artificial intelligence services. EXEMPLARY NETWORK ENVIRONMENTS Network environments suitable for implementing embodiments of the disclosure may include one or more client devices, servers, network-attached storage (NAS) devices, other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be implemented on one or more instances of the computing devices 400 from Fig. 4—e.g., each device may have similar components, features, and / or functions to the computing devices 400. Furthermore, backend devices (e.g., servers, NAS, etc.), if implemented, may be included as part of a data center 500, an example of which is described in more detail here in connection with Fig. 5. Components in a network environment can communicate with each other over one or more networks, which can be wired, wireless, or both. The network can comprise multiple networks or a network of networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the internet and / or a public switched telephone network (PSTN), and / or one or more private networks. If the network includes a wireless telecommunications network, components such as a base station, a communications mast, or even access points (as well as other components) can provide wireless connectivity. Compatible network environments can include one or more peer-to-peer network environments—in which case no server may be present—and one or more client-server network environments—in which case one or more servers may be present. In peer-to-peer network environments, the functionality described here with respect to one or more servers can be implemented on any number of client devices. In at least one embodiment, a network environment can comprise one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. A cloud-based network environment can comprise a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which can include one or more core network servers and / or edge servers. A framework layer can comprise a framework for supporting software of a software layer and / or one or more applications of an application layer. The software or applications can each comprise web-based service software or applications. In embodiments, one or more of the client devices can utilize the web-based service software or applications (e.g.,through access to the service software and / or applications via one or more application programming interfaces (APIs). The framework layer can be, but is not limited to, a type of web application framework for free and open-source software, such as a distributed file system for processing large amounts of data (e.g., "Big Data"). A cloud-based network environment can provide cloud computing and / or cloud storage, performing any combination (or one or more parts) of the computing and / or data storage functions described herein. Each of these various functions can be distributed across multiple locations by central or core servers (e.g., one or more data centers, which may be distributed across a state, region, country, the globe, etc.). If a connection to a user (e.g., a client device) is located relatively close to one or more edge servers, one or more core servers can delegate at least some of the functionality to the edge server(s). A cloud-based network environment can be private (e.g., restricted to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment). The client device(s) may include at least some of the components, features and functions of the example computer devices 400 described herein with reference to Fig. 4.For example, and without limitation, a client device may be a personal computer (PC), laptop, mobile device, smartphone, tablet computer, smartwatch, portable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, video camera, surveillance device or system, vehicle, boat, aircraft, virtual machine, drone, robot, portable communications device, hospital device, gaming device or system, entertainment system, vehicle computer system, embedded control system, remote control, household appliance, consumer electronics device, workstation, edge device, any combination of these devices, or any other suitable device. The disclosure can be described in the general context of computer code or machine-readable instructions, including computer-executable instructions such as program modules that are executed by a computer or other machine, such as a personal digital assistant or other handheld device. In general, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs specific tasks or implements certain abstract data types. The disclosure can be implemented in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The disclosure can also be implemented in distributed computing environments where tasks are performed by remote processing devices interconnected via a communication network. As used here, a statement of "and / or" in relation to two or more elements is to be interpreted as meaning only one element or a combination of elements. For example, "element A, element B and / or element C" can include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Furthermore, "at least one of element A or element B" can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Additionally, "at least one of element A and element B" can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. The subject matter of this disclosure is described in detail herein to satisfy legal requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have considered that the claimed subject matter could also be realized in other ways, incorporating various steps or combinations of steps similar to those described in this document, in conjunction with other present or future technologies. Furthermore, the terms "step" and / or "block," although used here to denote various elements of the methods employed, should not be interpreted as implying a particular sequence among or between the various steps disclosed herein, unless the sequence of each step is expressly described. It is understood that the aspects and embodiments described above serve only as examples and that detailed modifications can be made within the scope of the claims. Each device, method and feature disclosed in the description, as well as (where applicable) the claims and drawings, may be provided independently of each other or in any suitable combination. The reference numerals listed in the claims are for illustrative purposes only and do not restrict the scope of the claims.

Claims

One or more processors comprising: one or more circuits serving to: receive a request for an inference task from a device of a local network; select a first inference device from a plurality of inference devices connected to the local network, based at least on the inference task and one or more processing capabilities of the plurality of inference devices; and provide a specification of the first inference device to the device of the local network in response to the request. One or more processors according to claim 1, wherein the one or more circuits serve to: store a list of identifiers of the multiple inference devices in association with the one or more processing capabilities; and select the first inference device according to an order of the list of identifiers. One or more processors according to claim 2, wherein the one or more circuits serve to: update the list of identifiers of the plurality of inference devices at least on the basis of one or more responses to a multicast message transmitted over the local network. One or more processors according to any of the preceding claims, wherein the one or more circuits serve to: provide a network address of the first inference device as part of the specification. One or more processors according to one of the preceding claims, wherein the one or more circuits serve to: receive data relating to at least one machine learning model stored on the first inference device from the first inference device; and update the one or more processing capabilities of the plurality of inference devices at least based on the data relating to the at least one machine learning model. One or more processors according to any one of the preceding claims, wherein the one or more circuits serve to: determine that the first inference device comprises sufficient computing resources for the inference task; and select the first inference device in response to the determination that the first inference device comprises sufficient computing resources for the inference task. One or more processors according to any one of the preceding claims, wherein the one or more circuits serve to: transmit a message corresponding to the inference task to a second inference device from the plurality of inference devices; determine that the second inference device has failed to provide a response to the message within a time limit; and select the first inference device when it is determined that the second inference device has not provided the response within the time limit. One or more processors according to any of the preceding claims, wherein the one or more circuits serve to: receive data for the inference task from the device of the local network; provide the data for the inference task to the first inference device in response to the selection of the first inference device; and receive output information generated by the first inference device according to the inference task from the first inference device. One or more processors according to one of the preceding claims, wherein the one or more circuits serve to: create a schedule for a plurality of inference tasks received from a plurality of devices via the local network. One or more processors according to any one of the preceding claims, wherein the one or more processors are included in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing operations in the field of conversational AI; a system for performing generative AI operations using a small language model (SLM);a system for performing generative AI operations using a large language model (LLM); a system for performing generative AI operations using a picture language model (VLM); a system for performing generative AI operations using a multimodal language model (MMLM); a system for generating synthetic data; a system comprising one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is at least partially implemented using cloud computing resources. System comprising: a plurality of inference devices communicating with each other via a local network, each of the plurality of inference devices storing a respective machine learning model; a control device connected to the local network, the control device serving to: create a list of the multiple inference devices connected to the local network via communication with the local network; receive a specification of an inference task from a client device; select an inference device from the plurality of inference devices, based at least on the inference task and the respective machine learning model stored on each of the plurality of inference devices; and generate a response to the request, based at least on the selected inference device. System according to claim 11, wherein the control device serves to: provide a local network address of the selected inference device in response to the request. System according to claim 11 or 12, wherein the control device serves to: communicate data for the inference task to the selected inference device; monitor the execution of the inference task on the selected inference device; and provide an output of the inference task generated by the selected inference device to the client device. System according to one of claims 11 to 13, wherein the control device is included in the plurality of inference devices. System according to any one of claims 11 to 14, wherein the selected inference device is a first inference device and wherein the control device serves to: determine that a second inference device from the plurality of inference devices is not available for performing the inference task, based on at least one second network communication; and select the first inference device, based at least on the fact that the second inference device is not available. System according to any one of claims 11 to 15, wherein the system comprises at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing operations in the field of conversational AI; a system for performing generative AI operations using a small language model (SLM); a system for performing generative AI operations using a large language model (LLM);a system for performing generative AI operations using a visual language model (VLM); a system for performing generative AI operations using a multimodal language model (MMLM); a system for generating synthetic data; a system comprising one or more virtual machines (VMs); a system that is at least partially implemented in a data center; or a system that is at least partially implemented using cloud computing resources. A method comprising: receiving a request for an inference operation from a device of a local network using one or more processors; selecting a first inference device from a plurality of inference devices connected to the local network based on at least one of the inference operation or one or more processing capabilities of the plurality of inference devices using the one or more processors; and assigning at least part of the inference operation to the first inference device using the one or more processors. The method of claim 17, comprising: storing a list of identifiers of the plurality of inference devices in assignment to the one or more processing capabilities using the one or more processors; and selecting the first inference device according to an order of the list of identifiers using the one or more processors. The method of claim 18, further comprising: updating the list of identifiers of the plurality of inference devices based on at least one or more responses to a multicast message transmitted over the local network using the one or more processors. A method according to any one of claims 17 to 19, further comprising: providing a specification of the first inference device to the device of the local network in response to the request using one or more processors, wherein the specification includes a network address of the first inference device.