State-based distribution and routing of a stream for intelligent resource allocation in multi-sensor systems and applications

The SDR system addresses inefficiencies in processing continuous data streams by using a routing table to manage and allocate resources dynamically, ensuring efficient and scalable data processing with reduced errors and improved adaptability.

DE102025112840A1Pending Publication Date: 2025-10-23NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102025112840
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-22
Filing Date
2025-04-02
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Conventional data processing systems face challenges in efficiently managing and processing continuous, real-time data streams, particularly in distributed environments, leading to inefficiencies and potential errors when adding new features or services.

Method used

A power distribution and routing (SDR) system that utilizes a routing table to track and manage data streams, dynamically allocate resources, and integrate new microservices with reduced encoding requirements, ensuring efficient and scalable data processing by maintaining a state-based data structure and monitoring network anomalies.

Benefits of technology

The SDR system optimizes resource utilization, ensures fast scalability and adaptability, reduces the risk of errors, and maintains data integrity by dynamically managing resources and monitoring network health, enabling efficient processing of multiple continuous and real-time data streams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The disclosed approaches relate to the distribution and processing of stream data in a networked computing environment. A set of streaming data sources can be identified for a given cluster of nodes, and the data from these data sources can be tracked, managed, and / or routed to a number of different microservices. Stream distribution and routing (SDR) agents can be used to maintain a routing table that includes an identifier for each of the data sources. The individual agents can distribute information about the sensors to individual microservices, which can determine which data streams are relevant and then initiate individual processes, procedures, and functions using the relevant data streams.Output results from individual microservices can also be assigned an identifier, which can be further used by individual agents to route outputs to their specific microservices, which can then use the outputs from other microservices within their own workstreams.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Distributed computing offers a powerful approach to performing complex tasks, such as those involving the processing and management of large amounts of data. This approach leverages a network of interconnected computers working together to process data by distributing the workload across multiple nodes to improve efficiency and speed. A critical component of various distributed computing approaches is a container orchestration platform, such as Kubernetes. These platforms manage distributed computing by automating the deployment, scaling, and management of containerized applications within a distributed network.Distributed data processing can be used to process power data, such as data from sensor-based systems, which may require real-time analysis and coordination of power data from a multitude of sensors. In such systems, efficiently managing continuous data streams from each independent sensor presents a challenge that traditional data processing models struggle to address. BRIEF DESCRIPTION OF THE DRAWINGS

[0002] Various embodiments according to the present disclosure are described with reference to the drawings, in which: Fig. 1 An exemplary scenario for the use of a stream distribution and routing (SDR) system to provide rendered results is illustrated according to at least one embodiment; Fig. 2 a diagram illustrating components in an exemplary power distribution and routing system, according to at least one embodiment; Fig. 3 illustrates an exemplary SDR agent and an exemplary workload object, according to at least one embodiment; Fig. Four exemplary data flows in a power distribution and routing system are illustrated, according to at least one embodiment; Fig. 5 illustrates an exemplary process or method for distributing and routing currents using a current distribution and routing system, according to at least one embodiment; Fig. 6 illustrates an exemplary system environment that includes a power distribution and routing system, according to at least one embodiment; Fig. 7. An exemplary data center system is illustrated, according to at least one embodiment; Fig. 8 is a block diagram illustrating a computer system, according to at least one embodiment; Fig. 9 is a block diagram illustrating a computer system, according to at least one embodiment; Fig. 10 illustrates a computer system according to at least one embodiment; Fig. 11 illustrates a computer system according to at least one embodiment; Fig. 12 exemplary integrated circuits and associated graphics processors illustrated, according to at least one embodiment; Fig. 13A, Fig. 13B Illustrating exemplary integrated circuits and associated graphics processors according to at least one embodiment; Fig. 14 illustrates a computer system according to at least one embodiment; Fig. 15A illustrates a parallel processor according to at least one embodiment; Fig. 15B illustrates a partition unit according to at least one embodiment; and Fig. 16 at least sections of a graphics processor are illustrated, according to one or more embodiments. DETAILED DESCRIPTION

[0003] The following description details various embodiments. For illustrative purposes, specific configurations and details are presented to provide a thorough understanding of these embodiments. However, it is also obvious to those skilled in the art that the embodiments can be implemented without these specific details. Furthermore, known features may be omitted or simplified so as not to obscure the described embodiment.

[0004] The systems and procedures described herein may be used without restriction by non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more advanced driver assistance systems (ADAS)), controlled and uncontrolled robots or robotic platforms, warehouse vehicles, all-terrain vehicles, vehicles coupled to one or more trailers, flying airships, boats, shuttles, emergency vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, trains, underwater vehicles, remotely controlled vehicles such as drones, and / or other types of vehicles.Furthermore, the systems and methods described herein can be used for a variety of purposes, including but not limited to machine control, machine locomotion, machine driving, synthetic data generation, model training or updating, perception, augmented reality, virtual reality, mixed reality, robotics, security and surveillance, simulation and digital twin creation, autonomous or semi-autonomous machine applications, deep learning, environmental simulation, object or actor simulation and / or digital twin creation for an object or actor, data center processing, comparative AI, light transport simulation (e.g., ray tracing, path tracking, etc.), collaborative content creation for 3D assets, cloud computing, and / or any other suitable applications.

[0005] Disclosed embodiments may be included in a wide variety of different systems, such as automotive systems (e.g., a control system for an autonomous or semi-autonomous machine, a perception system for an autonomous or semi-autonomous machine), systems implemented using a robot, aviation systems, medical systems, boat systems, intelligent area monitoring systems, systems for performing deep learning operations, systems for performing simulation operations, systems for performing operations with digital twins, systems implemented using an edge device, systems incorporating one or more virtual machines (VMs), systems for performing synthetic data generation operations, systems implemented at least partially in a data center, systems for performing conversational AI operations, and systems for performing light transport simulation.Systems for performing collaborative content creation for 3D assets, systems that are implemented at least partially using cloud computing resources, and / or other types of systems.

[0006] Approaches according to various embodiments of the disclosure are directed toward the distribution and routing of data streams. For example, a stream distribution and routing (SDR) system according to at least one embodiment can distribute and process data streams in a networked computer or "cloud" environment. Such a system can identify a set of streaming data sources for a given cluster of nodes (such as computers, servers, or other computing devices) and then track, manage, and route the data from these data sources to a number of different microservices associated with the cluster of nodes. An initiator can publish information from individual data sources (e.g., sensors) over a bus and provide at least one message to one or more agents, each associated with one or more specific microservices.The agents can be used to manage a state-based data structure (e.g., maintain, update, etc.), such as a routing table (or something that could be called a route entry table), which contains an identifier for each of the data sources. The individual agents can then distribute information about the data sources to their respective microservices, which can determine which data streams are relevant to those microservices. The microservices can then initiate individual processes and functions using the relevant data streams. In certain embodiments, outputs from the individual microservices can also be assigned identifiers, which the individual agents can further use to route the outputs to their specific microservices, which can then use the outputs (from other microservices) within their own workflows.In this way, individual clusters of nodes can identify, track, and use different data sources within groups of microservices managed by individual agents.

[0007] Approaches according to at least one embodiment can provide several technical advantages and improvements. For example, the power distribution and routing system according to at least one embodiment can efficiently process multiple continuous and real-time data streams. Conventional systems, which are typically transaction-based, often face challenges in handling continuous data streams in real time. For example, conventional data processing models require coordination in distributed systems by sending and processing individual requests, which is not well-suited for continuous stream processing, which requires uninterrupted and ongoing data handling. The power distribution and routing system overcomes this challenge by managing multiple running processes simultaneously with a tracked state or status for each process and data stream.The system can dynamically identify and allocate resources for each data stream using a routing table. By employing such a routing table, the stream distribution and routing system can provide a comprehensive understanding of which process is handling which stream at any given time, advantageously facilitating or enabling parallel processing and tracking in a continuously operating environment.

[0008] Furthermore, according to at least one embodiment, the power distribution and routing system can provide an optimized process for adding new features, such as microservices, with reduced coding requirements. This can be achieved by integrating one or more SDR (power distribution and routing) agents into the SDR system. For example, if a user wants to add new features to an existing network cluster, the user can create a new feature, and the SDR system attaches an SDR agent to a corresponding workload object configured to execute the new feature. The system can simplify the deployment of various functionalities or microservices and acts as a robust backend component, enhancing the capabilities of large-scale data processing.Traditional systems, on the other hand, often involve coding and reconfiguring multiple interdependent components when introducing new features or services, making each addition a potential risk for introducing errors or incompatibilities. As a result, power distribution and routing approaches, as revealed herein, offer solutions for agile development and deployment with rapid scalability and adaptability in response to evolving data processing needs. Such approaches can ensure that new services can be dynamically integrated and existing ones efficiently updated.

[0009] Furthermore, approaches to power distribution and routing, as revealed herein, can optimize resource utilization efficiency by managing power density and resource allocation across the network. For example, the SDR system can achieve this by intelligently allocating data streams across the array of available microservices, or pods, within workload objects. The SDR system can dynamically allocate streams to workload objects based on the maximum capacity requirements of each assigned workload object by routinely recalculating performance metrics (such as utilized capacity), thus optimizing power density across the network. To support system reliability, the SDR system can maintain a minimum number of backup copies that can take over processing if an active pod fails.The SDR system can improve reliability by monitoring for anomalies or failures within the network, such as failed processes, offline nodes, or old stream identifiers. Additionally, the SDR system can include mechanisms for reconciling old stream identifiers with the initiator. This is achieved by requiring each initiator to maintain an up-to-date list of active stream identifiers, ensuring that the routing system remains responsive to the actual network state. Such mechanisms ensure that data processing is both current and accurate, reducing the risk of data loss or processing delays and maintaining the integrity and continuity of data streams within the distributed data processing environment.

[0010] Variations of these and other such functionalities can also be used within the scope of the various embodiments, as is obvious to the person skilled in the art in view of the teachings and proposals contained herein.

[0011] Fig. Figure 1 presents an illustrative diagram of a stream distribution and routing (SDR) system 110 according to at least one embodiment. In this example, the SDR system 110 is illustrated to connect to a network of sensors within a Kubernetes cluster. The SDR system 110 can receive real-time data streams from various sensors, such as Sensor A, Sensor B, and Sensor C. These sensors can correspond to any of a diverse range of devices, such as video cameras or radar systems, with each stream of data (live or otherwise) being processed by the SDR system 110. The SDR system 110 can dynamically distribute these streams to suitable workload objects (which may be referred to here as workers) within the cluster.These workload objects can correspond to different microservices selected to perform specific functions on the incoming data streams. In the [document / section]... Fig. In the illustrated example scenario 1, Sensor A, Sensor B, and Sensor C can represent high-resolution cameras used in a secured facility. Each sensor can be tasked with continuously monitoring different critical areas. In one embodiment, an initiator, such as a video storage toolkit (VST) in a security system, can activate or deactivate sensors (e.g., cameras) and their respective streams. Each camera can be assigned a unique identifier and can generate and stream real-time video data. The initiator can generate an event message in a specific format, including one or more unique identifiers, to activate or deactivate one or more corresponding sensors and announce the message via a configuration control bus.Upon detecting this message, SDR agents can capture the message and assign the stream to a suitable pod of a workload object capable of handling it. In one embodiment, the SDR system can process 110 video streams in real time and route them to designated microservices for specific processing, such as AI inferencing (e.g., motion detection or object detection). The SDR agents can then configure appropriate workload objects to process the video streams from the cameras and record mappings in a routing table. For example, the routing table can record a mapping of a unique identifier of a camera (e.g., the camera's stream ID) to a unique identifier of a pod (e.g., the pod's IP address).After processing by the workload objects, results are forwarded to various result generation or playback endpoints, such as result playback A and result playback B, which may be visual displays or other forms of result generation that provide users with processed information in real time, such as identified threats or tracking movements.

[0012] Fig. Figure 2 illustrates an exemplary SDR system within a cloud cluster 200 according to at least one embodiment. The exemplary SDR system can include an SDR controller 240, an initiator 210, a configuration control bus 250, an application data bus 260, and a number of SDR agents 221, 222, and 223, each SDR agent being associated with a respective workload object, such as workload objects 231, 232, and 233. These components of the SDR system are discussed in more detail below.

[0013] The initiator 210 can correspond to or execute a process, procedure, or function hosted in a cluster that maintains a list of sensors and manages a state or status corresponding to each sensor. The initiator 210 can enable or disable one, some, or all of the sensors (and their respective streams). In one embodiment, the initiator 210 can communicate with the SDR agents 221, 222, and 223 via the configuration control bus 250 using messages that follow or have a predefined message format (e.g., a JSON file with specific fields). The initiator 210 can send these structured messages, which encapsulate information associated with each data source.Each message can contain information such as the sensor type, one or more unique identifiers, and metadata that describes the nature and requirements of the data stream. An example message might follow or have the following format: “Sensor”: {“alert_type”: “camera_status_change”, “created_at”: “2023-03-10T00:45:16Z”, “event”: {“camera_id”: “302547db-d661-4fb0-9e46-62b7420da904”, “camera_name”: “front_door”, “camera_url”: “rtsp: / / . <camera-ipaddress> / ...", „change": „camera_add or camera_remove", „metadata": {„resolution": „1920 x1080", „codec": „h264", „framerate": 30}, „headers": {„source": „vst", „created_at": „2021 - 06 - 01 T14: 34: 13.417 Z"}}

[0014] For example, as illustrated in the sample message, when integrating a new set of surveillance cameras within a security network, the initiator 210 can publish a message containing information such as a camera type, a unique stream ID (identifier) ​​for each camera, resolution details, a frame rate, and the specific codec used for video encoding. In one embodiment, the initiator can assign a unique stream ID (such as a `camera_id`) to each data stream along with any other necessary metadata (e.g., the event field in the sample message). The message can be published to the configuration control bus 250, where it is made available for the SDR agents to receive. This message can be announced by the configuration control bus 250 to notify the SDR agents to receive the messages.Upon receiving the message, the SDR agents can route the stream data to microservices (e.g., pods in workload objects in a Kubernetes environment) that are suitable for specific tasks, such as AI inferencing, object detection, or motion tracking.

[0015] Workload objects 231, 232, and 233 can correspond to processing entities within a distributed cluster of nodes. As used here, a workload object can refer to a high-level abstraction representing a set of applications or microservices running on the cluster. These workload objects can be tasked with executing specific functions / microservices and generating output by processing data streams. In one embodiment, workload objects can be instantiated as process copies (commonly referred to as pods within the Kubernetes environment). That is, each workload object can manage a multitude of processes running from a multitude of pods corresponding to the workload object. Each pod can be uniquely identifiable within the network (e.g., associated with an IP address).In one embodiment, SDR agents can attach data streams to suitable workload objects and record the mapping of corresponding stream identifiers to corresponding pod IP addresses in a routing table that maps each data stream to its designated pod (and, by extension, workload object). Each workload object can further manage a configurable threshold for the maximum number of concurrent data streams that the workload object can efficiently handle. The SDR agents and workload objects are configured according to [reference to relevant documentation]. Fig. 3 discussed in more detail.

[0016] The SDR controller 240 can manage numerous aspects of the SDR agents 221, 222, and 223 within the SDR system. For example, the SDR controller 240 can manage the interactions between SDR agents and the distributed network of microservices. In another example, the SDR controller 240 can handle the initial startup of SDR agents in the cluster. The SDR controller 240 can also manage configuration data for each SDR agent, such as tracking information including connection ports for each SDR agent and the maximum capacity for each workload object associated with the agent. In one embodiment, the SDR controller 240 can dynamically adjust workload distribution based on real-time analytics and performance metrics.The SDR Controller 240 can continuously monitor the status of each SDR agent and its corresponding workload objects, such as monitoring the power density for each workload object. The SDR Controller 240 can also perform automatic scaling actions, such as provisioning additional pods when a particular service reaches its maximum capacity or when an influx of data streams is detected. The SDR Controller 240 can ensure that no single microservice is overloaded with excessive load, which could negatively impact processing efficiency and speed. For example, if there is a surge in video data from a fleet of newly deployed traffic drones, the SDR Controller 240 can analyze the current load and decide to bring up additional processing pods to keep pace with the increased demand.The SDR Controller 240 can also ensure redundancy by maintaining a pool of spare copies ready to be deployed if an active pod fails, thus providing fault tolerance. In addition to managing system resources, the SDR Controller 240 can also perform fault detection and recovery tasks. The SDR Controller 240 can actively search for and identify problems, such as pod failures or network anomalies, and immediately communicate with the Configuration Control Bus 250 to initiate corrective actions. By comparing the active stream identifiers with the initiator 210, the SDR Controller 240 can ensure that data processing capabilities are aligned with actual operational requirements.

[0017] The SDR agents 221, 222, and 223 can attach to workload objects and create or update a routing table that maintains the mapping of streams to workload objects. Upon receiving an activation message from the initiator 210, an SDR agent can identify a suitable workload object, or in some embodiments, a suitable pod, that has the capacity to handle the incoming stream. The SDR agents can manage (e.g., maintain, update, etc.) the routing entry table, mapping stream identifiers to pod addresses. The SDR agents can also facilitate communication between different workload objects by routing requests to the specific pods responsible for handling streams associated with specific stream identifiers. The SDR agents accomplish this by calling functions on the appropriate pod within the network.To facilitate this targeted routing, call functions can also include the stream identifier within the header of their requests. Functionalities relating to the SDR agents are defined according to [reference to relevant documentation]. Fig. 3 discussed in more detail.

[0018] Fig. Figure 3 illustrates an exemplary SDR agent attached to at least one workload object, according to at least one embodiment. The one in Fig. The example SDR agent shown in Figure 3 is SDR agent 221, which is associated with a workload object (i.e., workload object 231). Other SDR agents in the system can be configured similarly to SDR agent 221. SDR agent 221 can include an SDR stream distribution module 310, which distributes streams to appropriate workload objects, and an SDR request routing module 320, which routes requests to appropriate pods in the cluster. Workload object 231 can manage one or more pods, such as Pod A 330, Pod B 331, and Pod N 332. A pod, as used here, can refer to a deployable unit in container orchestration platforms (e.g., Kubernetes). A pod can contain one or more containers that share storage, networking, and a specification of how the containers are to run. Each workload object can manage one or more pods in which the actual processes or...Processes or microservices are executed. A workload object can manage the state and properties of these pods, such as the number of copies, the container images to use, and the network rules. Each workload object can be attached to an SDR agent, which manages communication between workload objects and other components in the system.

[0019] The SDR power distribution module 310 can perform the initial processing of data streams as the streams enter the system. This occurs upon receiving an activation message from an initiator (e.g., the initiator 210). Fig. 2) The SDR power distribution module 310 can identify the workload object 231 such that it has available capacity to process the incoming power and update the route entry table accordingly (e.g., with an entry specifying the power identifier and the address of a processing pod). In one embodiment, the route entry table can be implemented using a Redis database, enabling dynamic and efficient routing of power within the system.Various other message-oriented middleware (MOM) software or hardware infrastructure that support sending and receiving messages between distributed systems can also be used, such as Apache Kafka for high-throughput stream processing, RabbitMQ for a flexible message queue, ZeroMQ for low-latency asynchronous message delivery, and in-memory data grids like Hazelcast for distributed data management, providing scalable solutions for complex routing requirements. Additionally, hardware-based approaches, including programmable network switches or custom-designed FPGAs and ASICs, can provide efficient data routing capabilities.

[0020] The SDR stream distribution module 310 can then enable the transfer of a data payload to the workload object 231 and initiate processing. In one embodiment, workload objects can support RESTful (Representational State Transfer) operations for "add" and "delete" operations for dynamically modifying processing tasks. If such operations are not configured on a workload object, the SDR stream distribution module 310 can assign the stream to a pod and record this assignment in the route entry table to track a stream identifier (e.g., the stream identifier) ​​for ongoing processing activities. Additionally, to maintain service continuity and scalability, the SDR stream distribution module 310 can instantiate new workload processes (e.g., pods) when the current capacity is reached to ensure uninterrupted data stream processing capabilities of the system.After power data has been processed and results generated, the SDR power distribution module 310 can also assign a result identifier to the generated results and store this assignment in the route entry table. The generated results can be used by other microservices by retrieving the results using the result identifier. In one embodiment, the route entry table is a global database in which each SDR agent can update, query, and retrieve a related result associated with an identifier or an IP address. To ensure the integrity and consistency of data within the route entry table, particularly in scenarios involving concurrent write or read operations, the invention further employs a read / write lock mechanism for each record to prevent operational interruptions and maintain data consistency across the system.

[0021] To demonstrate the process of current distribution using an example, consider an SDR system where an initiator triggers the activation / deactivation of sensors (and their respective currents). The initiator can perform sensor addition and removal operations by sending messages (e.g., JSON messages) with a unique current identifier to a configuration control bus, which then announces these messages to the system. An SDR agent can receive such a message from the configuration control bus and, based on the message content, identify an available pod within the cloud environment that has the capacity to process the incoming data stream. Upon successful identification, the SDR agent can update a route entry table by adding a mapping of the current identifier to the IP address of the workload object.The corresponding workload object can be notified of the incoming data stream via a RESTful interface. The workload object is then configured to run the microservices and perform the necessary functions on the data stream.

[0022] The SDR Request Routing Module 320 can manage the dynamic allocation of data streams to workload objects or pods within a cloud-based infrastructure. The SDR Request Routing Module 320 can receive incoming service requests from other SDR agents and route the requests to the appropriate pods specified in the request. In one embodiment, the SDR Request Routing Module can leverage application programming interfaces (APIs) to route requests between microservices, such as Envoy XDS, which includes LDS (Listener Discovery Service), RDS (Route Discovery Service), CDS (Cluster Discovery Service), and EDS (Endpoint Discovery Service). These APIs allow Envoy XDS to dynamically receive configuration updates without requiring a restart. Each service request can be identified by a unique stream identifier encapsulated within the message header.The SDR request routing module 320 can use the stream identifier to locate a specific pod among potentially numerous pods within the Kubernetes cluster, with the specific pod being designated to process the specific data stream associated with the service request.

[0023] To illustrate the request routing process in an SDR system with an example, the service request routing operation is initiated within the SDR system when a service request is received, transmitted via HTTP (Hypertext Transfer Protocol) or RPC (Remote Procedure Calls, such as Google's gRPC). The request may include a specific stream identifier corresponding to the stream associated with the request. An SDR request routing module can consult a route entry table to identify the network address of the pod assigned to the stream identifier mentioned in the request header. The SDR request routing module can then assist in delivering a request payload to the identified pod responsible for processing the data stream.In one embodiment, the SDR request routing module can modify the routing configurations upon receiving a deactivation message from an initiator. In the scenario where a deactivation message is received, the SDR request routing module can remove the associated routing information from the route entry table. Removing unassigned entries from the route entry table can conserve system resources and optimize operational efficiency. In another embodiment, a request can be received to remove routing information associated with a data stream from the route entry table. Upon detecting that the data stream is being used by other workload objects, the SDR request routing module can determine not to delete the routing information and to retain it in the route entry table.If the SDR Request Routing Module 320 detects that the data stream is not assigned to any other workload objects, the SDR Request Routing Module 320 can delete the assigned entry from the route entry table.

[0024] Fig. Figure 4 illustrates an exemplary operational scenario with various components for processing and reproducing results based on data streams, made possible by an SDR system within a Kubernetes environment. As shown in Fig. As illustrated in Figure 4, the process or procedure can be initiated by an initiator 410, which is capable of triggering the activation / deactivation of sensors (e.g., Sensor ID1 and Sensor ID2) and their respective currents. The initiator 410 can send a message of a specific format to a configuration control bus, where the message is broadcast to SDR agents within the system. For example, the sensors could be security cameras in a facility managed by a VST, which can act as the initiator. The SDR system can include multiple SDR agents, such as SDR agents 421, 422, and 423. The SDR agents can receive the message broadcast by the configuration control bus and be configured to work with microservices.Upon receiving stream information from initiator 410, SDR agents use RESTful API calls to initiate and deploy microservice processing functions tailored to that specific stream. An SDR agent can find a workload object suitable for the stream data and record the mapping of the stream identifier to the workload object's IP address in a route entry table. For example, as shown in... Fig. Figure 4 illustrates that each SDR agent 421, 422, and 423 is assigned to a corresponding workload object, such as workload objects 431, 432, and 433, and such assignments are recorded in the route entry table. The result (e.g., output, result, etc.) of the processing tasks performed by the microservices is then transferred to the result generation or playback endpoints.

[0025] In one embodiment, each data stream can be processed by multiple microservices. For example, the result AB, ID1 440 and the result AB, ID2 441 are examples of results from combined processing. For stream ID1, generated by sensor ID1, the result (AB, ID1) is produced by the collaborative efforts of workload objects A and B. Similarly, for stream ID2, the result (AB, ID2) is produced by the collaborative operation of workload objects A and B. In one embodiment, microservices within a cluster can be managed by an SDR request routing module (e.g., the SDR request routing module 320 from [manufacturer name missing in original text]). Fig. 3) communicate with each other. For example, SDR agent C, associated with workload object C, can first interact with SDR agent B, also associated with workload object B, to obtain output from workload object B and generate result C, producing result C, ID1 450 and result C, ID2 451. For SDR agent C to capture results from SDR agent B, agent C can send gRPC requests with stream identifiers in the request headers. Upon receiving the request, SDR agent B can identify the specific process handling the specified stream identifier, such as through a routing table, and may be configured to forward the requested results to SDR agent C. After receiving the results, SDR agent C can then generate a composite result, such as result C, and forward this result to an endpoint (e.g., a server).send a client device) for playback.

[0026] Fig. Figure 5 illustrates an exemplary process or method for dynamically routing and processing stream data according to at least one embodiment. The process or method 500 can begin with an initiator 510 assigning one or more identifiers to one or more data sources in an SDR system. In one embodiment, the data source(s) can be sensors, such as cameras, that continuously stream data. The data source(s) can be made available to SDR agents within the SDR system via a control bus. The identification information, such as the identifiers for the one or more data sources, can be written to a control bus 520, which is made available to the SDR agents. One or more SDR agents can receive a request and locate a workload object for processing the data source.In one embodiment, the request can be received by the initiator as a workload object payload 530 to associate a data stream associated with a data source of one or more data sources with a workload object, wherein the workload object is available and capable of handling the specific task associated with the stream, such as performing AI inferencing like object detection or motion detection. The data source identifier can be obtained by an SDR agent and used 540 to identify a suitable workload object from a plurality of workload objects associated with the SDR agent that is available and capable of handling the specific task. In one embodiment, the SDR agent can make this mapping between the data source identifier and an identifier of the suitable workload object (e.g.,The appropriate workload object (IP address) is stored in a routing table. After the appropriate workload object has processed the stream data and generated a result, the routing table can then be updated with a second identifier, such as a result identifier, that corresponds to the result generated by the appropriate workload object (550).

[0027] Fig. Figure 6 illustrates an exemplary system environment that includes a power distribution and routing system, according to various embodiments. As an example, it illustrates... Fig. 6 An exemplary networked system 600 that can be used to provide, generate, modify, encode, process, and / or transmit data or other content. The exemplary networked system 600 can include a client device 602, another client device 603, a network 614, a third-party service 660, and a provider environment 616 that includes a power distribution and routing system 630.

[0028] The client device 602 can generate or receive data for a session using components of an application 607 on the client device 602 and data stored locally on the client device 602. For example, a user can use the client device 602 to perform stream data processing using the application 607. Although only one client device 602 is illustrated in detail, the exemplary networked system 600 can include one or more other client devices 603 that can communicate with the provider environment 616 over the network 614.The client device 602 can be any suitable data processing device capable of enabling a user to perform tasks related to stream data processing, as discussed herein, such as a desktop computer, a notebook computer, a computer workstation, a game console, a set-top box, a streaming device, a smartphone, a tablet computer, a VR headset, AR glasses, a portable computer, or a smart TV. In at least one embodiment, a user can access stream data processing results using a user interface (UI) 606 running on the client device 602, although at least one functionality can also operate on a remote device, a networked device, or via a cloud computing platform.In at least one embodiment, a user can provide input to the UI 606, for example, via a touch-sensitive display 604 or by moving a mouse cursor displayed on a screen. In one embodiment, a user can provide input such as records and data to the application 607. The application 607 can be provided by the provider environment 616 so that the user can download it to the client device 602. In at least one embodiment, the client device 602 can include at least one processor 608 (e.g., a CPU or GPU), memory 612, and storage 610 to run the application 607 and / or perform tasks on behalf of the application 607.

[0029] In one embodiment, the client device 602 can transmit a request over at least one wired or wireless network, such as the internet, Ethernet, a local area network (LAN), or a cellular network, among other such options. In this example, these requests can be transmitted to an address associated with a cloud provider, which can operate or control one or more electronic resources in a cloud provider environment, such as a data center or server farm. In at least one embodiment, the request can be received or processed by at least one edge server located at a network edge and outside of at least one security layer associated with the cloud provider environment.This reduces latency by enabling client devices to interact with servers that are located closer together, while also improving the security of resources in the cloud provider environment.

[0030] The network 614 can represent the communication paths between the client device 602, the provider environment 616, the other client device 603, and the third-party service 660. The client device 602 can send input information related to stream data processing over the network 614. This information can be received by a remote computing system, which may be part of a resource provider environment 616. In one embodiment, the network 614 is the Internet. The network 614 can include any suitable network, including an intranet, the Internet, a cellular network, a local area network (LAN), or any other such network or combination thereof, and communication over a network can be enabled via wired and / or wireless connections.Network 614 can also utilize dedicated or private communication links that are not necessarily part of the Internet. In one embodiment, Network 614 uses standard communication technologies and / or protocols. Thus, Network 614 can include connections that use technologies such as Ethernet, Wi-Fi, Integrated Services Digital Network (ISDN), Digital Subscriber Lines (DSL), Asynchronous Transfer Mode (ATM), etc. Likewise, the network protocols used in Network 614 can include Multiprotocol Label Switching (MPLS), Transmission Control Protocol / Internet Protocol (TCP / IP), Hypertext Transport Protocol (HTTP), Simple Mail Transfer Protocol (SMTP), File Transfer Protocol (FTP), etc. In one embodiment, at least some of the connections use mobile network technologies, such as Long Term Evolution (LTE).The data exchanged over the 614 network can be represented using various technologies or formats, including Hypertext Markup Language (XML), Wireless Access Protocol (WAP), Short Message Service (SMS), and others. Additionally, all or some of the connections can be encrypted using conventional encryption technologies, such as Secure Sockets Layer (SSL), Secure HTTP, or Virtual Private Networks (VPNs). Alternatively, the 602 client device can use customer-specific and / or dedicated data communication technologies instead of, or in addition to, those described above.

[0031] The provider environment 616 can include any suitable components for receiving requests and sending back information or performing actions in response to those requests. In the Fig. In the embodiment illustrated in Figure 6, the provider environment 616 can include an interface 618 and a server 620, which contains various components for performing tasks associated with stream data processing. In at least one embodiment, the provider environment 616 can include web servers and / or application servers for receiving and processing requests and then sending back data or other content or information in response to a request.

[0032] Interface 618 can receive communications to server 620. In at least one embodiment, interface 618 can include application programming interfaces (APIs) or other exposed interfaces that allow a user to submit requests to server 620. In at least one embodiment, interface 618 can also include other components, such as at least one web server, routing components, or load balancers. In at least one embodiment, components of interface 618 can determine the type of request or communication and route a request to a suitable system or service, such as the power distribution and routing system 630.

[0033] The server 620 can include a transfer manager 622, a content application 624, an object repository 634, and a user database 636. The server 620 can receive requests and data from the client device 602, perform tasks associated with the requests, and send results or other data to the client device 602. In at least one embodiment, the content application 624, running on the server 620 (for example, a cloud server or edge server), can initiate a session associated with the client device 602, using a session manager and user data stored in a user database 636, and can cause content, such as one or more object representations, to be selected from an object repository 634 for processing by a content manager 626.At least one portion of the generated content, such as results from stream data processing, can be transferred to the client device 602 using a suitable transfer manager 622 for transmission by downloading, streaming, or another such transmission channel. An encoder can be used to encode and / or compress at least some of this data before it is transferred to the client device 602. In at least one embodiment, the client device 602 receiving such content can make this content available to a corresponding application (e.g., the application 607) for selecting, providing, synthesizing, modifying, or using the content for presentation (or other purposes) on or through the client device 602.A decoder can also be used to decode data received over the network 614 for presentation via the client device 602, such as image or video content displayed on a touchscreen 604. In at least one embodiment, at least part of the content can already be stored on, played back on, or accessible to the client device 602, so that transmission over the network 614 is not necessary for at least this part of the content, for example, if the content has been previously downloaded or stored locally on a hard disk or optical disk. In at least one embodiment, a transmission mechanism such as data streaming can be used to transfer the content from the server 620 or the user database 636 to the client device 602.In at least one embodiment, at least one section of this content can be obtained, enhanced, and / or streamed from another source, such as a third-party service 660 or another client device 603, which may also include a content application 662 for generating, enhancing, or providing content. In at least one embodiment, sections of this functionality can be performed using multiple data processing devices or multiple processors within one or more data processing devices, which may include a combination of CPUs and GPUs.

[0034] In at least one embodiment, the Server 620 can include a processor such as a central processing unit (CPU). However, in at least one embodiment, resources in such environments can utilize GPUs to process data for at least certain types of requirements. In at least one embodiment, GPUs with thousands of cores are designed to handle numerous parallel workloads and have therefore become popular in deep learning for training neural networks and generating predictions.While using GPUs for offline builds enables faster training of larger and more complex models in at least one embodiment, offline prediction generation implies that either request-time input features cannot be used or predictions for all permutations of features must be generated and stored in an offline lookup table to serve real-time requirements. If, in at least one embodiment, a deep learning framework supports a CPU mode and a model is small and simple enough to perform feed-forward on a CPU with reasonable latency, a service on a CPU instance can host a model. In at least one embodiment, training can be performed on a GPU and real-time inference on a CPU.If, in at least one embodiment, a CPU-based approach is not a practical option, a service can be run on a GPU instance. However, since GPUs have different performance and cost characteristics than CPUs, in at least one embodiment, running a service that offloads a runtime algorithm to a GPU may require it to be designed differently than a CPU-based service.

[0035] The server 620 can include the content application 624, which contains the content manager 626 and the power distribution and routing system 630. As discussed earlier, the content manager 626 can send objects, such as records and instructions, from the object repository 634, along with requests and other data, from the client device 602 to the power distribution and routing system 630 for power data processing. The power distribution and routing system 630 can process input data and provide the results to the transfer manager 622 for return to the client device 602. The power distribution and routing system 630 can also use local records or records provided by the third-party service 660 for power data processing. DATA CENTER

[0036] Fig. Figure 7 illustrates an exemplary data center 700, in which at least one embodiment can be used. In at least one embodiment, the data center 700 includes a data center infrastructure layer 710, a framework layer 720, a software layer 730, and an application layer 740.

[0037] In at least one embodiment, as in Fig. As shown in Figure 7, the data center infrastructure layer 710 can include a resource orchestrator 712, clustered computing resources 714 and node computing resources (“node CR”) 716(1) to 716(N), where “N” is any positive integer (which may be different from other integers “N” used in other figures). In at least one embodiment, the node CR 716(1)-716(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (“FPGAs”), graphics processors, etc.), memory devices 718(1)-718(N) (e.g., dynamic read-only memory), data storage devices (e.g., solid-state or hard disk drives), network input / output devices (“NW-I / O” devices), network switches, virtual machines (“VMs”), power modules, and cooling modules.In at least one embodiment, one or more Node-CRs of Node-CR 716(1)-716(N) can be a server comprising one or more of the aforementioned computing resources.

[0038] In at least one embodiment, the grouped compute resources 714 can comprise separate groupings of node RRs housed in one or more racks (not shown), or many racks housed in data centers at different geographic locations (also not shown). In at least one embodiment, separate groupings of node RRs within grouped compute resources 714 can comprise grouped compute, network, storage, or memory resources that can be configured or allocated to support one or more workloads. In at least one embodiment, multiple node RRs containing CPUs or processors can be grouped in one or more racks to provide compute resources to support one or more workloads.In at least one embodiment, one or more racks can also include any number of power modules, cooling modules and network switches in any combination.

[0039] In at least one embodiment, the resource orchestrator 712 can configure or otherwise control one or more node CR 716(1) to 716(N) and / or grouped computing resources 714. In at least one embodiment, the resource orchestrator 712 can include a management unit of a software design infrastructure (“SDI”) for the data center 700. In at least one embodiment, the resource orchestrator 712 can include hardware, software, or a combination thereof.

[0040] In at least one embodiment, as in Fig. As shown in Figure 7, the framework layer 720 includes a task scheduler 722, a configuration manager 724, a resource manager 726, and a distributed file system 728. In at least one embodiment, the framework layer 720 can include a framework to support software 732 of software layer 730 and / or one or more applications 742 of application layer 740. In at least one embodiment, the software 732 or the application(s) 742 can each be web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 720 can, without restriction, be a type of web application framework for free and open-source software, such as Apache Spark™ (hereinafter "Spark"), which can use the distributed file system 728 for large-scale data processing (e.g., "Big Data").In at least one embodiment, the task scheduler 722 can include a Spark driver to facilitate the scheduling of workloads supported by various layers of the data center 700. In at least one embodiment, the configuration manager 724 can be able to configure various layers, such as the software layer 730 and the framework layer 720, including Spark and the distributed file system 728, to support large-scale data processing. In at least one embodiment, the resource manager 726 can be able to manage clustered or grouped compute resources that are allocated or assigned to support the distributed file system 728 and the task scheduler 722. In at least one embodiment, the clustered or grouped compute resources can include a grouped compute resource 714 on the data center infrastructure layer 710.In at least one embodiment, the resource manager 726 can coordinate with the resource orchestrator 712 to manage these allocated or assigned computing resources.

[0041] In at least one embodiment, the software 732, which is included in software layer 730, may include software that is used by at least parts of node CR 716(1) to 716(N), grouped computing resources 714, and / or the distributed file system 728 of framework layer 720. In at least one embodiment, one or more types of software may include, but are not limited to, internet web browsing software, email virus scanning software, database software, and streaming video content software.

[0042] In at least one embodiment, the application(s) 742 included in the application layer 740 may include one or more types of applications used by at least parts of node CR 716(1) to 716(N), grouped compute resources 714, and / or the distributed file system 728 of the framework layer 720. In at least one embodiment, one or more types of applications may include any number of a genomics application, a cognitive computing application, and a machine learning application, including, but not limited to, training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0043] In at least one embodiment, the configuration manager 724, the resource manager 726, and the resource orchestrator 712 can implement any number and type of self-modifying actions based on any amount and type of data acquired in any technically feasible manner. In at least one embodiment, self-modifying actions can relieve a data center operator of the data center 700 from potentially making poor configuration decisions and avoid potentially underutilized and / or underperforming sections of a data center.

[0044] In at least one embodiment, the Data Center 700 may include tools, services, software, or other resources for training one or more machine learning models, or for predicting or deriving information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weighting parameters according to a neural network architecture using software and computing resources previously described with reference to the Data Center 700.In at least one embodiment, trained machine learning models corresponding to one or more neural networks can be used to infer or predict information using resources previously described in relation to the Computing Center 700 and weighting parameters calculated using one or more training techniques described in this document.

[0045] In at least one embodiment, the data center can use CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, or other hardware to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above can be configured as a service to allow users to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.

[0046] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, the inference and / or training logic 715 can be implemented in the system of Fig. 7. can be used to perform inferencing or prediction operations that are based at least partially on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0047] The embodiments presented here can enable a linear controller with one or more features for improving the PSRR in order to identify and correct voltage noise within a circuit. COMPUTER SYSTEMS

[0048] Fig. Figure 8 is a block diagram illustrating an exemplary computer system, which may be a system with interconnected devices and components, a system-on-a-chip (SoC), or a certain combination thereof, and is formed with a processor that may include execution units for executing an instruction, according to at least one embodiment. In at least one embodiment, the computer system may, without limitation, include a component, such as a processor, to employ execution units that include logic for performing algorithms on process data, according to the present disclosure, as in the embodiment described in this document.In at least one embodiment, the Computer System 800 may include processors such as the PENTIUM® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core™ or Intel® Nervana™ microprocessors available from Intel Corporation in Santa Clara, California, although other systems (including PCs with other microprocessors, technical workstations, set-top boxes, and the like) may also be used. In at least one embodiment, the Computer System 800 may run a version of the WINDOWS operating system available from Microsoft Corporation in Redmond, Washington, although other operating systems (for example, UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0049] The embodiments can be used on other devices, such as handheld devices and embedded applications. Some examples of handheld devices are mobile phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and portable PCs. In at least one embodiment, embedded applications can include a microcontroller, a digital signal processor (DSP), a system-on-a-chip, network computers (NetPCs), set-top boxes, network hubs, wide-area network switches (WANs), or any other system capable of executing one or more instructions according to at least one embodiment.

[0050] In at least one embodiment, the computer system 800 can, without limitation, include a processor 802, which can, without limitation, include one or more execution units 808 for training and / or inferring a machine learning model according to the techniques described in this document. In at least one embodiment, the computer system 800 is a single-processor desktop or server system, but in another embodiment, the computer system 800 can be a multi-processor system.In at least one embodiment, the processor 802 can, without restriction, include a complex instruction set computing (“CISC”) microprocessor, a reduced instruction set computing (“RISC”) microprocessor, a very long instruction word (“VLIW”) microprocessor, a processor implementing a combination of instruction sets, or any other processing device, such as a digital signal processor. In at least one embodiment, the processor 802 can be coupled to a processor bus 810, which can transmit data signals between the processor 802 and other components in the computer system 800.

[0051] In at least one embodiment, the processor 802 can, without limitation, include a Level 1 ("L1") internal cache memory ("cache") 804. In at least one embodiment, the processor 802 can have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory can be located outside the processor 802. Other embodiments, depending on the specific implementation and requirements, can also include a combination of both internal and external caches. In at least one embodiment, the register file 806 can store various types of data in various registers, including, without limitation, integer registers, floating-point registers, status registers, and instruction pointer registers.

[0052] In at least one embodiment, the execution unit 808, which includes without limitation logic for performing integer and floating-point operations, is also located in the processor 802. In at least one embodiment, the processor 802 may also include a microcode ("uCode") read-only memory (ROM) that stores microcode for certain macro instructions. In at least one embodiment, the execution unit(s) 808 may include logic for handling a packed instruction set 809. In at least one embodiment, by including a packed instruction set 809 in an instruction set of a general-purpose processor 802, together with associated circuitry for executing instructions, operations used by numerous multimedia applications can be performed using compressed data in a processor 802.In at least one embodiment, numerous multimedia applications can be accelerated and executed more efficiently by using the entire width of a processor's data bus to perform operations on compressed data, thereby perhaps eliminating the need to transfer smaller data units across the processor's data bus to perform one or more operations on only one data element at a time.

[0053] In at least one embodiment, the execution unit 808 can also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, the computer system 800 can include a memory 820 without restriction. In at least one embodiment, the memory 820 can be a dynamic random access memory (“DRAM”) device, a static random access memory (“SRAM”) device, a flash memory device, or another type of memory device. In at least one embodiment, the memory 820 can store instructions 819 and / or data 821 represented by data signals that can be executed by the processor 802.

[0054] In at least one embodiment, a system logic chip can be coupled to the processor bus 810 and the memory 820. In at least one embodiment, the system logic chip can, without restriction, include a memory controller hub (“MCH”) 816, and the processor 802 can communicate with the MCH 816 via the processor bus 810. In at least one embodiment, the MCH 816 can provide a high-bandwidth memory path 818 for the memory 820 for storing instructions and data, and for storing graphics instructions, data, and textures. In at least one embodiment, the MCH 816 can route data signals between the processor 802, the memory 820, and other components in the computer system 800, and bridge data signals between the processor bus 810, the memory 820, and a system I / O interface 822.In at least one embodiment, the system logic chip can provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 816 can be coupled to the memory 820 via a high-bandwidth memory path 818, and a graphics / video card 812 can be coupled to the MCH 816 via an Accelerated Graphics Port (“AGP”) interconnect 814.

[0055] In at least one embodiment, the computer system 800 can use the system I / O interface 822, which is a proprietary node interface bus, to couple the MCH 816 to the I / O control hub (“ICH”) 830. In at least one embodiment, the ICH 830 can provide direct connections to some I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus can, without restriction, include a high-speed I / O bus for connecting peripheral devices to the memory 820, the chipset, and the processor 802. Examples can include, without limitation, an audio controller 829, a firmware hub (“Flash BIOS”) 828, a wireless transceiver 826, a data storage device 824, a legacy I / O controller 823 containing user input and keyboard interfaces 825, a serial expansion port 827, such as a universal serial bus (“USB”) port, and a network controller 834.In at least one embodiment, the data storage device 824 may comprise a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device or another mass storage device.

[0056] Illustrated in at least one embodiment Fig. 8 a system that includes interconnected hardware devices or “chips”, whereas in other embodiments Fig. 8 illustrates an exemplary system on a chip (“SoC”). In at least one embodiment, the Fig. The devices shown in Figure 8 are interconnected using proprietary interconnects, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of the Computer System 800 are interconnected using Compute Express Link (CXL) interconnects.

[0057] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, the inference and / or training logic 715 can be implemented in the system of Fig. 8 can be used to perform inferencing or prediction operations that are based at least partially on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0058] The embodiments presented herein can enable a linear controller with one or more features for improving the PSRR in order to identify and correct voltage noise within a circuit.

[0059] Fig. Figure 9 is a block diagram illustrating an electronic device 900 for using a processor 910 according to at least one embodiment. In at least one embodiment, the electronic device 900 can be, for example, and without limitation, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop computer, a tablet, a mobile device, a telephone, an embedded computer, or any other suitable electronic device.

[0060] In at least one embodiment, the electronic device 900 can, without limitation, include the processor 910, which is communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, the processor 910 is coupled using a bus or interface, such as a 1°C bus, a system management bus (“SMBus”), a low-pin-count (LPC) bus, a serial peripheral interface (“SPI”), a high-definition audio (“HDA”) bus, a serial-advanced technology attachment (“SATA”) bus, a universal serial bus (“USB”) (version 1, 2, 3), or a universal asynchronous receiver / transmitter (“UART”) bus. In at least one embodiment, the following is illustrated: Fig. 9 a system that includes interconnected hardware devices or “chips”, whereas in other embodiments Fig. 9 can illustrate an exemplary system on a chip (“SoC”). In at least one embodiment, the Fig. The devices illustrated in Figure 9 are interconnected using proprietary interconnects, standardized interconnects (e.g., PCIe), or a certain combination thereof. In at least one embodiment, one or more components are made of Fig. 9 interconnected using Compute Express Link (CXL) interconnections.

[0061] In at least one embodiment, Fig. 9 a display 924, a touchscreen 925, a touchpad 930, a near field communications unit (NFC) 945, a sensor hub 940, a thermal sensor 946, an Express chipset (EC) 935, a trusted platform module (TPM) 938, a BIOS / firmware / flash memory (BIOS, FW flash) 922, a DSP 960, a drive 920, such as a solid state drive (SSD) or a hard disk drive (HDD), a wireless local area network (WLAN) 950, a Bluetooth unit 952, a wireless wide area network (WWAN) 956, a global positioning system (GPS) 955, a camera (“USB 3.0 camera”) 954, such as a USB 3.0 camera and / or a low-power double data rate (“LPDDR”) storage unit (“LPDDR3”) 915, which is implemented, for example, in an LPDDR3 standard.These components can each be implemented in any suitable way.

[0062] In at least one embodiment, other components can be communicatively coupled to the processor 910 via components described herein. In at least one embodiment, an accelerometer 941, an ambient light sensor (“ALS”) 942, a compass 943, and a gyroscope 944 can be communicatively coupled to the sensor hub 940. In at least one embodiment, the thermal sensor 939, a fan 937, a keyboard 936, and a touchpad 930 can be communicatively coupled to the EC 935. In at least one embodiment, one or more loudspeakers 963, headphones 964, and a microphone (“Mic”) 965 can be communicatively coupled to an audio unit (“Class D audio codec and amplifier”) 962, which in turn can be communicatively coupled to the DSP 960. In at least one embodiment, the audio unit 962 can, for example and without limitation, include an audio encoder / decoder (“codec”) and a class-D amplifier.In at least one embodiment, a SIM card (“SIM”) 957 can be communicatively coupled with the WWAN unit 956. In at least one embodiment, components such as the WLAN unit 950 and the Bluetooth unit 952, as well as the WWAN unit 956, can be implemented in a next-generation form factor (“NGFF”).

[0063] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, the inference and / or training logic 715 can be implemented in the system of Fig. 9 can be used to perform inferencing or prediction operations that are based at least partially on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0064] The embodiments presented herein can enable a linear controller with one or more features for improving the PSRR in order to identify and correct voltage noise within a circuit.

[0065] Fig. Figure 10 shows a computer system 1000 according to at least one embodiment. In at least one embodiment, the computer system 1000 is configured to implement various processes and procedures described in this disclosure.

[0066] In at least one embodiment, the computer system 1000 comprises, without limitation, at least one central processing unit (“CPU”) 1002 connected to a communication bus 1010 implemented using any suitable protocol, such as PCI (“Peripheral Component Interconnect”), Peripheral Component Interconnect Express (“PCI Express”), AGP (“Accelerated Graphics Port”), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, the computer system 1000 comprises, without limitation, main memory 1004 and control logic (e.g., implemented as hardware, software, or a combination thereof), and data is stored in the main memory 1004, which may take the form of random-access memory (“RAM”).In at least one embodiment, a network interface subsystem (“network interface”) 1022 provides an interface to other computing devices and networks to receive data from and transmit data to other systems with the computer system 1000.

[0067] In at least one embodiment, the computer system 1000 comprises, without limitation, input devices 1008, a parallel processing system 1012, and display devices 1006, which may be implemented using a conventional cathode ray tube (“CRT”), a liquid crystal display (“LCD”), a light-emitting diode display (“LED”), or another suitable display technology. In at least one embodiment, user input is received from input devices 1008 such as a keyboard, mouse, touchpad, microphone, etc. In at least one embodiment, each module described herein may be arranged on a single semiconductor platform to form a processing system.

[0068] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, the inference and / or training logic 715 can be implemented in the system of Fig. 10. can be used to perform inferencing or prediction operations that are based at least partially on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0069] The embodiments presented herein can enable a linear controller with one or more features for improving the PSRR in order to identify and correct voltage noise within a circuit.

[0070] Fig. Figure 11 shows a computer system 1100 according to at least one embodiment. In at least one embodiment, the computer system 1100 comprises, without limitation, a computer 1110 and a USB stick 1120. In at least one embodiment, the computer 1110 can have, without limitation, any number and any type of processors (not shown) and memory (not shown). In at least one embodiment, the computer 1110 comprises, without limitation, a server, a cloud instance, a laptop, and a desktop computer.

[0071] In at least one embodiment, the USB stick 1120 comprises, without limitation, a processing unit 1130, a USB interface 1140, and USB interface logic 1150. In at least one embodiment, the processing unit 1130 can be any instruction execution system, device, or assembly capable of executing instructions. In at least one embodiment, the processing unit 1130 can comprise, without limitation, any number and any type of processing cores (not shown). In at least one embodiment, the processing unit 1130 comprises an application-specific integrated circuit (“ASIC”) optimized for performing any number and any type of machine learning-related operations.For example, in at least one embodiment, the processing unit 1130 is a tensor processing unit (“TPC”) optimized for performing machine learning inference operations. In at least one embodiment, the processing unit 1130 is an image processing unit (“VPU”) optimized for performing machine image processing and machine learning inference operations.

[0072] In at least one embodiment, the USB interface 1140 can be any type of USB connector or USB socket. For example, in at least one embodiment, the USB interface 1140 is a USB 3.0 Type-C socket for data and power. In at least one embodiment, the USB interface 1140 is a USB 3.0 Type-A connector. In at least one embodiment, the USB interface logic 1150 can comprise any number and any type of logic that enables the processing unit 1130 to communicate with devices (e.g., the computer 1110) via the USB port 1140.

[0073] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, the inference and / or training logic 715 can be implemented in the system of Fig. 11. can be used to perform inferencing or prediction operations that are based at least partially on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0074] The embodiments presented herein can enable a linear controller with one or more features for improving the PSRR in order to identify and correct voltage noise within a circuit.

[0075] Fig. Figure 12 shows exemplary integrated circuits and associated graphics processors that may be manufactured using one or more IP cores according to various embodiments described herein. In addition to what is shown, at least one embodiment may include other logic and circuitry, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0076] Fig. Figure 12 is a block diagram representing an exemplary system-on-a-chip (SOC) integrated circuit 1200, which can be manufactured according to at least one embodiment using one or more IP cores. In at least one embodiment, the SOC integrated circuit 1200 includes one or more application processors 1205 (e.g., CPUs), at least one graphics processor 1210, and may additionally include an image processor 1215 and / or a video processor 1220, each of which may be a modular IP core. In at least one embodiment, the integrated circuit 1200 of the SOC includes peripheral or bus logic comprising a USB controller 1225, a UART controller 1230, an SPI / SDIO controller 1235, and an I 2 2S / I 2 The 2C controller 1240 is included. In at least one embodiment, the integrated circuit 1200 of the SOC can include a display device 1245 coupled to one or more high-resolution multimedia interfaces (HDMI) 1250 and mobile industrial processor interfaces (MIPI) 1255. In at least one embodiment, the memory can be provided by a flash memory subsystem 1260 comprising flash memory and a flash memory controller. In at least one embodiment, a memory interface can be provided via a memory controller 1265 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits additionally include an embedded security engine 1270.

[0077] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, the inference and / or training logic 715 can be used in the integrated circuit 1200 of the SOC to perform inference or prediction operations that are based at least partially on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0078] The embodiments presented herein can enable a linear controller with one or more features for improving the PSRR in order to identify and correct voltage noise within a circuit.

[0079] The Fig. Figures 13A-13B illustrate exemplary integrated circuits and associated graphics processors that may be manufactured using one or more IP cores according to various embodiments described herein. In addition to what is shown, at least one embodiment may include other logic and circuitry, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.

[0080] The Fig. Figures 13A-13B are block diagrams illustrating exemplary graphics processors for use in a SoC according to the embodiments described herein. Fig. Figure 13A illustrates an exemplary graphics processor 1310 of an integrated circuit of a system on a chip, which may be manufactured according to at least one embodiment using one or more IP cores. Fig. Figure 13B shows an additional exemplary graphics processor 1340 of an integrated system-on-chip circuit, which, according to at least one embodiment, can be manufactured using one or more IP cores. In at least one embodiment, the graphics processor 1310 is made of Fig. 13A is a low-performance graphics processor core. In at least one embodiment, the graphics processor 1340 is in Fig. 13B a ​​higher-performance graphics processor core. In at least one embodiment, each of the graphics processors 1310, 1340 variants of the computer system 1100 can be made from Fig. 11.

[0081] In at least one embodiment, the graphics processor 1310 comprises a vertex processor 1305 and one or more fragment processors 1315A-1315N (e.g., 1315A, 1315B, 1315C, 1315D to 1315N-1 and 1315N). In at least one embodiment, the graphics processor 1310 can execute different shader programs via separate logic, such that the vertex processor 1305 is optimized for executing operations for vertex shader programs, while one or more fragment processors 1315A-1315N perform fragment (e.g., pixel) shading operations for fragment or pixel shader programs. In at least one embodiment, the vertex processor 1305 performs a vertex processing stage of a 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the fragment processors 1315A-1315N use the primitives and vertex data generated by the vertex processor 1305 to create an image buffer that is displayed on a display device.In at least one embodiment, the 1315A-1315N fragment processors are optimized to execute fragment shader programs as provided in an OpenGL API, which can be used to perform operations similar to a pixel shader program as provided in a Direct 3D API.

[0082] In at least one embodiment, the graphics processor 1310 additionally comprises one or more memory management units (MMUs) 1320A-1320B, cache(s) 1325A-1325B, and circuit interconnects 1330A-1330B. In at least one embodiment, one or more MMUs 1320A-1320B provide a mapping of virtual to physical addresses for the graphics processor 1310, including for the vertex processor 1305 and / or fragment processors 1315A-1315N, which, in addition to the vertex or image / texture data stored in one or more caches 1325A-1325B, can reference vertex or image / texture data stored in memory. In at least one embodiment, one or more MMUs 1320A-1320B can be synchronized with other MMUs within a system, including one or more MMUs that are connected to one or more application processors 1305, image processors 1315 and / or video processors 1320. Fig. 13A are assigned so that each processor 1305-1320 can participate in a common or unified virtual memory system. In at least one embodiment, one or more circuit connections 1330A-1330B enable the graphics processor 1310 to connect to other IP cores within the SoC either via an internal bus of the SoC or via a direct connection.

[0083] In at least one embodiment, the 1340 graphics processor has one or more shader cores 1355A-1355N (e.g., 1355A, 1355B, 1355C, 1355D, 1355E, 1355F, up to 1355N-1 and 1355N), as described in Fig. Figure 13B shows a unified shader core architecture in which a single core or type can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 1340 includes an inter-core task manager 1345, which acts as a thread dispatcher to distribute execution threads to one or more shader cores 1355A-1355N, and a tiling unit 1358 to accelerate tiling operations for tile-based rendering, in which rendering operations for a scene are distributed across image space to, for example, exploit local spatial coherence within a scene or optimize the use of internal caches.

[0084] The embodiments presented herein can enable a linear controller with one or more features for improving the PSRR in order to identify and correct voltage noise within a circuit.

[0085] Fig. Figure 14 is a block diagram representing a computing system 1400 according to at least one embodiment. In at least one embodiment, the computing system 1400 comprises a processing subsystem 1401 with one or more processors 1402 and a system memory 1404, which communicate via a connection path that may include a memory hub 1405. In at least one embodiment, the memory hub 1405 may be a separate component within a chipset component or be integrated into one or more processors 1402. In at least one embodiment, the memory hub 1405 is coupled to an I / O subsystem 1411 via a communication link 1406. In at least one embodiment, the I / O subsystem 1411 includes an I / O hub 1407, which may enable the computing system 1400 to receive input from one or more input devices 1408.In at least one embodiment, the I / O hub 1407 enables a display controller, which may be contained in one or more processors 1402, to provide outputs to one or more display devices 1410A. In at least one embodiment, one or more display devices 1410A coupled to the I / O hub 1407 may comprise a local, internal, or embedded display device.

[0086] In at least one embodiment, the processing subsystem 1401 comprises one or more parallel processors 1412 coupled to the memory hub 1405 via a bus or other communication link 1413. In at least one embodiment, the communication link 1413 can use any number of standards-based communication link technologies or protocols, such as, but not limited to, PCI Express, or be a vendor-specific communication interface or communication structure. In at least one embodiment, one or more parallel processors 1412 form a computationally focused parallel processing or vector processing system, which can have a large number of processor cores and / or processor clusters, such as a processor with many integrated cores (MIC).In at least one embodiment, some or all of the parallel processors 1412 form a graphics processing subsystem that can output pixels to one or more display devices 1410A coupled via the I / O hub 1407. In at least one embodiment, the parallel processors 1412 can also include a display controller and a display interface (not shown) to enable a direct connection to one or more display devices 1410B. In at least one embodiment, the parallel processors 1412 have one or more cores, such as the graphics cores 1400 discussed herein.

[0087] In at least one embodiment, a system storage unit 1414 can be connected to the I / O hub 1407 to provide a storage module for the computing system 1400. In at least one embodiment, an I / O switch 1416 can be used to provide an interface mechanism that enables connections between the I / O hub 1407 and other components, such as a network adapter 1418 and / or a wireless network adapter 1419, which can be integrated into the platform and various other devices that can be added via one or more auxiliary devices 1420. In at least one embodiment, the network adapter 1418 can be an Ethernet adapter or another wired network adapter.In at least one embodiment, the wireless network adapter 1419 may include one or more of the following components: Wi-Fi, Bluetooth, Near Field Communication (NFC) or another network device that includes one or more wireless radios.

[0088] In at least one embodiment, the computing system 1400 may include other components not explicitly shown, including USB or other connectors, optical storage drives, video recording devices, and the like, which may also be connected to the I / O hub 1407. In at least one embodiment, communication paths connecting various components in Fig. 14 connect to each other, implemented using any suitable protocols, such as PCI (Peripheral Component Interconnect)-based protocols (e.g. PCI-Express) or other bus or point-to-point communication interfaces and / or protocols, such as NV-Link high-speed links or interconnection protocols.

[0089] In at least one embodiment, the parallel processors 1412 include circuits optimized for graphics and video processing, including, for example, a video output circuit, and form a graphics processing unit (GPU), wherein, for example, the parallel processors 1412 include a graphics processing core 1400. In at least one embodiment, the parallel processors 1412 include circuits optimized for general-purpose processing. In at least one embodiment, components of the computing system 1400 can be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, the parallel processor(s) 1412, the memory hub 1405, the processor(s) 1402, and the I / O hub 1407 can be integrated into an integrated circuit (SoC).In at least one embodiment, components of the 1400 computing system can be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least one section of components of the 1400 computing system can be integrated into a multi-chip module (MCM) that can be connected to other multi-chip modules to form a modular computing system.

[0090] The inference and / or training logic 715 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, the inference and / or training logic 715 can be implemented in the system of Fig. 14 can be used to perform inferencing or prediction operations that are based at least partially on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0091] The embodiments presented herein can enable a linear controller with one or more features for improving the PSRR in order to identify and correct voltage noise within a circuit. PROCESSORS

[0092] Fig. Figure 15A shows a parallel processor 1500 according to at least one embodiment. In at least one embodiment, various components of the parallel processor 1500 can be implemented using one or more integrated circuit devices, such as programmable processors, application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). In at least one embodiment, the parallel processor 1500 shown is a variant of one or more parallel processors 1412, which are implemented in Fig. Figure 14 shows an exemplary embodiment. In at least one embodiment, a parallel processor 1500 has one or more graphics cores 1400.

[0093] In at least one embodiment, the parallel processor 1500 comprises a parallel processing unit 1502. In at least one embodiment, the parallel processing unit 1502 comprises an I / O unit 1504, which enables communication with other devices, including other instances of the parallel processing unit 1502. In at least one embodiment, the I / O unit 1504 can be directly connected to other devices. In at least one embodiment, the I / O unit 1504 connects to other devices using a hub or switch interface, such as a memory hub 1505. In at least one embodiment, connections between the memory hub 1505 and the I / O unit 1504 form a communication link 1513.In at least one embodiment, the I / O unit 1504 is connected to a host interface 1506 and a memory crossbar 1516, wherein the host interface 1506 receives commands to perform processing operations and the memory crossbar 1516 receives commands to perform memory operations.

[0094] In at least one embodiment, when the host interface 1506 receives an instruction buffer via the I / O unit 1504, it can forward work operations to a frontend 1508 for the execution of these instructions. In at least one embodiment, the frontend 1508 is coupled to a scheduler 1510 (which can be referred to as a sequencer) configured to distribute instructions or other work items to a processing cluster array 1512. In at least one embodiment, the scheduler 1510 ensures that the processing cluster array 1512 is properly configured and in a valid state before tasks are distributed to a cluster of the processing cluster array 1512. In at least one embodiment, the scheduler 1510 is implemented via firmware logic running on a microcontroller.In at least one embodiment, the scheduler 1510, implemented on a microcontroller, is configurable to perform complex scheduling and workload distribution operations with coarse and fine granularity, thereby enabling fast preemption and context switching of threads running on the processing array 1512. In at least one embodiment, the host software can detect workloads for scheduling on the processing cluster array 1512 via one of several graphics processing paths. In at least one embodiment, workloads can then be automatically distributed across the processing array cluster 1512 by the logic of the scheduler 1510 within a microcontroller that contains the scheduler 1510.

[0095] In at least one embodiment, the processing cluster array 1512 can comprise up to "N" processing clusters (e.g., cluster 1514A, cluster 1514B to cluster 1514N), where "N" is a positive integer (which may be a different integer "N" than in other figures). In at least one embodiment, each cluster 1514A-1514N of the processing cluster array 1512 can execute a large number of concurrent threads. In at least one embodiment, the scheduler 1510 can allocate work to the clusters 1514A-1514N of the processing cluster array 1512 by using different scheduling and / or workload distribution algorithms, which can vary depending on the workload encountered for each type of program or computation. In at least one embodiment, the scheduling orScheduling can be handled dynamically by the scheduler 1510 or partially supported by compiler logic during the compilation of the program logic, which is configured for execution by the processing cluster array 1512. In at least one embodiment, different clusters 1514A-1514N of the processing cluster array 1512 can be assigned for processing different types of programs or for executing different types of calculations.

[0096] In at least one embodiment, the processing cluster array 1512 can be configured to perform various types of parallel processing operations. In at least one embodiment, the processing cluster array 1512 is configured to perform general parallel computing operations. For example, in at least one embodiment, the processing cluster array 1512 can include logic for performing processing tasks that involve filtering video and / or audio data, performing modeling operations, including physical operations, and performing data transformations.

[0097] In at least one embodiment, the processing cluster array 1512 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing cluster array 1512 may include additional logic to support the execution of such graphics processing operations, including, but not limited to, texture sampling logic for performing texture operations, as well as tessellation logic and other vertex processing logic. In at least one embodiment, the processing cluster array 1512 may be configured to execute graphics processing-related shader programs, such as, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 1502 may transfer data from system memory for processing via the I / O unit 1504.In at least one embodiment, data transmitted during processing can be stored in an on-chip memory (e.g., a parallel processor memory 1522) during processing and subsequently written back to the system memory.

[0098] In at least one embodiment, when the parallel processing unit 1502 is used to perform graphics processing, the scheduler 1510 can be configured to divide a processing workload into approximately equal tasks to allow for better distribution of graphics processing operations across multiple clusters 1514A-1514N of the processing cluster array 1512. In at least one embodiment, sections of the processing cluster array 1512 can be configured to perform different types of processing.For example, in at least one embodiment, a first section can be configured to perform vertex shading and topology generation, a second section can be configured to perform tessellation and geometry shading, and a third section can be configured to perform pixel shading or other screenspace operations to generate a rendered image for display. In at least one embodiment, intermediate data generated by one or more of the clusters 1514A-1514N can be stored in buffers to allow the transfer of intermediate data between the clusters 1514A-1514N for further processing.

[0099] In at least one embodiment, the processing cluster array 1512 can receive processing tasks to be processed via the scheduler 1510, wherein the scheduler receives instructions from the front-end unit 1508 that define the processing tasks. In at least one embodiment, processing tasks can include indices of data to be processed, e.g., surface data (patch data), primitive data, vertex data, and / or pixel data, as well as state parameters and instructions that define how data is to be processed (e.g., which program is to be executed). In at least one embodiment, the scheduler 1510 can be configured to retrieve indices according to the tasks or to receive indices from the front-end unit 1508.In at least one embodiment, the front-end unit 1508 can be configured to ensure that the processing cluster array 1512 is brought into a valid state before a workload specified by an incoming instruction buffer (e.g., batch buffer, push buffer, etc.) is initiated.

[0100] In at least one embodiment, each of the one or more instances of the parallel processing unit 1502 can be coupled to a parallel processor memory 1522. In at least one embodiment, the parallel processor memory 1522 can be accessed via a memory crossbar 1516, which can receive memory requests from both the processing cluster array 1512 and the I / O unit 1504. In at least one embodiment, the memory crossbar 1516 can access the parallel processor memory 1522 via a memory interface 1518. In at least one embodiment, the memory interface 1518 can have several partitioning units (e.g., partitioning unit 1520A, partitioning unit 1520B to partitioning unit 1520N), each of which can be coupled to a section (e.g., a memory unit) of the parallel processor memory 1522.In at least one embodiment, the number of partitioning units 1520A-1520N is configured to equal the number of storage units, such that a first partitioning unit 1520A corresponds to a corresponding first storage unit 1524A, a second partitioning unit 1520B corresponds to a corresponding storage unit 1524B, and an Nth partitioning unit 1520N corresponds to a corresponding Nth storage unit 1524N. In at least one embodiment, the number of partitioning units 1520A-1520N cannot equal the number of storage units.

[0101] In at least one embodiment, the memory units 1524A-1524N can comprise various types of memory devices, including dynamic random-access memory (DRAM) or graphics random-access memory, such as synchronous graphics random-access memory (SGRAM), including double-data-rate graphics memory (GDDR). In at least one embodiment, the memory units 1524A-1524N can also comprise 3D stacked memory, including, but not limited to, high-bandwidth memory (HBM), HBM2e, or HDM3. In at least one embodiment, render targets, such as image memory or texture maps, can be stored on the memory units 1524A-1524N, allowing the partitioning units 1520A-1520N to write sections of each render target in parallel to efficiently utilize the available bandwidth of the parallel processor memory 1522.In at least one embodiment, a local instance of the parallel processor memory 1522 can be excluded in favor of a unified memory design that uses the system memory in conjunction with a local cache memory.

[0102] In at least one embodiment, each of the clusters 1514A-1514N of the processing cluster array 1512 can process data written to one of the storage units 1524A-1524N within the parallel processor memory 1522. In at least one embodiment, the memory crossbar 1516 can be configured to transfer an output from each cluster 1514A-1614N to any partitioning unit 1520A-1520N or to another cluster 1514A-1514N, which can perform additional processing operations on an output. In at least one embodiment, each cluster 1514A-1514N can communicate with the memory interface 1518 via the memory crossbar 1516 to read from or write to various external storage devices.In at least one embodiment, the memory crossbar 1516 has a connection to the memory interface 1518 for communication with the I / O unit 1504, as well as a connection to a local instance of the parallel processor memory 1522, enabling processing units within different processing clusters 1514A-1514N to communicate with system memory or other memory that is not local to the parallel processing unit 1502. In at least one embodiment, the memory crossbar 1516 can use virtual channels to separate traffic flows between the clusters 1514A-1514N and the partitioning units 1520A-1520N.

[0103] In at least one embodiment, multiple instances of the parallel processing unit 1502 can be provided on a single add-on card, or multiple add-on cards can be interconnected. In at least one embodiment, different instances of the parallel processing unit 1502 can be configured to work together, even if different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. For example, in at least one embodiment, some instances of the parallel processing unit 1502 can have higher-precision floating-point units compared to other instances.In at least one embodiment, systems containing one or more instances of the Parallel Processing Unit 1502 or the Parallel Processor 1500 can be implemented in a variety of configurations and form factors, including but not limited to desktop, laptop or handheld personal computers, servers, workstations, game consoles and / or embedded systems.

[0104] Fig. Figure 15B is a block diagram of a partitioning unit 1520 according to at least one embodiment. In at least one embodiment, the partitioning unit 1520 is an instance of one of the partitioning units 1520A-1520N. Fig. 15A. In at least one embodiment, the partitioning unit 1520 comprises an L2 cache 1521, an image buffer interface 1525, and a ROP 1526 (raster operation unit). In at least one embodiment, the L2 cache 1521 is a read / write cache configured to perform load and store operations received from the memory crossbar 1516 and the ROP 1526. In at least one embodiment, read errors and urgent write return requests are issued from the L2 cache 1521 to the image buffer interface 1525 for processing. In at least one embodiment, updates can also be sent to an image buffer for processing via the image buffer interface 1525. In at least one embodiment, the image buffer interface 1525 is connected to one of the storage units in the parallel processor memory, such as the storage units 1524A-1524N in Fig. 15A (e.g., within the parallel processor memory 1522), connected.

[0105] In at least one embodiment, the ROP 1526 is a processing unit that performs raster operations such as stenciling, Z-testing, blending, etc. In at least one embodiment, the ROP 1526 then outputs processed graphics data, which is stored in graphics memory. In at least one embodiment, the ROP 1526 includes compression logic for compressing depth or color data written to memory and for decompressing depth or color data read from memory. In at least one embodiment, the compression logic can be lossless compression logic that uses one or more of several compression algorithms. In at least one embodiment, the type of compression performed by the ROP 1526 can vary based on statistical properties of the data to be compressed.For example, in at least one embodiment, delta color compression is performed on depth and color data on a per-tile basis.

[0106] In at least one embodiment, the ROP 1526 is in each processing cluster (e.g., clusters 1514A-1514N in Fig. 15A) and not included in the partitioning unit 1520. In at least one embodiment, read and write requests for pixel data are transmitted via the memory crossbar 1516 instead of pixel fragment data. In at least one embodiment, processed graphics data can be displayed on a display device, for example, one of one or more display devices 1410. Fig. 14, displayed, forwarded for further processing by processors 1402 or for further processing by one of the processing units within the parallel processor 1500 Fig. 15A will be forwarded.

[0107] The embodiments presented herein can enable a linear controller with one or more features for improving the PSRR in order to identify and correct voltage noise within a circuit.

[0108] Fig. Figure 16 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, the system 1600 includes one or more processors 1602 and one or more graphics processors 1608 and can be a single-processor desktop system, a multi-processor workstation system, or a server system comprising a large number of processors 1602 or processor cores 1607. In at least one embodiment, the processing system 1600 is a processing platform integrated into an integrated circuit as a system-on-a-chip (SoC) for use in mobile, portable, or embedded devices. In at least one embodiment, one or more graphics processors 1608 comprise one or more graphics cores 1400.

[0109] In at least one embodiment, the System 1600 may comprise or be integrated into a server-based gaming platform, a game console comprising a game and media console, a mobile gaming console, a handheld game console, or an online game console. In at least one embodiment, the System 1600 is a mobile phone, a smartphone, a tablet computing device, or a mobile internet device. In at least one embodiment, the Processing System 1600 may also include, be coupled to, or be integrated into a wearable device, such as a smartwatch wearable device, a smart eyewear device, an augmented reality device, or a virtual reality device.In at least one embodiment, the processing system 1600 is a television or set-top box device comprising one or more processor(s) 1602 and a graphical user interface generated by one or more graphics processor(s) 1608.

[0110] In at least one embodiment, one or more processor(s) 1602 each include one or more processor core(s) 1607 for processing instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor core(s) 1607 is configured to process a specific instruction sequence 1609. In at least one embodiment, the instruction sequence 1609 can enable Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or computation using a Very Long Instruction Word (VLIW). In at least one embodiment, the processor core(s) 1607 can each process a different instruction sequence 1609, which may include instructions to facilitate the emulation of other instruction sequences. In at least one embodiment, the processor core(s) 1607 can...The processor core(s) 1607 also include other processing devices, such as a digital signal processor (DSP).

[0111] In at least one embodiment, the processor(s) 1602 include a cache memory 1604. In at least one embodiment, the processor(s) 1602 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory is shared by different components of the processor(s) 1602. In at least one embodiment, the processor(s) 1602 also use an external cache (e.g., a Level 3 ("L3") cache or Last-Level Cache ("LLC")) (not shown), which can be shared by the processor core(s) 1607 using known cache coherence techniques. In at least one embodiment, the processor(s) 1602 additionally includes a register file 1606, which may contain different types of registers for storing different types of data (e.g.,(Integer register, floating-point register, status register, and an instruction pointer register). In at least one embodiment, the register file 1606 may contain universal registers or other registers.

[0112] In at least one embodiment, one or more processors 1602 are coupled to one or more interface buses 1610 to transmit communication signals, such as address, data, or control signals, between the processor 1602 and other components in the system 1600. In at least one embodiment, the interface bus(s) 1610 can be a processor bus, such as a version of a Direct Media Interface (DMI) bus. In at least one embodiment, the interface bus(s) 1610 is not limited to a DMI bus and can include one or more peripheral component interconnection buses (e.g., PCI, PCI Express), memory buses, or other types of interface buses. In at least one embodiment, the processor(s) 1602 include an integrated memory controller 1616 and a platform control hub 1630.In at least one embodiment, the storage controller 1616 enables communication between a storage device and other components of the system 1600, while the platform controller hub (“PCH”) 1630 provides connections to input / output (“I / O”) devices via a local I / O bus.

[0113] In at least one embodiment, the storage device 1620 can be a dynamic random-access memory (DRAM) device, a static random-access memory (SRAM) device, a flash memory device, a phase-change memory device, or some other storage device that has suitable performance to serve as process memory. In at least one embodiment, the storage device 1620 can operate as system memory for the system 1600 to store data 1622 and instructions 1621 that are used when one or more processor(s) 1602 execute an application or process. In at least one embodiment, the memory controller 1616 is also coupled with an optional external graphics processor 1612 that can communicate with one or more graphics processor(s) 1608 in the processors 1602 to perform graphics and media operations.In at least one embodiment, a display device 1611 can be connected to the processor(s) 1602. In at least one embodiment, the display device 1611 can include one or more internal displays, such as those in a mobile electronic device or a laptop, or external displays connected via a display interface (e.g., DisplayPort, etc.). In at least one embodiment, the display device 1611 can include a head-mounted display (HMD), such as a stereoscopic display for use in virtual reality (VR) or augmented reality (AR) applications.

[0114] In at least one embodiment, the platform control hub 1630 enables peripheral devices to connect to the storage device 1620 and the processor 1602 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, without limitation, an audio controller 1646, a network controller 1634, a firmware interface 1628, a wireless transceiver 1626, touch sensors 1625, and a data storage device 1624 (e.g., a hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 1624 can be connected via a storage interface (e.g., SATA) or via a peripheral bus, such as a peripheral component interconnection bus (e.g., PCI, PCI Express). In at least one embodiment, the touch sensors 1625 can include touchscreen sensors, pressure sensors, or fingerprint sensors.In at least one embodiment, the wireless transceiver 1626 can be a WiFi transceiver, a Bluetooth transceiver, or a mobile network transceiver, such as a 3G, 4G, or Long-Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 1628 enables communication with the system firmware and can, for example, be a Unified Expandable Firmware Interface (UEFI). In at least one embodiment, the network controller 1634 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus(s) 1610. In at least one embodiment, the audio controller 1646 is a high-resolution, multi-channel audio controller.In at least one embodiment, the system 1600 includes an optional legacy I / O controller 1640 for coupling conventional devices (e.g., Personal System 2 (PS / 2)) to the system. In at least one embodiment, the platform control hub 1630 can also be connected to one or more universal serial bus (“USB”) controller(s) 1642, connection input devices such as keyboard and mouse 1643 combinations, a camera 1644, or other USB input devices.

[0115] In at least one embodiment, an instance of the memory controller 1616 and the platform control hub 1630 can be integrated into a discrete external graphics processor, such as an external graphics processor 1612. In at least one embodiment, the platform control hub 1630 and / or the memory controller 1616 can be external to the one or more processors 1602. For example, in at least one embodiment, the system 1600 can include an external memory controller 1616 and a platform control hub 1630, which can be configured as a memory control hub and peripheral control hub within a system chipset that communicates with the processor(s) 1602.

[0116] The embodiments presented here can enable a linear controller with one or more features for improving the PSRR in order to identify and correct voltage noise within a circuit.

[0117] Different embodiments can be described by the following sentences: 1. Computer-implemented method, including: Assigning an identifier to one or more data sources, wherein the identifier of the one or more data sources is made available via a control bus; Receiving a request to connect a data stream associated with a data source of one or more data sources to an available workload object from a variety of workload objects that is communicatively coupled to the control bus; Obtaining the identifier of one or more data sources and an identifier that is assigned to a specific workload object; Populating a routing table with the identifier of one or more data sources and the identifier assigned to the specific workload object; Performing a computational operation using the data stream through the specified workload object; and Updating the routing table with a second identifier that corresponds to a result generated by the specified workload object. 2. Computer-implemented method according to sentence 1, further comprising: Receiving a second request to connect to the result, where the second request is associated with a second workload object; and Providing routing information based on the second identifier of the result. 3. Computer-implemented method according to sentence 1, wherein at least one data source of the one or more data sources is a streaming data source. 4. Computer-implemented method according to sentence 1, further comprising: Receiving a third identifier from a second data source of one or more data sources; Determine that the third identifier is assigned to an entry in the routing table; and Replacing the third identifier with an assigned identifier for the entry in the routing table. 5. Computer-implemented method according to sentence 1, further comprising: Receiving a delete request from a workload object that is associated with the data stream; Determine that one or more other workload objects are associated with the data stream; and Maintaining routing information for the data stream. 6. Computer-implemented method according to sentence 1, further comprising: Identifying one or more data sources; and Writing identification information for one or more data sources to the control bus. 7. Processor comprehensive: one or more processing units for: Storing a data stream identifier in a routing table in response to an activation message for the data stream; Receiving a request for the data stream; Identifying an available workload object based on the data stream identifier in response to the request; Generating a result identifier for an output of the workload object; and Storing the result identifier in the routing table, with the output being accessible based on the result identifier. 8. Processor according to sentence 7, wherein the activation message is provided via a communication bus by a data stream initiator. 9. Processor according to sentence 7, wherein the one or more processing units are further configured to: Determine that a capacity for at least one workload object is at or above a threshold; and In response to the determination that the capacity of at least one workload object is at or above the threshold, a new workload process is initiated. 10. Processor according to sentence 7, wherein the one or more processing units are further configured to: Receiving a second request from a second workload object to connect to the result; and Providing routing information based on the result identifier for the second workload object. 11. Processor according to sentence 7, wherein one or more circuits are further configured to: Receiving a third identifier for a second data stream; Determine that the third identifier is assigned to an entry in the routing table; and Replacing the third identifier with an assigned identifier for the entry in the routing table. 12. Processor according to sentence 7, wherein one or more circuits are further configured to: Receiving a delete request from a worker assigned to the data stream; Determine that one or more other workers are connected to the data stream; and Maintaining routing information for the data stream. 13. Processor according to sentence 7, wherein the processor is present in a system comprising at least one of: a system for performing simulation operations; a system for performing simulation operations to test or validate autonomous machine applications; a system for performing digital twin operations; a system for performing a light transport simulation; a system for displaying a graphical output; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or displaying virtual reality (VR) content; a system for generating or displaying augmented reality (AR) content; a system for generating or displaying mixed-reality (MR) content; a system that includes one or more virtual machines, VMs; a system that is at least partially implemented in a data center; a system for performing hardware tests using a simulation; a system for synthetic data generation; a system for performing generative AI operations; a system for performing operations using a large language model, LLM; a system for performing operations using a video language model, VLM; a collaborative content creation platform for 3D assets; or a system that is implemented at least partially using cloud computing resources. 14. System comprehensive: one or more processing units for updating a routing table with identifier information associated with an output received from a workload object that requests access to a data stream from a variety of data streams from a variety of data sources using identifier information for the data stream. 15. System according to sentence 14, wherein the identifier information for the data stream is a stream identifier, and wherein the output is assigned a result identifier. 16. System according to sentence 15, wherein the data stream is requested by a request with a specific format, the request including the stream identifier. 17. System according to sentence 15, wherein the one or more processing units are further configured to: Receiving a second request from a second workload object to connect to the output; and Providing routing information based on the result identifier for the second workload object. 18. System according to sentence 15, wherein the one or more processing units are further configured to: Receiving a third identifier for a second data stream; Determine that the third identifier is assigned to an entry in the routing table; and Replacing the third identifier with an assigned identifier for the entry in the routing table. 19. System according to sentence 14, wherein the one or more processing units are further configured to: Determine that a capacity for the workload object is at or above a threshold; and In response to the determination that the capacity of the workload object is at or above the threshold, a new workload process is initiated. 20. System according to sentence 14, wherein the system comprises at least one of: a system for performing simulation operations; a system for performing simulation operations to test or validate autonomous machine applications; a system for performing digital twin operations; a system for performing a light transport simulation; a system for displaying a graphical output; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or displaying virtual reality (VR) content; a system for generating or displaying augmented reality (AR) content; a system for generating or displaying mixed-reality (MR) content; a system that includes one or more virtual machines, VMs; a system that is at least partially implemented in a data center; a system for performing hardware tests using a simulation; a system for synthetic data generation; a system for performing generative AI operations; a system for performing operations using a large language model, LLM; a system for performing operations using a video language model, VLM; a collaborative content creation platform for 3D assets; or a system that is implemented at least partially using cloud computing resources.

[0118] In at least one embodiment, a single semiconductor platform can refer to a single, unified, semiconductor-based integrated circuit or chip. In at least one embodiment, multi-chip modules with enhanced connectivity can be used, simulating on-chip operation and providing significant improvements over the use of a conventional central processing unit (CPU) and a bus implementation. In at least one embodiment, different modules can also be arranged separately or in various combinations of semiconductor platforms, depending on the user's requirements.

[0119] In at least one embodiment, which is based on Fig. As referenced in Section 10, computer programs in the form of machine-readable executable code or computer control logic algorithms are stored in main memory 1004 and / or secondary storage. Computer programs enable the computer system 1000 to perform various functions according to at least one embodiment when executed by one or more processors. In at least one embodiment, main memory 1004, storage, and / or any other storage medium are possible examples of computer-readable media. In at least one embodiment, secondary storage can refer to any suitable device or system for storage, such as a hard disk drive and / or a removable storage device, which may be a floppy disk drive, a magnetic tape drive, a CD drive, a DVD drive, a recording device, a USB flash drive, etc.In at least one embodiment, the architecture and / or functionality of various previous . Fig. 1-6 in the context of a CPU 1002, a parallel processing system 1012, an integrated circuit that incorporates at least some of the capabilities of a CPU 1002 and / or a parallel processing system 1012, a chipset (e.g., a group of integrated circuits designed to function as a unit and sold to perform related functions, etc.) and / or any suitable combination of integrated circuits.

[0120] In at least one embodiment, the architecture and / or functionality of various previous Fig. 1-6 in the context of a general computer system, a system with circuits, a game console system for entertainment purposes, an application-specific system, and more. In at least one embodiment, the computer system 1000 can be in the form of a desktop computer, a laptop computer, a tablet computer, servers, supercomputers, a smartphone (e.g., a wireless handheld device), a personal digital assistant (“PDA”), a digital camera, a vehicle, a head-mounted display, an electronic handheld device, a mobile phone, a television, a workstation, game consoles, an embedded system, and / or any other type of logic.

[0121] In at least one embodiment, the parallel processing system 1012 comprises, among other things, a plurality of parallel processing units (“PPUs”) 1014 and associated memory 1016. In at least one embodiment, the PPUs 1014 are connected to a host processor or other peripheral devices via a connection 1018 and a switch 1020 or multiplexer. In at least one embodiment, the parallel processing system 1012 distributes computational tasks to PPUs 1014 that can be parallelized, for example, as part of distributing computational tasks across multiple threads of graphics processing units (“GPUs”). In at least one embodiment, the memory is shared by some or all of the PPUs 1014 and is accessible to them (e.g.,for read and / or write accesses), although such shared memory may result in performance degradation compared to using local memory and registers resident in a PPU 1014. In at least one embodiment, the operation of PPUs 1014s is synchronized by using an instruction such as __syncthreads(), whereby all threads in a block (e.g., running across multiple PPUs 1014s) reach a specific execution point of the code before proceeding.

[0122] In at least one embodiment, one or more of the techniques or methods described herein use a oneAPI programming model. In at least one embodiment, a oneAPI programming model refers to a programming model for interacting with different compute accelerator architectures. In at least one embodiment, oneAPI refers to an application programming interface (API) designed for interacting with different compute accelerator architectures. In at least one embodiment, a oneAPI programming model uses a DPC++ programming language. In at least one embodiment, a DPC++ programming language refers to a high-level language for the productivity of data-parallel programming. In at least one embodiment, a DPC++ programming language is based at least partially on C and / or C++ programming languages.In at least one embodiment, a oneAPI programming model is a programming model such as that developed by Intel Corporation in Santa Clara, CA.

[0123] In at least one embodiment, oneAPI and / or the oneAPI programming model is used to interact with various accelerator, GPU, and processor architectures and / or variants thereof. In at least one embodiment, oneAPl comprises a set of libraries that implement various functionalities. In at least one embodiment, oneAPl includes at least a oneAPI DPC++ library, a oneAPI math kernel library, a oneAPl data analysis library, a oneAPI deep neural network library, a oneAPI collective communication library, a oneAPl threading component library, a oneAPl video processing library, and / or variations thereof.

[0124] In at least one embodiment, a oneAPI DPC++ library, also referred to as oneDPL, is a library that implements algorithms and functions for accelerating DPC++ kernel programming. In at least one embodiment, oneDPL implements one or more Standard Template Library (STL) functions. In at least one embodiment, oneDPL implements one or more parallel STL functions. In at least one embodiment, oneDPL provides a set of library classes and functions, such as parallel algorithms, iterators, function object classes, a range-based API, and / or variations thereof. In at least one embodiment, oneDPL implements one or more classes and / or functions of a C++ standard library. In at least one embodiment, oneDPL implements one or more random number generator functions.

[0125] In at least one embodiment, a oneAPI mathematics kernel library, also referred to as oneMKL, is a library that implements various optimized and parallelized routines for different mathematical functions and / or operations. In at least one embodiment, oneMKL implements one or more BLAS (Basic Linear Algebra Subprograms) and / or LAPACK (Linear Algebra Package) routines for dense linear algebra. In at least one embodiment, oneMKL implements one or more BLAS routines for sparse matrix linear algebra. In at least one embodiment, oneMKL implements one or more random number generators (RNGs). In at least one embodiment, oneMKL implements one or more vector mathematics (VM) routines for mathematical operations on vectors. In at least one embodiment, oneMKL implements one or more Fast Fourier Transform (FFT) functions.

[0126] In at least one embodiment, a oneAPI data analysis library, also referred to as oneDAL, is a library that implements various data analysis applications and distributed computations. In at least one embodiment, oneDAL implements various algorithms for preprocessing, transformation, analysis, modeling, validation, and decision-making for data analysis in batch, online, and distributed processing modes. In at least one embodiment, oneDAL implements various C++ and / or Java APIs and various connectors to one or more data sources. In at least one embodiment, oneDAL implements DPC++ API extensions to a conventional C++ interface and enables the use of a GPU for various algorithms.

[0127] In at least one embodiment, a oneAPI library for deep neural networks, also referred to as oneDNN, is a library that implements various deep learning functions. In at least one embodiment, oneDNN implements various functions, algorithms, and / or variations thereof for neural networks, machine learning, and deep learning.

[0128] In at least one embodiment, a oneAPl collective communication library, also referred to as oneCCL, is a library that implements various applications for deep learning and machine learning workloads. In at least one embodiment, oneCCL is based on lower-level communication middleware, such as Message Passing Interface (MPI) and libfabrics. In at least one embodiment, oneCCL enables a number of deep learning-specific optimizations, such as prioritization, persistent operations, execution orders, and / or variations thereof. In at least one embodiment, oneCCL implements various CPU and GPU functions.

[0129] In at least one embodiment, a oneAPl library with building blocks for threads, also referred to as oneTBB, is a library that implements various parallelized processes for different applications. In at least one embodiment, oneTBB is used for task-based, shared parallel programming on a host. In at least one embodiment, oneTBB implements generic parallel algorithms. In at least one embodiment, oneTBB implements concurrently existing containers. In at least one embodiment, oneTBB implements a scalable memory allocator. In at least one embodiment, oneTBB implements a work-stealing task scheduler. In at least one embodiment, oneTBB implements low-level synchronization primitives.In at least one embodiment, oneTBB is compiler-independent and usable on various processors, such as GPUs, PPUs, CPUs and / or variations thereof.

[0130] In at least one embodiment, a oneAPL video processing library, also referred to as oneVPL, is a library used to accelerate video processing in one or more applications. In at least one embodiment, oneVPL implements various video decoding, encoding, and processing functions. In at least one embodiment, oneVPL implements various functions for media pipelines on CPUs, GPUs, and other accelerators. In at least one embodiment, oneVPL implements device detection and selection in media-centric and video analytics workloads. In at least one embodiment, oneVPL implements API primitives for buffer sharing without copying.

[0131] In at least one embodiment, a oneAPI programming model uses a DPC++ programming language. In at least one embodiment, a DPC++ programming language is a programming language that, without restriction, contains functionally similar versions of CUDA mechanisms for defining device code and distinguishing between device code and host code. In at least one embodiment, a DPC++ programming language may contain a subset of the functionality of a CUDA programming language. In at least one embodiment, one or more operations of the CUDA programming model are executed using a oneAPI programming model with a DPC++ programming language.

[0132] In at least one embodiment, each application programming interface (API) described herein is compiled by a compiler, interpreter, or other software tool into one or more instructions, operations, or other signals. In at least one embodiment, the compilation includes generating one or more machine-executable instructions, operations, or other signals from the source code. In at least one embodiment, an API compiled with one or more instructions, operations, or other signals, upon execution, causes one or more processors, such as the 1310 graphics processor, the 1340 graphics processor, the 1400 graphics core, the 1500 parallel processor, or other logic circuitry further described herein, to perform one or more arithmetic operations.

[0133] It should be noted that although the embodiments described here may refer to a CUDA programming model, the techniques and procedures described here can be used with any suitable programming model, such as HIP, oneAPl and / or variations thereof.

[0134] Other variations are within the spirit of the present disclosure. Thus, while various modifications and alternative constructions can be made with respect to the disclosed methods, certain illustrated embodiments are shown in the drawings and have been described in detail above. However, it is understood that the intention is not to limit the disclosure to the specific disclosed form or forms, but rather, on the contrary, to cover all modifications, alternative constructions, and equivalents that fall within the spirit and scope of the disclosure as defined in the attached claims.

[0135] The use of the terms "a," "an," "the," and similar referents in the context of describing disclosed embodiments (particularly in the context of the following claims) is to be interpreted as covering both the singular and the plural unless otherwise specified herein or the context clearly contradicts this, and not as defining an expression. The terms "comprising," "having," "including," and "containing" are to be interpreted as open expressions (i.e., in the sense of "including without being limited") unless otherwise specified. The term "connected" is to be interpreted as partially or completely contained within, attached to, or joined to one another when used unmodified and referring to physical connections, even if an element is inserted between them.The mention of value ranges herein is intended merely as a quick method of individually referring to each separate value falling within the range, unless otherwise stated herein, and each separate value is included in the description as if it were individually reproduced herein. In at least one embodiment, the use of the term "set" (e.g., "a set of objects") or "subset" is to be understood as a non-empty compilation comprising one or more elements, unless otherwise noted or the context contradicts it. Furthermore, unless otherwise stated or the context contradicts it, the term "subset" of a corresponding set does not necessarily mean a proper subset of the corresponding set, but the subset and the corresponding set may be the same.

[0136] Unless specifically stated otherwise or the context clearly contradicts it, connective language, such as phrases of the form "at least one of A, B, and C" or "at least one of A, B, and C," is otherwise to be understood in the context in which it is generally used to indicate that an object, expression, etc., can be either A, B, or C, or any non-empty subset of the sentence consisting of A, B, and C. For example, in the illustrated example of a sentence containing three elements, the connective phrases "at least one of A, B, and C" and "at least one of A, B, and C" refer to any one of the following sentences: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}.Thus, such connecting expressions are generally not intended to express that certain embodiments require the presence of at least one of A, at least one of B, and at least one of C. Additionally, unless otherwise stated or contradicted by the context, the term "plurality" also denotes a state of plurality (e.g., "a plurality of elements" denotes multiple elements). In at least one embodiment, the number of elements in a plurality is at least two elements, but may be more if either explicitly stated or indicated by the context. Furthermore, unless otherwise stated or evident from the context, the phrase "based on" means "at least partially based on" and not "exclusively based on."

[0137] The operations of processes described herein may be performed in any suitable order, unless otherwise specified herein or the context clearly precludes it. In at least one embodiment, a process, such as the processes described herein (or variations and / or combinations thereof), is carried out under the control of one or more computing systems configured with executable instructions, and is implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed together on one or more processors, by hardware or combinations thereof. In at least one embodiment, code is stored on a computer-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions that can be executed by one or more processors.In at least one embodiment, a computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transitory signals (e.g., a propagating transient electrical or electromagnetic transmission) but includes non-transitory data storage circuits (e.g., buffers, caches, and queues) within transient signal senders / receivers. In at least one embodiment, the code (e.g., executable code or source code) is stored on a set of one or more non-transitory computer-readable storage media containing executable instructions (or other storage for executable instructions) which, when executed (i.e., as a result of execution) by one or more processors of a computing system, cause the computing system to perform the operations described herein.In at least one embodiment, a set of nontransitory computer-readable storage media comprises multiple nontransitory computer-readable storage media. One or more of the individual nontransitory storage media do not contain the entire code, while multiple nontransitory computer-readable storage media collectively store the entire code. In at least one embodiment, the executable instructions are executed such that different instructions are executed by different processors—for example, a nontransitory computer-readable storage medium stores instructions, and a central processing unit (CPU) executes some of the instructions, while a graphics processing unit (GPU) executes other instructions.In at least one embodiment, different components of a computing system have separate processors, and different processors execute different subsets of instructions.

[0138] In at least one embodiment, an arithmetic logic unit is a set of combinational logic circuits that accept one or more inputs to produce a result. In at least one embodiment, an arithmetic logic unit is used by a processor to implement mathematical operations such as addition, subtraction, or multiplication. In at least one embodiment, an arithmetic logic unit is used to implement logical operations such as logical AND / OR or XOR. In at least one embodiment, an arithmetic logic unit is stateless and consists of physical switching components such as semiconductor transistors arranged to form logic gates. In at least one embodiment, an arithmetic logic unit can operate internally as a stateful logic circuit with an associated clock.In at least one embodiment, an arithmetic logic unit can be constructed as an asynchronous logic circuit with an internal state that is not maintained in an associated register set. In at least one embodiment, an arithmetic logic unit is used by a processor to combine operands stored in one or more registers of the processor and to generate an output that can be stored by the processor in another register or in memory location.

[0139] In at least one embodiment, the processor, as a result of processing an instruction retrieved from the processor, presents one or more inputs or operands to an arithmetic logic unit (ALU), causing the ALU to generate a result that is at least partially based on an instruction code provided to the ALU's inputs. In at least one embodiment, the instruction codes provided by the processor to the ALU are at least partially based on the instruction executed by the processor. In at least one embodiment, combinational logic within the ALU processes the inputs and generates an output that is placed on a bus within the processor.In at least one embodiment, the processor selects a destination register, memory location, device, or output memory location on the output bus, such that the clocking of the processor causes the results generated by the ALU to be sent to the desired location.

[0140] In this application, the term arithmetic logic unit or ALU is used to refer to any computational logic circuit that processes operands to produce a result. For example, the term ALU in this document may refer to a floating-point unit, a DSP, a tensor core, a shader core, a coprocessor, or a CPU.

[0141] Accordingly, in at least one embodiment, computing systems are configured to implement one or more services that, individually or collectively, perform operations of the processes described herein, and such computing systems are configured with applicable hardware and / or software that enables the execution of operations. Furthermore, a computing system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment, it is a distributed computing system comprising several devices that operate differently, such that the distributed computing system performs the operations described herein and such that a single device does not perform all operations.

[0142] The use of any and all examples or illustrative language (e.g., "such as") provided in this document is intended solely to better clarify the embodiments of the disclosure and does not constitute a limitation of the scope of the disclosure unless claimed otherwise. No wording in the description should be interpreted as indicating any unclaimed element as essential to the implementation of the disclosure.

[0143] Any references, including publications, patent applications and patents mentioned in this document, are hereby incorporated by reference to the same extent as if each reference had been individually and specifically indicated as being included by reference and set forth in this document in its entirety.

[0144] In the description and claims, the terms "coupled" and "connected," along with their derivatives, may be used. It is understood that these terms are not intended to be synonymous. Rather, in specific examples, "connected" or "coupled" can be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" can also mean that two or more elements are not in direct contact with each other, but nevertheless interact or work together.

[0145] Unless expressly stated otherwise, terms such as "processing", "calculating", "calculating", "determining" or the like throughout this description are understood to refer to actions and / or processes of a computer or computing system or similar electronic computing device that manipulate and / or convert data represented as physical, e.g. electronic, quantities in the registers and / or memory of the computing system into other data represented in a similar manner as physical quantities in the memory, registers or other such information storage, transmission or display devices of the computing system.

[0146] Similarly, the term "processor" can refer to any device or section of a device that processes electronic data from registers and / or memories and converts that electronic data into other electronic data that can be stored in registers and / or memories. As non-restrictive examples, the "processor" can be a CPU or a GPU. A "computing platform" can include one or more processors. As used herein, "software" processes can include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Furthermore, each process can refer to multiple processes for executing instructions sequentially or in parallel, continuously or intermittently.In at least one embodiment, the terms “system” and “method” are used interchangeably in this document insofar as a system can embody one or more methods and the methods can be considered as a system.

[0147] This document may refer to the acquisition, capture, reception, or input of analog or digital data into a subsystem, computer system, or computer-implemented machine. In at least one embodiment, a process step of acquiring, capturing, receiving, or inputting analog and digital data can be performed in various ways, such as receiving data as a parameter of a function call or a call to an application programming interface. In at least one embodiment, the process step of acquiring, capturing, receiving, or inputting analog or digital data can be achieved by transmitting data via a serial or parallel interface. In at least one embodiment, the processes can be...The process steps of acquiring, capturing, receiving, or inputting analog or digital data are carried out by transmitting data over a computer network from the providing entity to the receiving entity. In at least one embodiment, reference may also be made to providing, outputting, transmitting, sending, or displaying analog or digital data. In various examples, the processes or process steps of providing, outputting, transmitting, sending, or displaying analog or digital data can be carried out by transmitting data as input or output parameters of a function call, a parameter of an application programming interface, or an interprocess communication mechanism.

[0148] Although the descriptions presented here are exemplary implementations of the described techniques, other architectures may also be used to implement the described functionality, and they are intended to be within the scope of this disclosure. Furthermore, although specific distributions of responsibilities are defined above for the purpose of discussion, various functions and responsibilities may be distributed and divided differently depending on the circumstances.

[0149] Although the subject matter has been further described in language specific to structural features and / or process steps, it is understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or steps described. Rather, specific features and steps are disclosed as exemplary ways of implementing the claims.

Claims

[1] Computer-implemented method, comprising: Assigning an identifier to one or more data sources, wherein the identifier of the one or more data sources is made available via a control bus; Receiving a request to connect a data stream associated with a data source of one or more data sources to an available workload object from a variety of workload objects that is communicatively coupled to the control bus; Obtaining the identifier of one or more data sources and an identifier that is assigned to a specific workload object; Populating a routing table with the identifier of one or more data sources and the identifier assigned to the specific workload object; Performing a computational operation using the data stream through the specified workload object; and Updating the routing table with a second identifier that corresponds to a result generated by the specified workload object. [2] Computer-implemented method according to claim 1, further comprising: Receiving a second request to connect to the result, where the second request is associated with a second workload object; and Providing routing information based on the second identifier of the result. [3] Computer-implemented method according to claim 1 or 2, wherein at least one data source of the one or more data sources is a streaming data source. [4] Computer-implemented method according to any one of the preceding claims, further comprising: Receiving a third identifier from a second data source of one or more data sources; Determine that the third identifier is assigned to an entry in the routing table; and Replacing the third identifier with an assigned identifier for the entry in the routing table. [5] Computer-implemented method according to any one of the preceding claims, further comprising: Receiving a delete request from a workload object that is associated with the data stream; Determine that one or more other workload objects are associated with the data stream; and Maintaining routing information for the data stream. [6] Computer-implemented method according to any one of the preceding claims, further comprising: Identifying one or more data sources; and Writing identification information for one or more data sources to the control bus. [7] Processor including: one or more processing units for: Storing a data stream identifier in a routing table in response to an activation message for the data stream; Receiving a request for the data stream; Identifying an available workload object based on the data stream identifier in response to the request; Generating a result identifier for an output of the workload object; and Storing the result identifier in the routing table, with the output being accessible based on the result identifier. [8] Processor according to claim 7, wherein the activation message is provided via a communication bus by a data stream initiator. [9] Processor according to claim 7 or 8, wherein the one or more processing units are further configured to: Determine that a capacity for at least one workload object is at or above a threshold; and In response to the determination that the capacity of at least one workload object is at or above the threshold, a new workload process is initiated. [10] Processor according to any one of claims 7 to 9, wherein the one or more processing units are further configured to: Receiving a second request from a second workload object to connect to the result; and Providing routing information based on the result identifier for the second workload object. [11] Processor according to one of claims 7 to 10, wherein one or more circuits are further configured to: Receiving a third identifier for a second data stream; Determine that the third identifier is assigned to an entry in the routing table; and Replacing the third identifier with an assigned identifier for the entry in the routing table. [12] Processor according to any one of claims 7 to 11, wherein the one or more circuits are further configured to: Receiving a delete request from a worker assigned to the data stream; Determine that one or more other workers are connected to the data stream; and Maintaining routing information for the data stream. [13] Processor according to any one of claims 7 to 12, wherein the processor is present in a system comprising at least one of: a system for performing simulation operations; a system for performing simulation operations to test or validate autonomous machine applications; a system for performing digital twin operations; a system for performing a light transport simulation; a system for displaying a graphical output; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or displaying virtual reality (VR) content; a system for generating or displaying augmented reality (AR) content; a system for generating or displaying mixed-reality (MR) content; a system that includes one or more virtual machines, VMs; a system that is at least partially implemented in a data center; a system for performing hardware tests using a simulation; a system for synthetic data generation; a system for performing generative AI operations; a system for performing operations using a large language model, LLM; a system for performing operations using a video language model, VLM; a collaborative content creation platform for 3D assets; or a system that is implemented at least partially using cloud computing resources. [14] System encompassing: one or more processing units for updating a routing table with identifier information associated with an output received from a workload object that requests access to a data stream from a multitude of data streams from a multitude of data sources using identifier information for the data stream. [15] System according to claim 14, wherein the identifier information for the data stream is a stream identifier, and wherein the output is assigned a result identifier. [16] System according to claim 15, wherein the data stream is requested by a request with a specific format, the request including the stream identifier. [17] System according to claim 15 or 16, wherein the one or more processing units are further configured to: Receiving a second request from a second workload object to connect to the output; and Providing routing information based on the result identifier for the second workload object. [18] System according to any one of claims 15 to 17, wherein the one or more processing units are further configured to: Receiving a third identifier for a second data stream; Determine that the third identifier is assigned to an entry in the routing table; and Replacing the third identifier with an assigned identifier for the entry in the routing table. [19] System according to any one of claims 14 to 18, wherein the one or more processing units are further configured to: Determine that a capacity for the workload object is at or above a threshold; and In response to the determination that the capacity of the workload object is at or above the threshold, a new workload process is initiated. [20] System according to any one of claims 14 to 19, wherein the system comprises at least one of: a system for performing simulation operations; a system for performing simulation operations to test or validate autonomous machine applications; a system for performing digital twin operations; a system for performing a light transport simulation; a system for displaying a graphical output; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or displaying virtual reality (VR) content; a system for generating or displaying augmented reality (AR) content; a system for generating or displaying mixed-reality (MR) content; a system that includes one or more virtual machines, VMs; a system that is at least partially implemented in a data center; a system for performing hardware tests using a simulation; a system for synthetic data generation; a system for performing generative AI operations; a system for performing operations using a large language model, LLM; a system for performing operations using a video language model, VLM; a collaborative content creation platform for 3D assets; or a system that is implemented at least partially using cloud computing resources.