Dynamic throttling of workload-based video processing functions using machine learning

By monitoring the resource usage and performance of the video system and dynamically adjusting parameters, the performance degradation problem caused by insufficient resources in the live video system was solved, achieving stable video stream transmission and processing, and maintaining user experience and service quality.

CN115955470BActive Publication Date: 2026-04-24NVIDIA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NVIDIA CORP
Filing Date
2022-09-28
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Live video systems suffer from performance degradation due to insufficient resources during video stream transmission and processing, which affects user experience and is difficult to solve effectively with existing technologies.

Method used

The system management unit monitors resource usage and performance, dynamically adjusts video processing and streaming parameters, including adjustable parameters for storage, processing, and streaming, and performs dynamic throttling based on policies and workload data to maintain service quality.

Benefits of technology

It achieves deterministic degradation of video streaming and processing under resource constraints, ensuring the stability of user experience and service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115955470B_ABST
    Figure CN115955470B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to dynamic throttling of workload-based video processing functions using machine learning. Systems and methods are disclosed for dynamically throttling video processing and / or streaming based on workload. Live video is captured from one or more sources (e.g., cameras) and stored. The video is then provided to a video processing engine and a video streaming engine. The video processing engine can perform one or more operations such as object detection, object tracking, and object classification to produce characterization data (e.g., bounding boxes, object tracks, alerts, object labels, object counts, boundary crossings, intersection highlights, etc.). System resource usage and performance of the video processing and streaming are monitored to produce workload data (e.g., metrics). Based on a policy and the workload data, the video streaming and / or processing is dynamically reconfigured by adjusting parameters provided to the video streaming and processing engines.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] A live video system receives video captured by one or more cameras and streams the video to one or more remote clients, while simultaneously processing the video, performing object detection, classification, and / or tracking. System performance may be affected depending on the combined video streaming and processing workload. Specifically, video streaming performance may degrade in terms of measured metrics, including frame rate, resolution, and / or bitrate. Video processing performance may simply fail or degrade (e.g., reduced image resolution, dropped frames, etc.) when computational and / or storage resources are insufficient to perform the processing operations. These issues and / or other problems related to the prior art need to be addressed. Summary of the Invention

[0002] Embodiments of this disclosure relate to dynamic throttling of video processing functions based on workload. Systems and methods for dynamically throttle video processing and / or streaming based on workload are disclosed. Live video is captured and stored from one or more sources (e.g., cameras). The video is then provided to a video processing engine and a video streaming engine. The video processing engine may perform one or more operations, such as object detection, object tracking, object classification, segmentation, pose detection, and face recognition, to generate representational data (e.g., bounding boxes, object trajectories, alarms, object labels, object counts, boundary intersections, intersection highlighting, etc.). The video streaming engine transmits the video to one or more streaming clients according to playback controls per client.

[0003] Unlike conventional systems, such as those described above that partition system resources (e.g., CPU, GPU, memory, network bandwidth) for exclusive use by processing or streaming components, the system resources of one or more embodiments may be shared. System resource usage and performance for video processing and streaming are monitored to generate workload data (e.g., metrics). Performance metrics for streaming may include the number of clients and frame rate. A non-limiting example of a performance metric for processing may be whether the processing was successfully completed. For simple frames, processing may succeed, but for more complex frames (e.g., frames with an increased number of moving objects), processing may fail.

[0004] Adjustable parameters for video streaming can include the number of clients, frame rate, resolution, and bitrate. Adjustable parameters for video processing can control frame rate (frame sampling), frame resolution, computational precision (inference resolution), tracking distance (tracking window), classification priority, the number of video sources to be processed, and the frequency of clip uploads. Parameters can be adjusted based on policies that include the priority of processing relative to streaming, thresholds for each metric, and priority for each processing operation. Based on policies and workload data, video streaming and / or processing can be dynamically reconfigured by adjusting the parameters provided to the video streaming and processing engines. Controlling performance degradation based on defined policies provides deterministic degradation and maintains Quality of Service (QoS).

[0005] A system, method, and computer-readable medium for dynamically reconfiguring video streaming transmission and / or processing are described. In one embodiment, a video stream causing at least one capture is transmitted by a video system to multiple remote clients. At least partially simultaneously with the transmission, the at least one captured video stream is processed by the video system to produce characterization data corresponding to a scene depicted in the at least one captured video stream, wherein the transmission and processing contribute to system workload. In response to determining the system workload triggering a policy-based action, processing parameters are dynamically adjusted during the processing of the at least one captured video stream, the processing parameters controlling at least one function performed by the video system. In one embodiment, the at least one captured video stream is transmitted using a streaming protocol. Attached Figure Description

[0006] The system and method for dynamic throttling of workload-based video processing functions are described in detail below with reference to the accompanying drawings, wherein:

[0007] Figure 1 A block diagram of an example video system is shown, which dynamically throttles video processing and / or streaming based on workloads suitable for implementing some embodiments of this disclosure.

[0008] Figure 2 A flowchart of a method for dynamic throttling of a workload-based video processing function, according to an embodiment, is shown.

[0009] Figure 3 A block diagram of another example video system is shown, which dynamically throttles video processing and / or streaming based on workloads suitable for implementing some embodiments of this disclosure.

[0010] Figure 4 Example parallel processing units suitable for implementing some embodiments of this disclosure are shown.

[0011] Figure 5AThis is suitable for use in implementing some embodiments of this disclosure. Figure 4 A conceptual diagram of the processing system implemented by the PPU.

[0012] Figure 5B An exemplary system is shown in which various architectures and / or functions of the various previous embodiments can be implemented.

[0013] Figure 5C Components of an exemplary system that can be used to train and utilize machine learning in at least one embodiment are shown.

[0014] Figure 6 An exemplary streaming system suitable for implementing some embodiments of this disclosure is shown. Detailed Implementation

[0015] Systems and methods related to dynamic throttling of workload-based video processing functions are disclosed. The workload of a video processing and streaming system varies based on the number of clients streaming video and the video content being processed. When the workload exceeds the available resources of the video system, uncontrolled performance degradation of video streaming and / or processing may occur. Uncontrolled degradation can negatively impact the user experience. For example, if frames cannot be stored, transmitted, transferred, or processed due to bandwidth, capacity, and / or processing limitations, the frame rate and / or image resolution may decrease.

[0016] To avoid uncontrolled performance degradation, the video system monitors workload data associated with video processing and streaming and dynamically adjusts the operating parameters controlling these functions. In other words, processing and / or streaming operations are dynamically throttled based on the video system workload. In one embodiment, throttling is achieved by adjusting operating parameters based on policies defined by the user and / or system designer. Policy-based degradation control provides deterministic degradation and maintains QoS.

[0017] Figure 1A block diagram of an example video system 100 is shown, which dynamically throttles video processing and / or streaming based on workloads suitable for implementing some embodiments of this disclosure. It should be understood that such and other arrangements described herein are illustrative by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, commands, function groups, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in combination with other components, and implemented in any suitable combination and location. The various functions described herein as being performed by entities can be performed by hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory. Moreover, those skilled in the art will understand that any system performing the operation of video system 100 is within the scope and spirit of embodiments of this disclosure.

[0018] Video system 100 includes a video input and storage unit 110, a video processing unit 120, a video streaming unit 130, and a system management unit 140. Instead of partitioning system resources (e.g., CPU, GPU, memory, network bandwidth) for exclusive use by processing or streaming components, system resources are shared for both processing and streaming operations. System management unit 140 monitors system resource usage and the performance of video processing and streaming to generate workload data (e.g., metrics). In one embodiment, workload data is provided by each of the video input and storage unit 110, video processing unit 120, and video streaming unit 130.

[0019] Parameters can be adjusted according to strategies to throttle video processing and streaming based on system workload. These parameters control the functions performed by the video input and storage unit 110, video processing unit 120, and video streaming unit 130. Based on strategies and workload data, the system management unit 140 can dynamically reconfigure one or more of the video input and storage operations, video streaming operations, and / or video processing operations by adjusting the parameters provided to the video input and storage unit 110, video processing unit 120, and video streaming unit 130, respectively. These strategies may include processing priority relative to streaming, thresholds for each metric, and priority for each processing operation.

[0020] Live video can be captured from one or more sources (e.g., cameras) and received at the video input and storage unit 110. For example... Figure 1 As shown, N >A captured video stream (video 1, ..., video N) is input to the video input and storage unit 110. In one embodiment, each captured video stream is compressed and stored within or by the video input and storage unit 110. In one embodiment, the number of captured video streams that are compressed and / or stored is controlled by storage parameters. In one embodiment, each source capture may include a video stream of pixel data for a sequence of video frames, wherein the pixel data represents color, grayscale, luminance, infrared, depth, or other types of data.

[0021] The video input and storage unit 110 provides storage performance data 114 to the system management unit 140. Storage performance data 114 may include one or more of the following: storage resource usage, number of video streams, pixel resolution, frame rate, frame complexity data, etc. In one embodiment, frame complexity data corresponds to image complexity and may be determined based on the compression level (e.g., compression rate or compression ratio) implemented for one or more frames. Storage performance data 114 may be provided for each source or combined for one or more captured video streams. In one embodiment, storage parameters 116, dynamically adjusted by the system management unit 140, may control the use of storage resources and / or the processing of captured video streams. For example, but not limited to, storage parameters 116 may control the number of captured video streams that are compressed and / or stored, the number and / or frequency of frames that are compressed and / or stored, the pixel resolution of each frame that is compressed and / or stored, and the compression level. Individual storage parameters 116 may be provided for each captured video stream, or a single storage parameter 116 may be used for all captured video streams.

[0022] At least one captured video stream is output from video input and storage unit 110 to video stream transmission unit 130 and video processing unit 120. At least one captured video stream can be read from storage and decompressed before being output as one or more video input streams 105 and / or 115. In one embodiment, at least one captured video stream is simultaneously output to both video stream transmission unit 130 and video processing unit 120 as one or more video input streams 105 and 115. In one embodiment, at least one captured video stream is output to video stream transmission unit 130 and video processing unit 120 as one or more video input streams 105 and 115, respectively, so that video stream transmission unit 130 and video processing unit 120 can receive the same frames at different times. In one embodiment, a first captured video stream is output to video stream transmission unit 130 as one or more video input streams 105, but not to video processing unit 120. In one embodiment, a first captured video stream is not output to video stream transmission unit 130, but is output to video processing unit 120 as one or more video input streams 115. In one embodiment, the first captured video stream is output to the video stream transmission unit 130 as one or more video input streams 105 at a different frame rate and / or resolution than the first captured video stream is output to the video processing unit 120 as one or more video input streams 115.

[0023] Video processing unit 120 processes at least one captured video stream to generate representational data 122. In one embodiment, video processing unit 120 may perform one or more operations, such as (but not limited to) object detection, object tracking, object classification, segmentation, pose detection, and object or face recognition, to generate representational data (e.g., bounding boxes, object trajectories, alerts, object labels, object counts, semantic masks, boundary crossings, intersection highlighting, etc.). Video processing unit 120 may include one or more neural network models for performing the operations listed above, such as (but not limited to) object detection, object tracking, object classification, segmentation, pose detection, and object or face recognition.

[0024] In one embodiment, probability or accuracy data may be associated with characterization data 122. In one embodiment, probability or accuracy data may include processing performance data 124. In one embodiment, when probability or accuracy data falls below a defined confidence level, one or more frames of the video input stream 105 processed to generate characterization data 122 associated with the low confidence level are identified and stored as low-confidence video clips. In one embodiment, characterization data 122 may also include frame clips associated with anomalous behavior detected by the video processing unit 120. In one embodiment, low-confidence and / or anomalous video clips may be associated with alarms, used as training data, or transmitted for further analysis.

[0025] The video processing unit 120 provides processing performance data 124 to the system management unit 140. In one embodiment, the processing performance data 124 includes clip generation rate and / or data corresponding to the workload of video processing resources. In one embodiment, the processing performance data indicates whether the video processing function performed by the video processing unit 120 was successfully completed. For simple frames, processing may succeed, but for more complex frames (with an increased number of objects), processing may fail. The processing performance data 124 may be provided for each video input stream 105 or may be combined for one or more video input streams 105.

[0026] The adjustable processing parameters 126 provided by the system management unit 140 to the video processing unit 120 can control the functions performed by the video processing unit 120 to generate representation data 122. For example, the processing parameters 126 can control the computational accuracy in terms of the resolution of the feature maps processed during inference. The adjustable processing parameters 126 provided by the system management unit 140 to the video processing unit 120 can also control the number of one or more video input streams 105 processed, the number and / or frequency of frames processed (e.g., frame rate or frame sampling), the pixel resolution of each frame processed, the tracking distance (tracking window) used during processing, and the clip generation frequency. In one embodiment, when classifying multiple object types, the adjustable processing parameters 126 provided by the system management unit 140 to the video processing unit 120 can also control the classification priority. For example, the classification priority can be used to control the maximum number of objects identified and classified in each frame, or to prioritize the classification of a first object type over other object types (i.e., identifying humans before animals). Individual processing parameters 126 can be provided for each video input stream 105, or a single processing parameter 126 can be used for all video input streams 105. The processing parameter 126 can be adjusted based on processing performance data 124 and strategies related to workload data.

[0027] The video streaming unit 130 transmits (e.g., transmits or streams) one or more video input streams 115 to one or more video streaming clients 125 according to the playback control of each client. In one embodiment, a Real-Time Streaming Protocol (RTSP) can be used to perform the transmission. Examples of real-time streaming protocols include User Datagram Protocol (UDP), Secure and Reliable Transport (SRT), and Web Real-Time Communication (WebRTC).

[0028] In one embodiment, the number of video streaming clients 125 is dynamic and can increase or decrease at any time. Video streaming unit 130 provides streaming performance data 134 to system management unit 140. In one embodiment, streaming performance data 134 includes transmission bandwidth data, such as available and / or consumed bandwidth. In one embodiment, streaming performance data 134 includes the number of video streaming clients 125. In one embodiment, the streaming workload increases linearly with the number of video streaming clients 125. Streaming performance data 134 can be provided for each video input stream 115, or streaming performance data 134 can be combined for one or more video input streams 115. Similarly, streaming performance data 134 can be provided for each video streaming client 125, or streaming performance data 134 can be combined for video streaming clients 125.

[0029] The adjustable streaming parameters 136 provided by the system management unit 140 to the video streaming unit 130 may include the maximum number of video streaming clients 125, the number and / or frequency (e.g., frame rate or frame sampling) of frames transmitted to one or more video streaming clients 125, the pixel resolution of each frame transmitted to one or more video streaming clients 125, and the transmission bit rate of one or more video streaming clients 125.

[0030] In one embodiment, when transmission bandwidth is limited, the adjustable streaming parameters 136 provided by the system management unit 140 to the video streaming unit 130 can also control the priority of clients. Individual streaming parameters 136 can be provided for each video input stream 115, or a single streaming parameter 136 can be used for all video input streams 115. Similarly, streaming parameters 136 can be provided for each video streaming client 125, or streaming parameters 136 can be combined for each video streaming client 125. The streaming parameters 136 can be adjusted based on streaming performance data 134 and strategies related to workload data.

[0031] Based on policy and workload data, system management unit 140 dynamically reconfigures video input and storage units 110, video processing unit 120, and video streaming unit 130 by adjusting storage parameters 116, processing parameters 126, and / or streaming parameters 136. Workload data represents system workload and directly includes storage performance data 114, processing performance data 124, and streaming performance data 134, or is analyzed to generate metrics. In one embodiment, system management unit 140 may also adjust the number of enabled processors based on workload, power consumption policies, and environmental conditions (e.g., temperature) within the video system 100. For example, system management unit 140 may adjust power modes in response to system workload. In one embodiment, parameters (storage, processing, and / or streaming) can be adjusted at any time, including when video processing unit 120 is processing at least one captured video stream to generate characterization data 122 and / or when streaming processing unit 130 is transmitting one or more video input streams 115 to one or more video streaming clients 125.

[0032] Now, based on the user's needs, further illustrative information will be provided regarding the various optional architectures and features that can implement the aforementioned framework. It should be strongly noted that the following information is for illustrative purposes only and should not be construed as limiting in any way. Any of the following features may be optionally combined, excluding or not excluding the other features described.

[0033] Figure 2 A flowchart of a method 200 for dynamic throttling of a workload-based video processing function according to an embodiment is shown. Each block of the method 200 described herein includes a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in memory. The method can also be embodied as computer-usable instructions stored on a computer storage medium. The method can be provided by a standalone application, service, or hosting service (independently or in combination with another hosting service) or a plug-in to another product, to name a few. Furthermore, method 200 is illustrated by way of example. Figure 1 The video system 100 has been described. However, this method may be performed additionally or alternatively by any system or any combination of systems, including but not limited to those described herein. Furthermore, those skilled in the art will understand that any system performing method 200 is within the scope and spirit of the embodiments of this disclosure.

[0034] In step 210, at least one captured video stream is transmitted by the video system to multiple remote clients. In one embodiment, the video system includes video system 100. In step 220, simultaneously with the transmission, the video system processes at least one captured video stream to generate characterization data corresponding to the scene depicted in the at least one captured video stream, wherein the transmission and processing impose a system workload. In one embodiment, the processing performed by video processing unit 120 includes a processing portion of the system workload. In one embodiment, the processing performed by video streaming unit 130 includes a streaming portion of the system workload. In one embodiment, video input and storage performed by video input and storage unit 110 includes a storage portion of the system workload. In one embodiment, environmental conditions within the system, such as temperature, impose a system workload.

[0035] In one embodiment, the processing is implemented by a neural network configured to perform at least one of object detection, classification, or tracking. In one embodiment, the representation data includes at least one of bounding boxes, object trajectories, object labels, object counts, boundary intersections, or intersection highlighting. In one embodiment, the representation data includes video clips. In one embodiment, an alert is generated in response to the representation data.

[0036] In step 230, in response to determining that the system workload triggers a policy-based action, processing parameters are dynamically adjusted. These processing parameters control at least one function performed by the video system during the processing of at least one captured video stream. In one embodiment, the at least one function includes adjustments to at least one of frame rate, frame resolution, computational precision, tracking distance, classification priority, the number of video sources to be processed, and the frequency of video clip uploads. In one embodiment, computational precision includes the number of feature channels enabled for processing. In one embodiment, the power mode is adjusted in response to the system workload.

[0037] In one embodiment, in response to determining that the workload triggers a policy-based response, streaming parameters controlling at least one of the frame rate, bit rate, or resolution of at least one captured video stream are dynamically adjusted. In one embodiment, the system workload triggers a policy-based action when a change in the number of several remote clients, including the plurality of remote clients, is detected. In one embodiment, the plurality of remote clients includes video streaming client 125. In one embodiment, a standardized monitoring mechanism is employed to enable individual modules to expose their current workload and performance. For example, a framework can be used to enable video processing unit 120, video input and storage unit 110, and video streaming unit 130 to expose Hypertext Transfer Protocol (HTTP) endpoints that can be read in a standardized manner by system management unit 140 to retrieve associated metrics or performance data.

[0038] Figure 3 A block diagram of another example video system 300 is shown, which dynamically throttles video processing and / or streaming based on workload and is suitable for implementing some embodiments of this disclosure. The video system 300 includes a video input unit 310, a video processing unit 320, an analysis unit 335, a storage device 330, a gateway 340, and an alarm monitor 345. In one embodiment, the video input unit 310 includes video management software (VMS) that receives multiple captured video streams, from video 1 to video N. In one embodiment, the captured video streams are transmitted to a video streaming client 125 and / or the video processing unit 320 using a Real-time Streaming Protocol (RTSP). In one embodiment, the captured video streams are transmitted to the video streaming client 125 and / or the video processing unit 320 using an HTTP Live Streaming Protocol (HLS). The video input unit 310 also stores the captured video streams in the storage device 330. The video input unit 310 can perform one or more operations performed by the video input and storage unit 110, such as compressing the video streams.

[0039] Video processing unit 320 can be configured to perform one or more operations performed by video processing unit 120, including generating characterization data 122 and / or performance data. In one embodiment, characterization data 122 and performance data include metadata provided to analysis unit 335 via a remote dictionary server (REDIS). Analysis unit 335 can be configured to generate a dashboard representation of characterization data 122 and / or performance data for display via a web application programming interface (API). In one embodiment, analysis unit 335 performs one or more operations performed by system management unit 140. Analysis unit 335 can retrieve video clips stored in storage device 330 based on metadata.

[0040] Gateway 340 receives playback control initiated by video streaming client 125, and in response, one or more video streams stored in storage device 330 can be transmitted from video input unit 310 to video streaming client 125. In one embodiment, video streaming client 125 communicates with gateway 340 using the WebRTC protocol. In response to an alarm notification received from analysis unit 335, alarm monitor 345 retrieves and analyzes video clips stored in storage device 330.

[0041] Video systems 100 and 300 perform simultaneous video streaming and processing, as well as workload monitoring, enabling policy-based dynamic reconfiguration of video streaming and / or processing. Dynamic reconfiguration is achieved by adjusting processing, streaming, and storage parameters, with the goal of influencing system resource usage to provide deterministic performance degradation based on defined policies and maintaining QoS for video streaming clients.

[0042] Parallel processing architecture

[0043] Figure 4 A parallel processing unit (PPU) 400 according to one embodiment is illustrated. PPU 400 can be used to implement video systems 100 and / or 300. PPU 400 can be used to implement one or more of a video input and storage unit 110, a video processing unit 120, a video streaming unit 130, and a system management unit 140. In one embodiment, a processor such as PPU 400 can be configured to implement a neural network model. The neural network model can be implemented as software instructions executed by the processor, or in other embodiments, the processor can include a matrix of hardware elements configured to process a set of inputs (e.g., electrical signals representing values) to generate a set of outputs that can represent activations of the neural network model. In other embodiments, the neural network model can be implemented as a combination of processing and software instructions executed by the hardware element matrix. Implementing a neural network model can include determining a set of parameters for the neural network model through, for example, supervised or unsupervised training of the neural network model, and, alternatively, performing inference using that parameter set to process new input sets.

[0044] In one embodiment, PPU 400 is a multi-threaded processor implemented on one or more integrated circuit devices. PPU 400 is a latency-hiding architecture designed to process many threads in parallel. A thread (e.g., an execution thread) is an instantiation of an instruction set configured to be executed by PPU 400. In one embodiment, PPU 400 is a graphics processing unit (GPU) configured to implement a graphics rendering pipeline for processing three-dimensional (3D) graphics data to generate two-dimensional (2D) image data for display on a display device. In other embodiments, PPU 400 may be used to perform general-purpose computing. While an exemplary parallel processor is provided herein for illustrative purposes, it should be specifically noted that such a processor is illustrated for illustrative purposes only, and any processor may be employed to complement and / or replace this processor.

[0045] One or more PPU 400s can be configured to accelerate thousands of high-performance computing (HPC), data center, cloud computing, and machine learning applications. PPU 400s can be configured to accelerate numerous deep learning systems and applications used in autonomous vehicles, simulations, computational graphics such as ray or path tracing, deep learning, high-precision speech, image, and text recognition systems, intelligent video analytics, molecular simulations, drug discovery, disease diagnosis, weather forecasting, big data analytics, astronomy, molecular dynamics simulations, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations, among others.

[0046] like Figure 4 As shown, PPU 400 includes an input / output (I / O) unit 405, a front-end unit 415, a scheduler unit 420, a job allocation unit 425, a hub 430, a crossbar (Xbar) 470, one or more general purpose processing clusters (GPCs) 450, and one or more memory partitioning units 480. PPU 400 can be connected to a host processor or other PPU 400 via one or more high-speed NVLink 410 interconnects. PPU 400 can be connected to a host processor or other peripheral devices via interconnect 402. PPU 400 can also be connected to local memory 404, which includes multiple memory devices. In one embodiment, local memory may include multiple dynamic random access memory (DRAM) devices. The DRAM devices may be configured as a high-bandwidth memory (HBM) subsystem, wherein multiple DRAM dies are stacked within each device.

[0047] The NVLink 410 interconnect enables the system to expand and include one or more PPUs 400 in conjunction with one or more CPUs, supporting cache coherency between the PPUs 400 and the CPU, as well as the CPU controller. Data and / or commands can be sent from or from the NVLink 410 to other units of the PPU 400 via hub 430, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown). Figure 5B A more detailed description of the NVLink 410.

[0048] I / O unit 405 is configured to send and receive communications (e.g., commands, data, etc.) from a host processor (not shown) via interconnect 402. I / O unit 405 may communicate directly with the host processor via interconnect 402, or via one or more intermediate devices such as memory bridges. In one embodiment, I / O unit 405 may communicate with one or more other processors, such as one or more PPUs 400, via interconnect 402. In one embodiment, I / O unit 405 implements a Peripheral Component Interconnect High Speed ​​(PCIe) interface for communication via a PCIe bus, and interconnect 402 is a PCIe bus. In alternative embodiments, I / O unit 405 may implement other types of known interfaces for communication with external devices.

[0049] I / O unit 405 decodes data packets received via interconnect 402. In one embodiment, the data packets represent commands configured to cause PPU 400 to perform various operations. I / O unit 405 transmits the decoded commands to various other units of PPU 400 that these commands may specify. For example, some commands may be transmitted to front-end unit 415. Other commands may be transmitted to hub 430 or other units of PPU 400, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown). In other words, I / O unit 405 is configured to route communication between and among the various logical units of PPU 400.

[0050] In one embodiment, a program executed by the host processor encodes a command stream in a buffer that provides a workload to the PPU 400 for processing. The workload may include instructions and data to be processed by those instructions. The buffer is an area of ​​memory accessible (e.g., read / write) by both the host processor and the PPU 400. For example, I / O unit 405 may be configured to access a buffer in system memory connected to interconnect 402 via a memory request transmitted through interconnect 402. In one embodiment, the host processor writes a command stream to the buffer and then transmits a pointer to the start of the command stream back to the PPU 400. Front-end unit 415 receives pointers to one or more command streams. Front-end unit 415 manages the one or more streams, reads commands from these streams, and forwards the commands to the respective units of the PPU 400.

[0051] Front-end unit 415 is coupled to scheduler unit 420, which configures various GPCs 450 to process tasks defined by the one or more streams. Scheduler unit 420 is configured to track status information related to the various tasks managed by scheduler unit 420. Status can indicate which GPC 450 a task is assigned to, whether the task is active or inactive, the priority associated with the task, etc. Scheduler unit 420 manages the execution of multiple tasks on the one or more GPCs 450.

[0052] Scheduler unit 420 is coupled to job allocation unit 425, which is configured to dispatch tasks for execution on GPC 450. Job allocation unit 425 can track several scheduled tasks received from scheduler unit 420. In one embodiment, job allocation unit 425 manages a pending task pool and an active task pool for each GPC 450. When GPC 450 completes the execution of a task, the task is evicted from the active task pool of GPC 450, and one of the other tasks from the pending task pool is selected and scheduled for execution on GPC 450. If an active task on GPC 450 is idle, for example, while waiting for data dependencies to be resolved, then the active task can be evicted from GPC 450 and returned to the pending task pool, while another task in the pending task pool is selected and scheduled for execution on GPC 450.

[0053] In one embodiment, the host processor executes a driver kernel that implements an application programming interface (API), enabling one or more applications executing on the host processor to be scheduled for operations to be performed on the PPU 400. In one embodiment, multiple computing applications are executed concurrently by the PPU 400, and the PPU 400 provides isolation, Quality of Service (QoS), and independent address spaces for the multiple computing applications. Applications may generate instructions (e.g., API calls) that cause the driver kernel to generate one or more tasks to be executed by the PPU 400. The driver kernel outputs the tasks to one or more streams being processed by the PPU 400. Each task may include one or more associated thread groups, referred to herein as warps. In one embodiment, a warp includes 32 associated threads that can execute in parallel. Cooperative threads may refer to multiple threads that include instructions for performing tasks and can exchange data via shared memory. These tasks may be assigned to one or more processing units within the GPC 450, and instructions are scheduled for execution by at least one warp.

[0054] The work allocation unit 425 communicates with one or more GPCs 450 via an XBar 470. The XBar 470 is an interconnect network that couples a plurality of units of the PPU 400 to other units of the PPU 400. For example, the XBar 470 can be configured to couple the work allocation unit 425 to a specific GPC 450. Although not explicitly shown, one or more other units of the PPU 400 may also be connected to the XBar 470 via a hub 430.

[0055] Tasks are managed by scheduler unit 420 and dispatched to GPC 450 by work allocation unit 425. GPC 450 is configured to process tasks and generate results. Results can be consumed by other tasks within GPC 450, routed to different GPCs 450 via XBar 470, or stored in memory 404. Results can be written to memory 404 via memory partitioning unit 480, which implements a memory interface for reading data from and writing data to memory 404.

[0056] The results can be transferred to another PPU 400 or CPU via NVLink 410. In one embodiment, PPU 400 includes U memory partition units 480, the number of which is equal to the number of individual and distinct memory devices coupled to memory 404 of PPU 400. Each GPC 450 may include a memory management unit to provide virtual address to physical address translation, memory protection, and arbitration of memory requests. In one embodiment, the memory management unit provides one or more Translation Backing Buffers (TLBs) for performing virtual address to physical address translation in memory 404.

[0057] In one embodiment, memory partitioning unit 480 includes a raster operation (ROP) unit, a secondary (L2) cache, and a memory interface coupled to memory 404. The memory interface can implement 32-bit, 64-bit, 128-bit, or 1024-bit data buses for high-speed data transfer. PPU 400 can connect to up to Y memory devices, such as high-bandwidth memory stacks or graphics dual data rate, version 5, synchronous dynamic random access memory, or other types of persistent storage devices. In one embodiment, the memory interface implements an HBM2 memory interface, and Y equals half of U. In one embodiment, the HBM2 memory stack is located on the same physical package as PPU 400, providing significant power and area savings compared to conventional GDDR5 SDRAM systems. In one embodiment, each HBM2 stack includes four memory dies and Y equals 4, wherein each HBM2 stack includes two 129-bit channels per die, for a total of eight channels, and the data bus width is 1024 bits.

[0058] In one embodiment, memory 404 supports single error correction double detection (SECDED) error correction codes (ECC) to protect data. ECC provides a high level of reliability for computing applications sensitive to data corruption. Reliability is particularly important where the PPU 400 is handling very large datasets and / or in large-scale cluster computing environments where applications run for extended periods.

[0059] In one embodiment, PPU 400 implements a multi-level memory hierarchy. In one embodiment, memory partitioning unit 480 supports unified memory that provides a single, unified virtual address space for the CPU and PPU 400 memory, allowing data sharing between virtual memory systems. In one embodiment, the frequency with which PPU 400 accesses memory located on other processors is tracked to ensure that memory pages are moved to the physical memory of the PPU 400 that is accessing these pages more frequently. In one embodiment, NVLink 410 supports address translation services, allowing PPU 400 to directly access the CPU's page table and providing PPU 400 with full access to the CPU's memory.

[0060] In one embodiment, the replication engine transfers data between multiple PPUs 400 or between a PPU 400 and a CPU. The replication engine can generate page faults for addresses that are not mapped to a page table. The memory partitioning unit 480 can then repair the page faults, mapping these addresses to the page table, after which the replication engine can perform the transfer. In conventional systems, memory is fixed (e.g., non-pageable) for multiple replication engine operations across multiple processors, significantly reducing available storage. In the event of a hardware page fault, addresses can be passed to the replication engine without concern for whether memory pages reside, and the replication process is transparent.

[0061] Data from memory 404 or other system memory can be fetched by memory partitioning unit 480 and stored in an on-chip L2 cache 460 shared among the various GPCs 450. As shown, each memory partitioning unit 480 includes a portion of the L2 cache associated with the corresponding memory 404. Low-level caches can then be implemented in various units within the GPC 450. For example, each processing unit within the GPC 450 can implement a Level 1 (L1) cache. The L1 cache is a private memory dedicated to a specific processing unit. The L2 cache 460 is coupled to memory interface 470 and XBar 470, and data from the L2 cache can be fetched and stored in each of the L1 caches for processing.

[0062] In one embodiment, the processing unit within each GPC 450 implements a SIMD (Single Instruction Multiple Data) architecture, where each thread in a thread group (e.g., a thread bundle) is configured to process a different data set based on the same set of instructions. All threads in the thread group execute the same instructions. In another embodiment, the processing unit implements a SIMT (Single Instruction Multiple Thread) architecture, where each thread in the thread group is configured to process a different data set based on the same set of instructions, but individual threads in the thread group are allowed to diverge during execution. In one embodiment, maintaining a program counter, call stack, and execution state for each thread bundle allows for concurrency between thread bundles and serial execution within a thread bundle when threads diverge. In another embodiment, maintaining a program counter, call stack, and execution state for each individual thread allows for equal concurrency among all threads within and between thread bundles. When maintaining an execution state for each individual thread, threads executing the same instructions can aggregate and execute in parallel for maximum efficiency.

[0063] Cooperative groups are a programming model for organizing groups of communicating threads. They allow developers to express the granularity at which threads are communicating, enabling richer and more efficient expressions of parallel decomposition. The Cooperative Startup API supports synchronization between blocks of threads executing parallel algorithms. Conventional programming models provide a single, simple construct for synchronizing cooperative threads: a barrier across all threads in a thread block (e.g., the `syncthreads()` function). However, programmers often prefer to define thread groups smaller than thread blocks in the form of a collective group-wide function interface, and synchronize within the defined group to allow for greater performance, design flexibility, and software reuse.

[0064] Collaboration groups enable programmers to explicitly define thread groups (as small as a single thread) at the sub-block and multi-block granularity, and perform collective operations such as synchronization on threads within the collaboration group. This programming model supports clean composition across software boundaries, allowing libraries and utility functions to be safely synchronized within their local contexts without having to make assumptions about aggregation. The collaboration group primitive allows for the implementation of new cooperative parallelism patterns, including producer-consumer parallelism, opportunistic parallelism, and global synchronization across the entire mesh of thread blocks.

[0065] Each processing unit comprises a large number (e.g., 128, etc.) of different processing cores (e.g., functional units), which may be fully pipelined, single-precision, double-precision, and / or mixed-precision, and include floating-point arithmetic logic units and integer arithmetic logic units. In one embodiment, the floating-point arithmetic logic unit implements the IEEE 754-2008 standard for floating-point arithmetic. In one embodiment, the core comprises 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.

[0066] Tensor cores are configured to perform matrix operations. Specifically, tensor cores are configured to perform deep learning matrix arithmetic, such as GEMM (matrix-matrix multiplication), for convolution operations during neural network training and inference. In one embodiment, each tensor core operates on a 4x4 matrix and performs matrix multiplication and accumulation operations, D = A × B + C, where A, B, C, and D are 4x4 matrices.

[0067] In one embodiment, matrix multiplication inputs A and B can be integer, fixed-point, or floating-point matrices, while accumulation matrices C and D can be integer, fixed-point, or floating-point matrices of equal or higher bit width. In one embodiment, the Tensor Core operates on 1-bit, 4-bit, or 8-bit integer input data using 32-bit integer accumulation. An 8-bit integer matrix multiplication requires 1024 operations and results in a full-precision product, which is then accumulated with other intermediate multiplications using 32-bit integer addition for an 8x8x16 matrix multiplication. In one embodiment, the Tensor Core operates on 16-bit floating-point input data using 32-bit floating-point accumulation. A 16-bit floating-point multiplication requires 64 operations and results in a full-precision product, which is then accumulated with other intermediate multiplications using 32-bit floating-point addition for a 4x4x4 matrix multiplication. In practice, the Tensor Core is used to perform operations on much larger two-dimensional or higher-dimensional matrices composed of these smaller elements. APIs such as the CUDA 9 C++ API expose specialized matrix loading, matrix multiplication and accumulation, and matrix storage operations for efficient use with Tensor Core from CUDA-C++ programs. At the CUDA level, the thread bundle-level interface takes a 16x16 matrix spanning all 32 threads of the thread bundle.

[0068] Each processing unit may also include M Special Function Units (SFUs) that perform special functions (e.g., attribute evaluation, reciprocal square root, etc.). In one embodiment, an SFU may include a tree traversal unit configured to traverse a hierarchical tree data structure. In one embodiment, an SFU may include a texture unit configured to perform texture map filtering operations. In one embodiment, a texture unit is configured to load texture maps (e.g., a 2D texture array) from memory 404 and sample these texture maps to produce sampled texture values ​​for use by a shader program executed by the processing unit. In one embodiment, the texture maps are stored in shared memory that may include or contain an L1 cache. The texture units use mip maps (e.g., texture maps with varying levels of detail) to implement texture operations such as filtering. In one embodiment, each processing unit includes two texture units.

[0069] Each processing unit also includes N Load Memory Units (LSUs) that implement load and store operations between shared memory and the register file. Each processing unit includes an interconnect network that connects each of the cores to the register file and connects the LSUs to the register file and the shared memory. In one embodiment, the interconnect network is a cross switch that can be configured to connect any of the cores to any register in the register file and connect the LSUs to memory locations in the register file and the shared memory.

[0070] Shared memory is an on-chip memory array that allows data storage and communication between processing units and between threads within a processing unit. In one embodiment, the shared memory includes 128KB of storage capacity and is located on the path from each of the processing units to memory partition unit 480. The shared memory can be used for caching reads and writes. One or more of the shared memory, L1 cache, L2 cache, and memory 404 serve as a backup cache.

[0071] Combining data caching and shared memory functionality into a single memory block provides optimal overall performance for both types of memory access. This capacity can be used as a cache by programs that do not use shared memory. For example, if shared memory is configured to use half its capacity, then texture and load / store operations can use the remaining capacity. Integration into shared memory allows it to function as a high-throughput conduit for streaming data, while providing high-bandwidth and low-latency access to frequently reused data.

[0072] When configured for general-purpose parallel computing, a simpler configuration can be used compared to graphics computing. Specifically, bypassing fixed-function graphics processing units (GPUs) creates a much simpler programming model. In this general-purpose parallel computing configuration, the work allocation unit 425 directly dispatches and assigns thread blocks to processing units within the GPC 450. Threads execute the same program using a unique thread ID during computation to ensure that each thread uses the executor program and the processing unit performing the computation, the shared memory for communication between threads, and the LSU (Local Subsystem for Memory) for reading and writing global memory via the shared memory and memory partitioning unit 480 to generate unique results. When configured for general-purpose parallel computing, processing units can also write commands, which the scheduler unit 420 can use to start new work on the processing unit.

[0073] Each of the PPUs 400 may include one or more processing cores and / or components thereof, such as a Tensor Core (TC), a Tensor Processing Unit (TPU), a Pixel Vision Core (PVC), a Ray Tracing (RT) Core, a Vision Processing Unit (VPU), a Graphics Processing Cluster (GPC), a Texture Processing Cluster (TPC), a Streaming Multiprocessor (SM), a Tree Traversal Unit (TTU), an Artificial Intelligence Accelerator (AIA), a Deep Learning Accelerator (DLA), an Arithmetic Logic Unit (ALU), an Application-Specific Integrated Circuit (ASIC), a Floating-Point Unit (FPU), Input / Output (I / O) Elements, Peripheral Component Interconnect (PCI) or Peripheral Component Interconnect High Speed ​​(PCIe) Elements and / or the like, and / or be configured to perform its functions.

[0074] The PPU 400 may be included in desktop computers, laptop computers, tablet computers, servers, supercomputers, smartphones (e.g., wireless, handheld devices), personal digital assistants (PDAs), digital cameras, vehicles, head-mounted displays, handheld electronic devices, and the like. In one embodiment, the PPU 400 is implemented on a single semiconductor substrate. In another embodiment, the PPU 400 is included in a system-on-a-chip (SoC) along with one or more other devices such as an additional PPU 400, memory 404, a reduced instruction set computer (RISC) CPU, a memory management unit (MMU), a digital-to-analog converter (DAC), and the like.

[0075] In one embodiment, PPU 400 may be included on a graphics card that includes one or more memory devices. The graphics card may be configured to interface with a PCIe slot on a desktop computer's motherboard. In yet another embodiment, PPU 400 may be an integrated graphics processing unit (iGPU) or a parallel processor included in the chipset of the motherboard. In yet another embodiment, PPU 400 may be implemented in reconfigurable hardware. In yet another embodiment, a portion of PPU 400 may be implemented in reconfigurable hardware.

[0076] Exemplary computing system

[0077] As developers expose and leverage more parallelism in applications such as artificial intelligence computing, systems with multiple GPUs and CPUs are being used across various industries. High-performance GPU-accelerated systems with tens to thousands of compute nodes are being deployed in data centers, research facilities, and supercomputers to solve increasingly complex problems. With the increasing number of processing devices within high-performance systems, communication and data transmission mechanisms need to be scaled to support the increased bandwidth.

[0078] Figure 5A According to one embodiment, the use Figure 4 A conceptual diagram of a processing system 500 implemented by a PPU 400. The exemplary system 500 can be configured to implement... Figure 2 The method 200 shown in the figure. The processing system 500 includes a CPU 530, a switch 510, multiple PPUs 400, and various memories 404.

[0079] The NVLink 410 provides a high-speed communication link between each of the PPUs 400. Although Figure 5BThe diagram illustrates a specific number of NVLink 410 and interconnect 402 connections, but the number of connections to each PPU 400 and CPU 530 can vary. Switch 510 forms an interface between interconnect 402 and CPU 530. PPU 400, memory 404, and NVLink 410 can reside on a single semiconductor platform to form a parallel processing module 525. In one embodiment, switch 510 supports two or more protocols to interface between various connections and / or links.

[0080] In another embodiment (not shown), NVLink 410 provides one or more high-speed communication links between each PPU 400 and CPU 530, and switch 510 forms an interface between interconnect 402 and each PPU 400. PPU 400, memory 404, and interconnect 402 may reside on a single semiconductor platform to form a parallel processing module 525. In yet another embodiment (not shown), interconnect 402 provides one or more communication links between each PPU 400 and CPU 530, and switch 510 uses NVLink 410 to form an interface between each PPU 400 to provide one or more high-speed communication links between PPUs 400. In another embodiment (not shown), NVLink 410 provides one or more high-speed communication links between PPU 400 and CPU 530 via switch 510. In yet another embodiment (not shown), interconnect 402 directly provides one or more communication links between each PPU 400. One or more of the NVLink 410 high-speed communication links can be implemented as physical NVLink interconnects or on-chip or die-on interconnects using the same protocol as the NVLink 410.

[0081] In the context of this specification, a single semiconductor platform can refer to a single semiconductor-based integrated circuit fabricated on a bare die or chip. It should be noted that the term "single semiconductor platform" can also refer to a multi-chip module with increased connectivity, simulating on-chip operation and representing a significant improvement over conventional bus implementations. Of course, the various circuits or devices can also be located individually within the semiconductor platform or in various combinations thereof, as desired by the user. Alternatively, the parallel processing module 525 can be implemented as a circuit board substrate, and each PPU 400 and / or memory 404 can be a packaged device. In one embodiment, the CPU 530, switch 510, and parallel processing module 525 reside on a single semiconductor platform.

[0082] In one embodiment, the signaling rate of each NVLink 410 is 20-25 gigabits per second, and each PPU400 includes six NVLink 410 interfaces (e.g., Figure 5A As shown, each PPU 400 includes five NVLink 410 interfaces. Each NVLink 410 provides a data transfer rate of 25 gigabits per second in each direction, and the six links provide 400 gigabits per second. NVLink 410 can be used as follows: Figure 5A It is used exclusively for PPU-to-PPU communication, or for a combination of PPU-to-PPU and PPU-to-CPU when the CPU 530 also includes one or more NVLink 410 interfaces.

[0083] In one embodiment, NVLink 410 allows direct load / store / atomic access from CPU 530 to memory 404 of each PPU 400. In one embodiment, NVLink 410 supports coherent operation, allowing data read from memory 404 to be stored in the cache hierarchy of CPU 530, reducing cache access latency of CPU 530. In one embodiment, NVLink 410 includes support for Address Translation Service (ATS), allowing PPU 400 to directly access page tables within CPU 530. One or more of NVLink 410 may also be configured to operate in low-power mode.

[0084] Figure 5B An exemplary system 565 is illustrated, in which various architectures and / or functions of various prior embodiments can be implemented. The exemplary system 565 can be configured to implement... Figure 2 Method 200 is shown in the figure.

[0085] As shown in the figure, a system 565 is provided, which includes at least one central processing unit 530 connected to a communication bus 575. The communication bus 575 may directly or indirectly couple one or more of the following devices: main memory 540, network interface 535, CPU 530, display device 545, input device 560, switch 510, and parallel processing system 525. The communication bus 575 may be implemented using any suitable protocol and may represent one or more links or buses, such as address bus, data bus, control bus, or combinations thereof. The communication bus 575 may include one or more bus or link types, such as Industry Standard Architecture (ISA) bus, Extended Industry Standard Architecture (EISA) bus, Video Electronics Standards Association (VESA) bus, Peripheral Component Interconnect (PCI) bus, Peripheral Component Interconnect High Speed ​​(PCIe) bus, HyperTransport, and / or another type of bus or link. In some embodiments, direct connections exist between components. As an example, CPU 530 may be directly connected to main memory 540. Furthermore, CPU 530 may be directly connected to parallel processing system 525. In cases where there is a direct or point-to-point connection between components, the communication bus 575 may include a PCIe link that implements the connection. In these examples, the PCI bus need not be included in the system 565.

[0086] Although using lines Figure 5B The different blocks are shown connected via a communication bus 575, but this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component such as a display device 545 can be considered an I / O component, such as an input device 560 (e.g., if the display is a touchscreen). As another example, the CPU 530 and / or the parallel processing system 525 may include memory (e.g., main memory 540 may represent storage devices other than the parallel processing system 525, the CPU 530, and / or other components). In other words, Figure 5B The term "computing device" is merely illustrative. No distinction is made between categories such as "workstation," "server," "laptop," "desktop," "tablet," "client device," "mobile device," "handheld device," "game console," "electronic control unit (ECU)," "virtual reality system," and / or other device or system types, as all of these are expected to fall under [the relevant category]. Figure 5B Within the scope of computing devices.

[0087] System 565 also includes main memory 540. Control logic (software) and data are stored in main memory 540, which can take the form of a variety of computer-readable media. Computer-readable media can be any available medium that can be accessed by system 565. Computer-readable media can include volatile and non-volatile media, as well as removable and non-removable media. For example and without limitation, computer-readable media can include computer storage media and communication media.

[0088] Computer storage media may include volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, main memory 540 may store computer-readable instructions such as an operating system (e.g., representing programs and / or program elements). Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage devices, magnetic cartridges, magnetic tape, disk storage devices or other magnetic storage devices, or any other medium that may be used to store desired information and that can be accessed by system 565. When used herein, computer storage media does not include the signal itself.

[0089] Computer storage media may contain computer-readable instructions, data structures, program modules, or other data types in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information transport medium. The term "modulated data signal" may refer to a signal whose characteristics are set or altered in a manner that encodes information into that signal. For example and without limitation, computer storage media may include wired media such as wired networks or direct wired connections, and wireless media such as sound, RF, infrared, and other wireless media. Any combination of the above should also be included within the scope of computer-readable media.

[0090] When executed, the computer program enables system 565 to perform various functions. CPU 530 may be configured to execute at least some of the computer-readable instructions to control one or more components of system 565 to perform one or more of the methods and / or processes described herein. Each of CPUs 530 may include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing numerous software threads simultaneously. Depending on the type of system 565 implemented, CPU 530 may include any type of processor and may include different types of processors (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of system 565, the processor may be an advanced RISC machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors such as math coprocessors, system 565 may include one or more CPUs 530.

[0091] In addition to or alternatively to CPU 530, parallel processing module 525 may be configured to execute at least some of the computer-readable instructions to control one or more components of system 565 to perform one or more of the methods and / or processes described herein. Parallel processing module 525 may be used by system 565 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, parallel processing module 525 may be used for general-purpose computing on a GPU (GPGPU). In embodiments, CPU 530 and / or parallel processing module 525 may execute any combination of the methods, processes, and / or portions thereof, discretely or jointly.

[0092] System 565 also includes input device 560, parallel processing system 525, and display device 545. Display device 545 may include a display (e.g., a monitor, touchscreen, television screen, head-up display (HUD), other display types, or combinations thereof), speakers, and / or other presentation components. Display device 545 may receive data from other components (e.g., parallel processing system 525, CPU 530, etc.) and output that data (e.g., images, video, sound, etc.).

[0093] Network interface 535 enables system 565 to be logically coupled to other devices, including input device 560, display device 545, and / or other components, some of which may be embedded (e.g., integrated into) system 565. Illustrative input device 560 includes microphone, mouse, keyboard, joystick, gamepad, game controller, satellite dish, scanner, printer, wireless device, etc. Input device 560 can provide a natural user interface (NUI) that processes user-generated air gestures, voice, or other physiological input. In some instances, input can be transmitted to appropriate network elements for further processing. NUI can implement voice recognition, stylus recognition, facial recognition, biometric recognition, on-screen and adjacent-screen gesture recognition, air gestures, head-eye tracking, and touch recognition associated with the display of system 565 (described in more detail below). System 565 may include depth cameras for gesture detection and recognition, such as stereo camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations thereof. In addition, system 565 may include an accelerometer or gyroscope that allows motion detection (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope may be used by system 565 to render immersive augmented reality or virtual reality.

[0094] Furthermore, system 565 can be coupled to a network (e.g., a telecommunications network, a local area network (LAN), a wireless network, a wide area network (WAN) such as the Internet, a peer-to-peer network, a cable network, etc.) via network interface 535 for communication purposes. System 565 can be included in a distributed network and / or cloud computing environment.

[0095] Network interface 535 may include one or more receivers, transmitters, and / or transceivers, enabling system 565 to communicate with other computing devices via electronic communication networks, including wired and / or wireless communications. Network interface 535 may be implemented as a network interface controller (NIC) including one or more data processing units (DPUs) for performing operations such as (e.g., but not limited to) packet parsing and accelerating network processing and communication. Network interface 535 may include components and functions that allow communication over any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or InfiniBand), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.

[0096] System 565 may also include an auxiliary storage device (not shown). The auxiliary storage device includes, for example, a hard disk drive and / or a removable storage drive representing a floppy disk drive, magnetic tape drive, compact disc drive, digital versatile disc (DVD) drive, recording device, Universal Serial Bus (USB) flash memory. The removable storage drive reads from and / or writes to the removable storage unit in a well-known manner. System 565 may also include a hard-wired power supply, a battery power supply, or a combination thereof (not shown). This power supply can supply power to System 565 to enable the components of System 565 to operate.

[0097] Each of the aforementioned modules and / or devices may even reside on a single semiconductor platform to form system 565. Alternatively, various different modules may be placed individually or located in various combinations of semiconductor platforms as desired by the user. Although various different embodiments have been described above, it should be understood that they are given by way of example only and without limitation. Therefore, the breadth and scope of preferred embodiments should not be limited to any of the exemplary embodiments described above, but should be defined only by the following claims and their equivalents.

[0098] Example network environment

[0099] A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network-attached storage devices (NAS), other back-end devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may be... Figure 5A Processing system 500 and / or Figure 5B Implemented on one or more instances of the exemplary system 565, for example, each device may include similar components, features and / or functions of the processing system 500 and / or the exemplary system 565.

[0100] Components of a network environment can communicate with each other via a network, which can be wired, wireless, or a combination of both. A network can include multiple networks or a network of networks. For example, a network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks—such as the Internet, and / or the Public Switched Telephone Network (PSTN), and / or one or more private networks. In cases where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (along with other components) can provide wireless connectivity.

[0101] A compatible network environment may include one or more peer-to-peer network environments—in which case the server may not be included in the network environment—and one or more client-server network environments—in which case one or more servers may be included in the network environment. In a peer-to-peer network environment, the functionality described herein regarding the server can be implemented on any number of client devices.

[0102] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework of one or more applications supporting a software layer and / or an application layer. The software or application may include web-based service software or applications, respectively. In embodiments, one or more client devices may use the web-based service software or application (e.g., by accessing the service software and / or application via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a free and open-source software web application framework type that can be used for large-scale data processing (e.g., "big data").

[0103] A cloud-based network environment can provide cloud computing and / or cloud storage to implement the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these different functions can be distributed across multiple locations from a central or core server (e.g., a central or core server in one or more data centers, which may be distributed across states, regions, countries, globally, etc.). If the connection to a user (e.g., a client device) is relatively close to an edge server, then the core server can assign at least a portion of the functions to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0104] Client devices may include Figure 5A Example processing system 500 and / or Figure 5BAt least some of the components, features, and functions of the exemplary system 565. For example and without limitation, the client device may be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smartwatch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, video camera, surveillance equipment or system, vehicle, ship, spacecraft, virtual machine, drone, robot, handheld communication device, hospital equipment, gaming equipment or system, entertainment system, vehicle computer system, embedded system controller, remote control, appliance, consumer electronics device, workstation, edge device, any combination of these defined devices, or any other suitable device.

[0105] Machine Learning

[0106] Deep neural networks (DNNs) developed on processors such as the PPU 400 have been used in a wide variety of use cases, from self-driving cars to faster drug development, from automatic image captioning in online image databases to intelligent real-time language translation in video chat applications. Deep learning is a technique that models the neural learning process of the human brain, which continuously learns, becomes smarter, and delivers more accurate results faster over time. Just as a child is initially taught by adults to correctly identify and classify various shapes, eventually becoming able to identify shapes without any guidance, a deep learning or neural learning system needs to be trained in object recognition and classification so that it becomes smarter and more efficient at identifying basic objects, occluded objects, and so on, while also attaching context to objects.

[0107] At its simplest level, neurons in the human brain receive various inputs, assigning a level of importance to each of these inputs, and the output is passed to other neurons to make a response. Artificial neurons, or perceptrons, are the most basic model of neural networks. In one example, a perceptron can receive one or more inputs representing various features of objects that the perceptron is being trained to recognize and classify, and each of these features is assigned a weight based on its importance in defining the shape of the object.

[0108] Deep neural network (DNN) models consist of multiple layers of numerous connected nodes (e.g., perceptrons, Boltzmann machines, radial basis functions, convolutional layers, etc.) that can be trained on massive amounts of input data to solve complex problems quickly and with high accuracy. In one example, the first layer of a DNN model breaks down an input image of a car into different segments and searches for basic patterns such as lines and angles. The second layer assembles these lines to find higher-level patterns such as wheels, windshields, and mirrors. The next layer identifies the type of vehicle, and the final few layers generate labels for the input image that identify the model of a specific car brand.

[0109] Once trained, a DNN can be deployed and used to identify and classify objects or patterns in a process called inference. Examples of inference (the process by which a DNN extracts useful information from a given input) include identifying handwritten digits on a check deposited into an ATM, identifying images of friends in a photograph, delivering movie recommendations to over 50 million users, identifying and classifying different types of cars, pedestrians, and road hazards in self-driving cars, or translating human language in real time.

[0110] During training, data flows through the DNN in the forward propagation phase until a prediction is produced indicating the label corresponding to the input. If the neural network does not correctly label the input, the error between the correct label and the predicted label is analyzed, and the weights are adjusted for each feature during the backpropagation phase until the DNN correctly labels the input as well as other inputs in the training dataset. Training complex neural networks requires significant parallel computing power, including floating-point multiplication and addition supported by a PPU400. Inference is less computationally intensive than training and is a latency-sensitive process where the trained neural network is applied to new inputs it has not seen before for tasks such as image classification, sentiment detection, label recommendation, language recognition and translation, and typically infers new information.

[0111] Neural networks rely heavily on matrix operations, and complex, multi-layered networks require massive amounts of floating-point performance and bandwidth for both efficiency and speed. Leveraging thousands of processing cores optimized for matrix operations and delivering tens to hundreds of TFLOPS of performance, the PPU 400 is a computing platform capable of providing the performance required for deep neural network-based artificial intelligence and machine learning applications.

[0112] Furthermore, images generated using one or more of the techniques disclosed herein can be used to train, test, or certify DNNs for recognizing real-world objects and environments. Such images can include driveways, factories, buildings, urban environments, rural environments, humans, animals, and any other physical objects or scenes of real-world environments. Such images can be used to train, test, or certify DNNs used in machines or robots to manipulate, process, or modify real-world physical objects. Additionally, such images can be used to train, test, or certify DNNs used in autonomous vehicles to navigate and move vehicles in the real world. Furthermore, images generated using one or more of the techniques disclosed herein can be used to communicate information to users of such machines, robots, and vehicles.

[0113] Figure 5C Components of an example system 555, which can be used to train and utilize machine learning according to at least one embodiment, are illustrated. As will be discussed, various components can be provided by a single computing system or various combinations of computing devices and resources, which may be under the control of a single entity or multiple entities. Furthermore, aspects may be triggered, initiated, or requested by different entities. In at least one embodiment, the training of the neural network may be guided by a vendor associated with vendor environment 506, while in at least one embodiment, training may be requested by a customer or other user who can access the vendor environment through client device 502 or other such resources. In at least one embodiment, training data (or data to be analyzed by the trained neural network) may be provided by a vendor, user, or third-party content provider 524. In at least one embodiment, client device 502 may be, for example, a vehicle or object to be navigated on behalf of a user, who can submit requests and / or receive instructions that aid in device navigation.

[0114] In at least one embodiment, a request can be submitted via at least one network 504 for receipt by a vendor environment 506. In at least one embodiment, the client device can be any suitable electronic and / or computing device that enables a user to generate and send such requests, such as, but not limited to, desktop computers, laptop computers, computer servers, smartphones, tablets, game consoles (portable or otherwise), computer processors, computing logic, and set-top boxes. One or more networks 504 can include any suitable network for transmitting requests or other such data, such as the Internet, intranet, Ethernet, cellular network, local area network (LAN), wide area network (WAN), personal area network (PAN), self-organizing network providing direct wireless connectivity between peers, etc.

[0115] In at least one embodiment, a request may be received at interface layer 508, which in this example may forward data to training and inference manager 532. Training and inference manager 532 may be a system or service including hardware and software for managing services and requests corresponding to data or content. In at least one embodiment, training and inference manager 532 may receive a request to train a neural network and may provide data for the request to training module 512. In at least one embodiment, if the request is not specified, training module 512 may select an appropriate model or neural network to use and may train the model using the associated training data. In at least one embodiment, training data may be a batch of data stored in training data repository 514, received from client device 502, or obtained from third-party vendor 524. In at least one embodiment, training module 512 may be responsible for training the data. The neural network may be any suitable network, such as a recurrent neural network (RNN) or convolutional neural network (CNN). Once the neural network is trained and successfully evaluated, the trained neural network may be stored in, for example, model repository 516, which may store different models or networks for users, applications, or services, etc. In at least one embodiment, there may be multiple models for a single application or entity, which can be utilized based on multiple different factors.

[0116] In at least one embodiment, at a subsequent point in time, a request for content (e.g., path determination) or data that is at least partially determined or influenced by a trained neural network can be received from client device 502 (or another such device). This request may include, for example, input data to be processed using the neural network to obtain one or more inference or other output values, classifications, or predictions. Alternatively, in at least one embodiment, the input data may be received by interface layer 508 and directed to inference module 518, although different systems or services may also be used. In at least one embodiment, if not already locally stored in inference module 518, inference module 518 may obtain a suitably trained network, such as a trained deep neural network (DNN) as discussed herein, from model repository 516. Inference module 518 may provide data as input to the trained network, which may then generate one or more inferences as outputs. This may, for example, include the classification of instances of input data. In at least one embodiment, the inference may then be transmitted to client device 502 for display to a user or for other communication with the user. In at least one embodiment, user context data may also be stored in a user context data repository 522, which may include data about the user that can be used as network input to generate inference or determine data returned to the user after obtaining an instance. In at least one embodiment, relevant data, including at least some of the input or inference data, may also be stored in a local database 534 for processing future requests. In at least one embodiment, the user may use account information or other information to access resources or functions of the vendor environment. In at least one embodiment, user data may also be collected and used to further train the model, if permitted and available, to provide more accurate inference for future requests. In at least one embodiment, requests to a machine learning application 526 executed on a client device 502 may be received via a user interface, and the results may be displayed via the same interface. The client device may include resources such as a processor 528 and a memory 562 for generating requests and processing results or responses, and at least one data storage element 552 for storing data for the machine learning application 526.

[0117] In at least one embodiment, processor 528 (or the processor of training module 512 or inference module 518) will be a central processing unit (CPU). However, as mentioned above, resources in such an environment can utilize GPUs to process data for at least some types of requests. GPUs, such as the PPU 400, have thousands of cores and are designed to handle large amounts of parallel workloads, thus becoming popular in deep learning for training neural networks and generating predictions. While using GPUs for offline building allows for faster training of larger, more complex models, offline prediction generation means that request-time input features cannot be used, or predictions must be generated for all features and stored in a lookup table for real-time service requests. If the deep learning framework supports CPU mode and the model is small and simple enough that the feedforward can be performed on the CPU with reasonable latency, then a service on a CPU instance can host the model. In this case, training can be done offline on the GPU and inference can be performed in real-time on the CPU. If the CPU approach is not feasible, the service can run on a GPU instance. However, due to the different performance and cost characteristics of GPUs compared to CPUs, running a service that offloads runtime algorithms to the GPU may require it to be designed differently from a CPU-based service.

[0118] In at least one embodiment, video data can be provided from client device 502 for enhancement in vendor environment 506. In at least one embodiment, the video data can be processed for enhancement on client device 502. In at least one embodiment, the video data can be streamed from third-party content provider 524 and enhanced by third-party content provider 524, vendor environment 506, or client device 502. In at least one embodiment, video data can be provided from client device 502 for use as training data in vendor environment 506.

[0119] In at least one embodiment, supervised and / or unsupervised training may be performed by client device 502 and / or vendor environment 506. In at least one embodiment, a set of training data 514 (e.g., classified or labeled data) is provided as input for use as training data. In at least one embodiment, the training data may include instances of at least one type of object for which the neural network is to be trained, and information identifying that object type. In at least one embodiment, the training data may include a set of images, each image including a representation of an object of a type, wherein each image also includes, or is associated with, tags, metadata, classification, or other information identifying or identifying the type of object represented in the corresponding image. Various other types of data may also be used as training data, which may include text data, audio data, video data, and so on. In at least one embodiment, training data 514 is provided as training input to training module 512. In at least one embodiment, training module 512 may be a system or service including hardware and software, such as one or more computing devices executing a training application for training a neural network (or other model or algorithm, etc.). In at least one embodiment, training module 512 receives instructions or requests indicating the type of model to be used for training. In at least one embodiment, the model can be any suitable statistical model, network, or algorithm useful for such a purpose, which may include artificial neural networks, deep learning algorithms, learning classifiers, Bayesian networks, etc. In at least one embodiment, training module 512 may select an initial model or other untrained models from an appropriate repository and train the model using training data 514 to generate a trained model (e.g., a trained deep neural network) that can be used to classify similar types of data or generate other such inference. In at least one embodiment where training data is not used, an initial model can still be selected to train on the input data of each training module 512.

[0120] In at least one embodiment, the model can be trained in several different ways, which may depend in part on the type of model chosen. In at least one embodiment, a training dataset can be provided to a machine learning algorithm, wherein the model is a model artifact created through a training process. In at least one embodiment, each instance of the training data contains the correct answer (e.g., classification) that may be referred to as the target or target attribute. In at least one embodiment, the learning algorithm finds patterns in the training data that map input data attributes to the target—the answer to be predicted—and the machine learning model is the output that captures these patterns. In at least one embodiment, the machine learning model can then be used to obtain predictions for new data without a specified target.

[0121] In at least one embodiment, the training and inference manager 532 may select from a set of machine learning models, including binary classification, multi-class classification, generative, and regression models. In at least one embodiment, the type of model to be used may depend at least in part on the type of target to be predicted.

[0122] In one embodiment, PPU 400 includes a graphics processing unit (GPU). PPU 400 is configured to receive commands specifying a shader program for processing graphics data. Graphics data can be defined as a set of primitives such as points, lines, triangles, quadrilaterals, triangle strips, etc. Typically, a primitive includes data specifying the number of vertices used for that primitive (e.g., in a model-space coordinate system) and attributes associated with each vertex of that primitive. PPU 400 can be configured to process graphics primitives to generate framebuffers (e.g., pixel data for each of the pixels in a display).

[0123] The application writes model data (e.g., attributes and vertex sets) for the scene into memory such as system memory or memory 404. The model data defines each of the objects that may be visible on the display. The application then makes API calls to the driver kernel, requesting that the model data be rendered and displayed. The driver kernel reads the model data and writes commands to one or more streams to perform operations that process the model data. These commands may reference different shader programs to be implemented on processing units within the PPU 400, including one or more vertex shaders, shell shaders, domain shaders, geometry shaders, and pixel shaders. For example, one or more of the processing units may be configured to execute a vertex shader program that processes a number of vertices defined by the model data. In one embodiment, these different processing units may be configured to execute different shader programs concurrently. For example, a first subset of processing units may be configured to execute a vertex shader program, while a second subset of processing units may be configured to execute a pixel shader program. The first subset of processing units processes the vertex data to produce processed vertex data and writes the processed vertex data into L2 cache 460 and / or memory 404. After the processed vertex data is rasterized (e.g., transformed from 3D data to 2D data in screen space) to produce fragment data, a second subset of processing units executes pixel shaders to produce processed fragment data, which is then mixed with other processed fragment data and written to the frame buffer in memory 404. Vertex shader and pixel shader programs can execute concurrently, pipelinedly processing different data from the same scene until all model data for that scene has been rendered to the frame buffer. The contents of the frame buffer are then transferred to the display controller for display on the display device.

[0124] Images generated using one or more of the techniques disclosed herein can be displayed on a monitor or other display device. In some embodiments, the display device may be directly coupled to the system or processor that generates or renders the image. In other embodiments, the display device may be indirectly coupled to the system or processor, for example, via a network. Examples of such networks include the Internet, mobile telecommunications networks, Wi-Fi networks, and any other wired and / or wireless networking systems. When the display device is indirectly coupled, images generated by the system or processor can be streamed to the display device over the network. Such streaming allows, for example, video games or other applications that render images to execute on servers, data centers, or cloud-based computing environments, and the rendered images are transmitted and displayed on one or more user devices (e.g., computers, video game consoles, smartphones, other mobile devices, etc.) physically separate from the server or data center. Therefore, the techniques disclosed herein can be applied to enhance streamed images and services that stream images, such as NVIDIA GeForce Now (GFN), Google Stadia, etc.

[0125] Example Streaming System

[0126] Figure 6 This is a schematic diagram of an example system 605 of a streaming system according to some embodiments of the present disclosure. Figure 6 Includes server 603 (which may include with Figure 5A Example processing system 500 and / or Figure 5B (Similar components, features and / or functions to exemplary system 565), client 604 (which may include similar ... Figure 5A Example processing system 500 and / or Figure 5B The exemplary system 565 has similar components, features, and / or functions to the network 606 (which may be similar to the network described herein). In some embodiments of this disclosure, system 605 may be implemented.

[0127] In one embodiment, the streaming system 605 is a game streaming system, and the server 604 is a game server. In system 605, for a game session, the client device 604 can simply receive input data in response to input from the input device 626, send the input data to the server 603, receive encoded display data from the server 603, and display the display data on the display 624. In this way, computationally intensive computation and processing are offloaded to the server 603 (e.g., rendering of the game session's graphics output, especially ray or path tracing, is performed by the GPU 615 of the server 603). In other words, the game session is streamed from the server 603 to the client device 604, thereby reducing the demands on the client device 604 for graphics processing and rendering.

[0128] For example, regarding the instantiation of a game session, client device 604 can display frames of the game session on display 624 based on display data received from server 603. Client device 604 can receive input from one of input devices 626 and generate input data in response. Client device 604 can send the input data to server 603 via communication interface 621 and over network 606 (e.g., the Internet), and server 603 can receive the input data via communication interface 618. CPU 608 can receive the input data, process the input data, and send the data to GPU 615, which causes GPU 615 to generate a rendering of the game session. For example, the input data can represent the movement of a user character in the game, such as firing a weapon, reloading, passing a ball, turning a vehicle, etc. Rendering component 612 can render the game session (e.g., representing the result of the input data), and rendering capture component 614 can capture the rendering of the game session as display data (e.g., image data as frames of the captured game session rendering). The rendering of a game session may include lighting and / or shadow effects computed using one or more parallel processing units of server 603 (e.g., a GPU, which may further employ one or more dedicated hardware accelerators or processing cores to perform ray or path tracing techniques). Encoder 616 can then encode the display data to generate encoded display data, which can be sent to client device 604 via communication interface 618 through network 606. Client device 604 can receive the encoded display data via communication interface 621, and decoder 622 can decode the encoded display data to generate display data. Client device 604 can then display the display data via display 624.

[0129] Server 603 may include video system 100, which receives one or more source video streams and transmits them to one or more client devices 604 via communication interface 618 and one or more networks 606. One or more CPUs 608 and / or one or more GPUs 615 may receive characterization data 122 and process alarms. One or more CPUs 608 and / or one or more GPUs 615 may generate dashboard components for display based on characterization data 122 and / or system workload.

[0130] It should be noted that the techniques described herein can be contained in executable instructions stored in a computer-readable medium for use by or in conjunction with a processor-based instruction execution machine, system, apparatus, or device. Those skilled in the art will appreciate that, for some embodiments, various different types of computer-readable media may be included for storing data. When used herein, “computer-readable medium” includes one or more of any suitable media for storing executable instructions of a computer program, such that an instruction execution machine, system, apparatus, or device can read (or retrieve) the instructions from the computer-readable medium and execute those instructions to implement the described embodiments. Suitable storage formats include one or more of electronic, magnetic, optical, and electromagnetic formats. A non-exhaustive list of conventional exemplary computer-readable media includes: portable computer disks; random access memory (RAM); read-only memory (ROM); erasable programmable read-only memory (EPROM); flash memory devices; and optical storage devices, including portable compact discs (CDs), portable digital video discs (DVDs), and the like.

[0131] It should be understood that the arrangement of components shown in the accompanying drawings is for illustrative purposes, and other arrangements are possible. For example, one or more of the elements described herein may be implemented wholly or partially as electronic hardware components. Other elements may be implemented in software, hardware, or a combination of software and hardware. Moreover, some or all of these other elements may be combined, some may be omitted entirely, and additional components may be added while still achieving the functionality described herein. Therefore, the subject matter described herein can be implemented in many different variations, and all such variations are contemplated to be within the scope of the claims.

[0132] To facilitate understanding of the topics described herein, many aspects are described in sequence of actions. Those skilled in the art will recognize that various actions can be performed by dedicated circuitry or circuit systems, by program instructions executed by one or more processors, or by a combination of both. The description of any sequence of actions herein is not intended to imply that a particular order in which the actions described for execution must be followed. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by the context.

[0133] In the context of describing the subject matter (especially in the context of the claims below), the use of the terms “a,” “an,” “this,” and similar designations should be interpreted to cover both the singular and plural, unless otherwise specified herein or obviously contradicted by the context. The use of the term “at least one” (e.g., at least one of A and B) followed by a list of one or more items should be interpreted to mean one item selected from the listed items (A or B), or any combination of two or more of the listed items (A and B), unless otherwise specified herein or obviously contradicted by the context. Furthermore, the foregoing description is for illustrative purposes only and not for limiting purposes, as the scope of protection sought is defined by the claims set forth thereafter with their equivalents. The use of any and all example or exemplary language provided herein (e.g., “such as”) is intended merely to better illustrate the subject matter and does not constitute a limitation on the scope of the subject matter, unless otherwise stated. The use of “based on,” and other similar phrases indicating conditions leading to the result, in both the claims and the written description, is not intended to exclude any other conditions leading to that result. The language in the description should not be interpreted as indicating that any unclaimed element is essential for the implementation of the claimed invention.

Claims

1. A computer-implemented method, comprising: The processor initiates the transmission of at least one captured video stream to multiple remote clients; At least partially simultaneously with the transmission, the processor processes the at least one captured video stream to generate characterization data corresponding to the scene depicted in the at least one captured video stream; Monitor the performance of the processing and the bandwidth consumed for the transmission to generate workload data associated with the system workload; In response to determining that the system workload triggers a policy-based action, processing parameters are dynamically adjusted to generate the characterization data, the processing parameters controlling at least one function executed by the processor, wherein the at least one function includes adjustment of at least one of computational accuracy, tracking distance, or classification priority; as well as Dynamically adjust at least one of the frame rate, bit rate, or resolution of the at least one captured video stream being transmitted.

2. The computer-implemented method of claim 1, wherein the transmission is performed using a real-time streaming protocol.

3. The computer-implemented method of claim 1, wherein the processing is implemented by a neural network, the neural network being configurable to perform at least one of object detection, object classification, object tracking, segmentation, pose detection, or object recognition.

4. The computer-implemented method of claim 3, wherein the representation data includes at least one of bounding box, object trajectory, object label, object count, boundary intersection, or intersection highlighting.

5. The computer-implemented method of claim 1, wherein the computational precision includes the resolution of the feature map processed during inference.

6. The computer-implemented method of claim 1, wherein the characterization data includes selected video clips based on a probability associated with a lower level of confidence than the characterization data.

7. The computer-implemented method of claim 1, wherein one or more of the memory capacity, memory bandwidth, processor capacity, or network bandwidth contribute to the system workload.

8. The computer-implemented method of claim 1, wherein one or more environmental conditions within the system result in a system workload.

9. The computer-implemented method of claim 6, further comprising: An alert is generated in response to the selection of the video clip.

10. The computer-implemented method of claim 1, wherein the streaming parameters are dynamically adjusted in response to determining that the workload triggers a policy-based response.

11. The computer-implemented method of claim 1, wherein the processing parameter controls at least one of the frame sampling rate, the number of the at least one captured video streams being processed, or the video clip generation frequency.

12. The computer-implemented method of claim 1, wherein the system workload triggers the policy-based action when a change in a plurality of remote clients, including the plurality of remote clients, is detected.

13. The computer-implemented method of claim 1, wherein the system workload triggers the policy-based action when a change in a plurality of captured video streams, including the at least one captured video stream, is detected.

14. The computer-implemented method of claim 1, wherein at least one of the steps of initiating transmission, processing, or dynamically adjusting is performed on a server or in a data center to generate an image, and the image is streamed to a user device.

15. The computer-implemented method of claim 1, wherein at least one of the steps of inducing transmission, the processing, or the dynamic adjustment is performed within a cloud computing environment.

16. The computer-implemented method of claim 1, wherein at least one of the steps of inducing transmission, the processing, or the dynamically adjusting is performed for training, testing, or validating at least one of the neural networks.

17. The computer-implemented method of claim 16, wherein the neural network comprises a neural network used in at least one of a machine, robot, or autonomous vehicle.

18. A system comprising: A memory that stores at least one captured video stream; as well as A processor connected to the memory, wherein the processor is configured to: This causes the at least one captured video stream to be transmitted to multiple remote clients; At least partially simultaneously with the transmission, the at least one captured video stream is processed to generate characterization data corresponding to the scene depicted in the at least one captured video stream; Monitor the performance of the processing and the bandwidth consumed for the transmission to generate workload data associated with the system workload; In response to determining that the system workload triggers a policy-based action, processing parameters are dynamically adjusted to generate the characterization data, the processing parameters controlling at least one function executed by the processor, wherein the at least one function includes adjustment of at least one of computational accuracy, tracking distance, or classification priority; as well as Dynamically adjust at least one of the frame rate, bit rate, or resolution of the at least one captured video stream being transmitted.

19. A non-transitory computer-readable medium storing computer instructions for throttling video processing functions, the computer instructions, when executed by one or more processors, causing the one or more processors to perform the following steps: This causes at least one captured video stream to be transmitted to multiple remote clients; At least partially simultaneously with the transmission, the one or more processors process the at least one captured video stream to generate characterization data corresponding to the scene depicted in the at least one captured video stream; Monitor the performance of the processing and the bandwidth consumed for the transmission to generate workload data associated with the system workload; In response to determining that the system workload triggers a policy-based action, processing parameters are dynamically adjusted to generate the characterization data, the processing parameters controlling at least one function executed by the one or more processors, wherein the at least one function includes adjustment of at least one of computational accuracy, tracking distance, or classification priority; as well as Dynamically adjust at least one of the frame rate, bit rate, or resolution of the at least one captured video stream being transmitted.

20. The non-transitory computer-readable medium of claim 19, wherein the streaming parameters are dynamically adjusted in response to determining that the workload triggers a policy-based response.

Citation Information

Patent Citations

  • Video processing method and device, electronic device, and storage medium

    CN109587560A

  • Adaptive video transmission configuration method and system

    CN113242469A