Tensor API Segmentation for CPU-PPU Resource Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Performing operations using tensors often requires significant memory, time, or computing resources, which can be inefficient.

Innovation Solution

An application programming interface (API) is developed to perform tensor operations efficiently by utilizing a computing system with a central processing unit (CPU) and a parallel processing unit (PPU), such as a graphics processing unit (GPU), and optimizing memory allocation and compiler options.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional methods are used to perform tensor operations, then the operations can be executed, but significant memory, time, and computing resources are required

Engineering Contradiction:
Improvetensor operation efficiencyVSAvoidcomputing resources
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent segments tensor operations into multiple stages: a host processor executes a first stage of operations on a first tensor, then transfers control to a co-processor for a second stage on a second tensor. This segmentation allows each processor to specialize in specific computation tasks, improving overall efficiency and reducing resource consumption compared to conventional single-processor approaches

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a co-processor as an intermediary between the host processor and the tensor computation system. The co-processor receives control signals from the host processor and executes specialized tensor operations, acting as a mediator that offloads computational burden and optimizes resource utilization without requiring direct manipulation of all tensor data by the host processor

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of time

If tensors are processed using traditional computing systems, then computations can be performed, but significant memory and time resources are consumed

Engineering Contradiction:
Improvecomputation timeVSAvoidmemory resources
Core Design Contradiction:
Loss of timeVSQuantity of substance

Solution Approach 1:

The computation process is divided into temporal stages executed by different processors. The host processor handles initial setup and control, while the co-processor handles intensive tensor operations. This temporal segmentation reduces the time each processor needs to allocate to tensor operations, decreasing overall computation time while reducing memory pressure on the host system

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a control signal mechanism that copies computational control logic from the host processor to the co-processor. This allows the co-processor to independently execute tensor operations without continuously requiring host processor involvement, reducing communication overhead and memory resource requirements while maintaining computation accuracy

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250156984A1Application programming interface to perform tensor operations
Publication Date: 2025.05.15 NVIDIA CORP
  • US20250156984A1 patent drawing
  • US20250156984A1 patent drawing
  • US20250156984A1 patent drawing

AI summary

Apparatuses, systems, and techniques to perform one or more operations using a tensor. In at least one embodiment, one or more circuits are to perform an application programming interface (API) to perform one or more operations using one or more tensors based on at least one or more indications of the plurality of operations by the API, one or more compiler options indicated by the API, and/or one of one or more indications of an amount of storage to be used to perform the one or more operations.