Serverless GPU Function Orchestration Across Multi-Cluster Environments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Configuring a service to execute code on a GPU can be difficult due to infrastructure, scalability, and security concerns, especially in a serverless architecture.

Innovation Solution

A serverless architecture that allows for the execution of GPU cloud functions without requiring users to manage infrastructure or services, utilizing a central controller and agents to deploy workers across clusters, managing queues for execution requests, and ensuring secure communication.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a service is configured to execute code on GPU for remote clients, then GPU computing efficiency is improved, but infrastructure complexity and security configuration become difficult

Engineering Contradiction:
ImproveGPU computing efficiencyVSAvoidinfrastructure configuration complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces a serverless computing platform as an intermediary layer between remote clients and GPU resources. This platform includes a controller that manages worker deployment, queue management, and resource allocation, abstracting away the complex infrastructure configuration from users while maintaining efficient GPU execution.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements automated service configuration where the serverless platform automatically provisions workers, manages queues, and configures security policies without requiring manual infrastructure setup by users. The platform self-manages the complexity of GPU resource orchestration while providing simple access to clients.

Inventive Principle:
Principle #25Self-service

2Productivity

If a service is configured to execute code on GPU for remote clients, then GPU computing efficiency is improved, but security concerns increase

Engineering Contradiction:
ImproveGPU computing efficiencyVSAvoidsecurity configuration
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The serverless platform acts as a secure intermediary between clients and GPU workers. The controller implements authentication, authorization, and security policies that protect GPU resources while enabling efficient client access. This intermediary layer handles security concerns centrally without compromising GPU execution efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If workers are deployed across multiple cluster environments, then scalability and adaptability are improved, but system complexity increases

Engineering Contradiction:
Improvemulti-environment deployment capabilityVSAvoidsystem management complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal serverless platform that can deploy and manage workers across diverse cluster environments (cloud providers, data centers, edge devices). The controller implements environment-agnostic worker deployment, queue management, and resource orchestration that works uniformly across different infrastructure types, enabling multi-environment adaptability without proportionally increasing management complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Productivity

If workers are deployed across multiple cluster environments, then scalability is improved, but management complexity increases

Engineering Contradiction:
ImprovescalabilityVSAvoidworker management complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The serverless platform implements self-service automation for worker deployment and management across multiple clusters. The controller automatically provisions workers in appropriate environments based on queue demands, manages lifecycle events, and coordinates resources without requiring manual intervention. This automation enables scalability while keeping management complexity contained within the platform itself.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250370771A1Data center resource orchestration using serverless application programming interfaces
Publication Date: 2025.12.04 NVIDIA CORP
  • US20250370771A1 patent drawing
  • US20250370771A1 patent drawing
  • US20250370771A1 patent drawing

AI summary

Disclosed are systems and techniques for a cloud function worker for executing code using graphics processing units (GPUs) in a serverless architecture. The techniques include receiving, at a cloud function worker, a cloud function execution request from a cloud function queue of a cloud function controller. The techniques include identifying, based on the cloud function execution request, a first artificial intelligence (AI) model of a plurality of AI models of the cloud function controller. The techniques include generating a cloud function execution result of the cloud function execution request using the first AI model and at least a graphics processing unit (GPU) of a cluster environment hosting the cloud function worker. The techniques include causing the cloud function execution result to be transmitted to the cloud function controller.