Foundation Model Deployment via AI Capacity Profiling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing FMaaS platforms struggle to interpret non-standardized, text-based, and/or data-based prompts for AI model deployment on resource-constrained edge devices, often requiring expert intervention or discarding such requests, and the limited computing resources of edge devices restrict model deployment options.

Innovation Solution

A system using generative AI through a pre-trained large language model to translate text-based client requirements into interpretable FMaaS requests, identifying resource-optimal AI models by generating model and data descriptions, and performing AI task capacity profiling to select compatible and efficient model variants.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If text-based service requests are directly processed by existing FMaaS platforms, then expert intervention is required and requests may be discarded, but this increases operational complexity and reduces automation

Engineering Contradiction:
Improveautomated model selectionVSAvoidplatform interpretation complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

A generative AI translation layer is introduced as an intermediary between text-based service requests and the FMaaS platform. This translation layer converts natural language requests into standardized model selection parameters, enabling automated processing without expert intervention while maintaining platform compatibility

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system enables self-service by allowing the generative AI model to autonomously interpret text-based requests and automatically select appropriate AI models based on the translated parameters, eliminating the need for expert human intervention in the model selection process

Inventive Principle:
Principle #25Self-service

2Reliability

If complex foundation models are deployed on edge devices, then AI task performance is improved, but resource consumption exceeds device capacity

Engineering Contradiction:
ImproveAI task performanceVSAvoidedge device resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system changes the parameters of AI models by generating multiple variants with different complexity levels, size, and resource requirements. The generative AI model selects the appropriate variant parameters based on the translated service request and edge device capacity, optimizing the balance between performance and resource consumption

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system dynamically adapts model selection to match edge device capabilities. The generative AI translation process considers real-time device resource constraints and dynamically selects or generates model variants that are optimized for the specific device capacity, enabling flexible deployment across heterogeneous edge devices

Inventive Principle:
Principle #15Dynamics

3Productivity

If multiple AI model variants are evaluated for deployment, then resource-optimal selection is achieved, but processing time and computational overhead increase

Engineering Contradiction:
Improvemodel deployment efficiencyVSAvoidmodel selection time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-generating multiple AI model variants with different characteristics before deployment. The generative AI model creates a portfolio of model variants with varying complexity and resource requirements, allowing for efficient selection without extensive real-time evaluation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The model selection process is segmented into distinct phases: translation of service request to parameters, generation of multiple model variants, and selection based on translated parameters. This segmentation allows each phase to be optimized independently, reducing overall processing time while maintaining thorough evaluation

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250307543A1Resource-efficient foundation model deployment on constrained edge devices
Publication Date: 2025.10.02 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20250307543A1 patent drawing
  • US20250307543A1 patent drawing
  • US20250307543A1 patent drawing

AI summary

Computer-implemented methods for efficiently deploying foundation models on resource-constrained edge devices are disclosed herein. Aspects include receiving a text-based service request for an artificial intelligence (AI) model for an edge client. Aspects further include generating model and data descriptions using the text-based service request. Aspects also include generating an AI task capacity profile. Aspects further include selecting a resource-optimal AI model for deployment on the edge device based on the AI task capacity profile.