Foundation Model Deployment via AI Capacity Profiling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing FMaaS platforms struggle to interpret non-standardized, text-based, and/or data-based prompts for AI model deployment on resource-constrained edge devices, often requiring expert intervention or discarding such requests, and the limited computing resources of edge devices restrict model deployment options.
Innovation Solution
A system using generative AI through a pre-trained large language model to translate text-based client requirements into interpretable FMaaS requests, identifying resource-optimal AI models by generating model and data descriptions, and performing AI task capacity profiling to select compatible and efficient model variants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If text-based service requests are directly processed by existing FMaaS platforms, then expert intervention is required and requests may be discarded, but this increases operational complexity and reduces automation
Solution Approach 1:
A generative AI translation layer is introduced as an intermediary between text-based service requests and the FMaaS platform. This translation layer converts natural language requests into standardized model selection parameters, enabling automated processing without expert intervention while maintaining platform compatibility
Solution Approach 2:
The system enables self-service by allowing the generative AI model to autonomously interpret text-based requests and automatically select appropriate AI models based on the translated parameters, eliminating the need for expert human intervention in the model selection process
2Reliability
If complex foundation models are deployed on edge devices, then AI task performance is improved, but resource consumption exceeds device capacity
Solution Approach 1:
The system changes the parameters of AI models by generating multiple variants with different complexity levels, size, and resource requirements. The generative AI model selects the appropriate variant parameters based on the translated service request and edge device capacity, optimizing the balance between performance and resource consumption
Solution Approach 2:
The system dynamically adapts model selection to match edge device capabilities. The generative AI translation process considers real-time device resource constraints and dynamically selects or generates model variants that are optimized for the specific device capacity, enabling flexible deployment across heterogeneous edge devices
3Productivity
If multiple AI model variants are evaluated for deployment, then resource-optimal selection is achieved, but processing time and computational overhead increase
Solution Approach 1:
The system performs preliminary action by pre-generating multiple AI model variants with different characteristics before deployment. The generative AI model creates a portfolio of model variants with varying complexity and resource requirements, allowing for efficient selection without extensive real-time evaluation
Solution Approach 2:
The model selection process is segmented into distinct phases: translation of service request to parameters, generation of multiple model variants, and selection based on translated parameters. This segmentation allows each phase to be optimized independently, reducing overall processing time while maintaining thorough evaluation
Data Source
AI summary
Computer-implemented methods for efficiently deploying foundation models on resource-constrained edge devices are disclosed herein. Aspects include receiving a text-based service request for an artificial intelligence (AI) model for an edge client. Aspects further include generating model and data descriptions using the text-based service request. Aspects also include generating an AI task capacity profile. Aspects further include selecting a resource-optimal AI model for deployment on the edge device based on the AI task capacity profile.


