Multi-Level AI Architecture for Scalable Generative Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current AI generative models and Large Language Models (LLMs) require significant supercomputing efforts, leading to high computational costs, slow response times, and limited scalability and expandability.

Innovation Solution

The system employs a multi-level generative AI and LLM architecture that utilizes derived requests, multiple h-LLMs with varying levels of accuracy, and a distributed architecture to efficiently process client requests, assign tasks, and combine results for improved performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If current AI generative models and LLMs are used, then language processing and content generation capabilities are achieved, but computational costs are high and response times are slow

Engineering Contradiction:
Improveresponse timeVSAvoidcomputational cost
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent segments the monolithic LLM into multiple smaller specialized models (h-LLMs) organized in a hierarchical structure. Each h-LLM is trained on specific datasets for particular tasks, allowing the system to process requests in parallel and reduce overall computational burden while maintaining response quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a hierarchical dimension to the model architecture, organizing h-LLMs into multiple levels where higher-level models coordinate lower-level specialized models. This dimensional transformation enables more efficient resource allocation and reduces computational costs through distributed processing.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If current AI generative models and LLMs are used, then content generation capability is achieved, but scalability and expandability are limited

Engineering Contradiction:
ImprovescalabilityVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal hierarchical framework where h-LLMs can be dynamically allocated to different tasks based on request requirements. The system can handle diverse generative tasks (text, code, images, video) using the same architectural pattern, enabling easy scaling and expansion without redesigning the core system.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements a dynamic request routing system that assigns incoming requests to appropriate h-LLMs based on task characteristics, model availability, and performance metrics. This dynamic allocation enables the system to adapt to changing workloads and scale efficiently as new models are added.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12299017B2Method and system for multi-level artificial intelligence supercomputer design
Publication Date: 2025.05.13 MADISETTI VIJAY
  • US12299017B2 patent drawing
  • US12299017B2 patent drawing
  • US12299017B2 patent drawing

AI summary

A method of generating responses responsive to input prompts including receiving a multimodal input prompt, the multi-modal input prompt comprising a first input prompt mode and a second input prompt mode, generating first and second derived prompts responsive to the first and second input prompt modes, transmitting the first and second derived prompts to first and second h-LLMs that are trained using data having a mode that is the same of the respective input prompt modes, receiving one or more first and second results from the first and second h-LLMs, generating at least one combined result from the one or more first and second results, and transmitting the combined result to the user.