Middleware Platform for Multi-LLM Response Aggregation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems are unable to efficiently support new language models as they are introduced at a rapid pace, often limiting organizations to onboard one model at a time, which hampers the ability to leverage the latest features and technologies in a timely manner.

Innovation Solution

A computer-implemented middleware platform that provides generative AI services by interacting with multiple large language models (LLMs), combining their responses, and integrating proprietary data to generate enhanced outputs, while also managing load and providing scalable and autonomous deployment of models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If current systems onboard one language model at a time, then system complexity is reduced and ease of operation is improved, but adaptability to support multiple rapidly introduced language models deteriorates

Engineering Contradiction:
Improveease of onboarding language modelsVSAvoidability to support multiple language models
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The middleware platform is designed as a universal system that can interact with multiple different language models through standardized interfaces. The platform provides multi-functional capabilities to support various LLMs simultaneously, allowing organizations to leverage multiple models without requiring separate onboarding processes for each model.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The middleware acts as an intermediary layer between the user interface and multiple language models. This mediator component manages communications with different LLMs, handling model selection, load distribution, and response aggregation, thereby enabling multi-model support while maintaining operational simplicity for end users.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If multiple language models are integrated simultaneously, then adaptability and productivity are improved, but device complexity and difficulty of managing model interactions increase

Engineering Contradiction:
Improvespeed of integrating new language modelsVSAvoidcomplexity of managing model interactions
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system architecture is segmented into distinct modular components: user interface layer, middleware platform layer, and language model layer. This segmentation allows new language models to be integrated at the model layer without affecting other components, enabling rapid integration while maintaining manageable system complexity through clear separation of concerns.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The middleware platform serves as a complexity-managing intermediary that handles all interactions with multiple language models. It provides centralized control over model selection, load balancing, and response aggregation, thereby enabling high productivity in model integration while containing management complexity within the middleware layer rather than propagating it throughout the entire system.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If responses from multiple language models are combined, then manufacturing precision of output quality is improved, but device complexity and processing time increase

Engineering Contradiction:
Improvequality of generated outputVSAvoidcomplexity of combining model responses
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The middleware platform acts as an intermediary that manages the combination of responses from multiple language models. It implements structured processes for aggregating, evaluating, and synthesizing multiple model outputs into a single high-quality response, thereby improving output quality while containing the complexity of response combination within the middleware layer.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system incorporates feedback mechanisms where responses from multiple language models are evaluated and compared. The middleware uses this feedback to select or synthesize the best response, improving output quality through iterative refinement while managing the complexity of combining multiple responses through systematic evaluation processes.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250053453A1System and method for implementing an advisory assistant to a generative artifical intelligence tool
Publication Date: 2025.02.13 KPMG LLP
  • US20250053453A1 patent drawing
  • US20250053453A1 patent drawing
  • US20250053453A1 patent drawing

AI summary

The invention relates to computer-implemented systems and methods that implement an innovative generative AI service based on proprietary expertise and industry knowledge. The generative AI service provides unique autonomous features, such as combining separate and distinct LLM responses and prompts to create unique results. Other autonomous features may include an ability to handle scaling and auto deployment of models and rerouting requests autonomously to ensure user load is balanced across the entire globally distributed Generative AI infrastructure estate. The generative AI service may further deploy new Production instances of models on demand by predefined system criteria as well as by explicit user request based on projected demand increase and/or the need for specific instance for further model fine-tuning.