Copilot Microservice Qualification for Domain-Restricted LLM Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large Language Models (LLMs) face challenges with increasing computational burden, latency, and generation of undesirable artifacts, limiting their practical application in various domains.

Innovation Solution

A microservice architecture utilizing small to mid-sized trained machine learning tools, each performing specialized functions, interconnected to form a network that can be efficiently deployed on single-node systems, with features like retrieval augmented generation and qualification services to enhance performance and reduce artifacts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If large LLMs with more parameters are used, then language processing capability is improved, but computational burden and inference time increase

Engineering Contradiction:
Improvelanguage processing capabilityVSAvoidinference time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the large language model into multiple smaller specialized microservices, each trained for specific tasks (e.g., code generation, natural language processing, data analysis). This segmentation maintains high capability in each domain while reducing the computational burden of any single model, thereby decreasing inference time without sacrificing overall language processing capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal microservice architecture where multiple specialized small LLMs work together to perform diverse language processing tasks. Each microservice is optimized for specific functions but the collective system provides universal language processing capabilities comparable to large LLMs, with reduced computational requirements for each individual component.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If large LLMs with more parameters are used, then language processing capability is improved, but computational resources required increase

Engineering Contradiction:
Improvelanguage processing capabilityVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent divides the computational workload across multiple smaller specialized microservices instead of using one large LLM. Each microservice requires fewer computational resources and can be deployed on less powerful hardware, reducing overall energy consumption while maintaining comprehensive language processing capabilities through the ensemble of specialized models.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs multiple small, inexpensive microservices that can be independently deployed and scaled. These smaller models require less computational power and energy to run compared to a single large LLM, making the system more energy-efficient while providing equivalent or superior specialized capabilities through their collective functionality.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Measurement precision

If large LLMs are used, then comprehensive knowledge is improved, but generation of undesirable artifacts increases

Engineering Contradiction:
Improveknowledge accuracyVSAvoidundesirable artifacts
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The patent segments the knowledge processing into multiple specialized microservices, each trained on specific domains and tasks. This specialization reduces the generation of undesirable artifacts because each model focuses on its specific domain rather than attempting to cover all knowledge areas, thereby improving knowledge accuracy within each domain while minimizing hallucinations and errors.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements feedback mechanisms where the output of one microservice can be validated or refined by other specialized microservices. This cross-validation process reduces undesirable artifacts by allowing multiple specialized models to check and refine each other's outputs, improving overall knowledge accuracy while filtering out errors and hallucinations.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250342171A1Copilot implementation: restricting operation to a domain of competence
Publication Date: 2025.11.06 THIA ST CO
  • US20250342171A1 patent drawing
  • US20250342171A1 patent drawing
  • US20250342171A1 patent drawing

AI summary

Apparatus and methods are disclosed for implementing a copilot as a network of microservices including specialized large language models (LLMs) or other trained machine learning (ML) tools. The microservice network architecture supports flexible, customizable, or dynamically determinable dataflow from client input to corresponding output. Compared to much larger competing LLMs, comparable or superior performance is achieved for certain tasks, while significantly reducing computation time and hardware requirements, even to a single compute node with a single GPU. Examples incorporate a qualification microservice to test data, destined for a downstream microservice, for conformance with the copilot's competency. A knowledge graph of a corpus of documents is built, visualized, and pruned. The data is tested for conformance with the pruned graph representation, and non-conforming data is excluded from the dataflow. Variations and additional techniques are disclosed.