Copilot Microservice Qualification for Domain-Restricted LLM Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large Language Models (LLMs) face challenges with increasing computational burden, latency, and generation of undesirable artifacts, limiting their practical application in various domains.
Innovation Solution
A microservice architecture utilizing small to mid-sized trained machine learning tools, each performing specialized functions, interconnected to form a network that can be efficiently deployed on single-node systems, with features like retrieval augmented generation and qualification services to enhance performance and reduce artifacts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If large LLMs with more parameters are used, then language processing capability is improved, but computational burden and inference time increase
Solution Approach 1:
The patent segments the large language model into multiple smaller specialized microservices, each trained for specific tasks (e.g., code generation, natural language processing, data analysis). This segmentation maintains high capability in each domain while reducing the computational burden of any single model, thereby decreasing inference time without sacrificing overall language processing capability.
Solution Approach 2:
The patent creates a universal microservice architecture where multiple specialized small LLMs work together to perform diverse language processing tasks. Each microservice is optimized for specific functions but the collective system provides universal language processing capabilities comparable to large LLMs, with reduced computational requirements for each individual component.
2Measurement precision
If large LLMs with more parameters are used, then language processing capability is improved, but computational resources required increase
Solution Approach 1:
The patent divides the computational workload across multiple smaller specialized microservices instead of using one large LLM. Each microservice requires fewer computational resources and can be deployed on less powerful hardware, reducing overall energy consumption while maintaining comprehensive language processing capabilities through the ensemble of specialized models.
Solution Approach 2:
The patent employs multiple small, inexpensive microservices that can be independently deployed and scaled. These smaller models require less computational power and energy to run compared to a single large LLM, making the system more energy-efficient while providing equivalent or superior specialized capabilities through their collective functionality.
3Measurement precision
If large LLMs are used, then comprehensive knowledge is improved, but generation of undesirable artifacts increases
Solution Approach 1:
The patent segments the knowledge processing into multiple specialized microservices, each trained on specific domains and tasks. This specialization reduces the generation of undesirable artifacts because each model focuses on its specific domain rather than attempting to cover all knowledge areas, thereby improving knowledge accuracy within each domain while minimizing hallucinations and errors.
Solution Approach 2:
The patent implements feedback mechanisms where the output of one microservice can be validated or refined by other specialized microservices. This cross-validation process reduces undesirable artifacts by allowing multiple specialized models to check and refine each other's outputs, improving overall knowledge accuracy while filtering out errors and hallucinations.
Data Source
AI summary
Apparatus and methods are disclosed for implementing a copilot as a network of microservices including specialized large language models (LLMs) or other trained machine learning (ML) tools. The microservice network architecture supports flexible, customizable, or dynamically determinable dataflow from client input to corresponding output. Compared to much larger competing LLMs, comparable or superior performance is achieved for certain tasks, while significantly reducing computation time and hardware requirements, even to a single compute node with a single GPU. Examples incorporate a qualification microservice to test data, destined for a downstream microservice, for conformance with the copilot's competency. A knowledge graph of a corpus of documents is built, visualized, and pruned. The data is tested for conformance with the pruned graph representation, and non-conforming data is excluded from the dataflow. Variations and additional techniques are disclosed.


