Multi-Agent Alignment With Observer Feedback for Reliable LLM Tasks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional large language models (LLMs) face challenges in performing complex tasks due to unpredictable output, lack of self-awareness, alignment with diverse human preferences, and inefficiencies in prompt engineering, leading to issues like AI hallucinations and resource-intensive fine-tuning, which are not scalable for dynamic environments.

Innovation Solution

Integrating generative AI models with adaptive machine learning processes and layered memory structures, using observer agents and Bayesian-inspired approaches to align agents with user-specific roles, reduce prompt engineering effort, and optimize resource use through hierarchical planning and asynchronous coordination.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional LLMs are used to perform complex tasks, then they can handle diverse tasks with general knowledge, but they produce unpredictable output and suffer from AI hallucinations

Engineering Contradiction:
Improvetask handling capabilityVSAvoidoutput predictability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces an observer agent as an intermediary component that monitors and validates the outputs of the primary LLM agent. This observer agent checks for hallucinations and validates responses against ground truth or reasoning traces, thereby improving output reliability without compromising the primary agent's task handling capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback loops where observer agents provide validation feedback to primary agents, and reasoning traces are used to feedback into the system for improving future responses. This feedback mechanism helps reduce hallucinations and improves output predictability while maintaining versatility.

Inventive Principle:
Principle #23Feedback

2Reliability

If fine-tuning is applied to improve alignment with user preferences, then alignment improves, but resource consumption and complexity increase significantly

Engineering Contradiction:
Improvealignment with user preferencesVSAvoidfine-tuning complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system enables agents to dynamically adapt to user preferences through self-service mechanisms where observer agents monitor user interactions and automatically adjust agent behavior or prompt strategies without requiring external fine-tuning. This reduces the complexity of manual fine-tuning processes while maintaining alignment.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent employs preliminary actions by pre-configuring observer agents and reasoning trace mechanisms that proactively prevent misalignment issues before they occur, rather than relying on reactive fine-tuning. This preliminary setup reduces the need for complex iterative fine-tuning processes.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If prompt engineering is intensified to improve task performance, then task accuracy improves, but time consumption and resource use increase

Engineering Contradiction:
Improvetask accuracyVSAvoidprompt engineering time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system implements self-service prompt optimization where observer agents automatically analyze task requirements and generate or adjust prompts dynamically based on the situation, eliminating the need for manual prompt engineering for each task while maintaining high accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent makes the prompt engineering process dynamic by allowing prompts to be automatically adjusted in real-time based on task context, user preferences, and observer agent feedback, rather than using static pre-engineered prompts. This dynamic adaptation improves accuracy without increasing time investment.

Inventive Principle:
Principle #15Dynamics

4Extent of automation

If autonomous agents are deployed to perform user-level tasks, then automation level improves, but control and alignment with human preferences become difficult

Engineering Contradiction:
Improveautonomous task executionVSAvoidalignment with human preferences
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The observer agent serves as an intermediary between the autonomous primary agent and human users, monitoring agent actions and ensuring alignment with human preferences. This intermediary layer maintains high automation while providing continuous alignment verification.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements multi-layer feedback mechanisms where observer agents provide feedback on agent actions, reasoning traces capture decision-making processes, and user feedback loops enable continuous alignment improvement. This feedback structure maintains autonomy while ensuring alignment with human preferences.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP4657315A1Dynamic agents with real-time alignment
Publication Date: 2025.12.03 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP4657315A1 patent drawingFigure 1
  • EP4657315A1 patent drawingFigure 2
  • EP4657315A1 patent drawingFigure 3

AI summary

An example may receive at least one input via at least one device. An example may use the at least one input to determine an objective. An example may use the objective, a multi-agent system, and an automated agent to cause at least one first sub-agent of the multi-agent system to generate and execute a first plan including one or more tasks to achieve the objective. An example may cause at least one second sub-agent of the multi-agent system to execute a second plan to supervise the at least one first sub-agent in accordance with a supervision level that indicates a level of supervision of the automated agent by an entity associated with the objective.