Dynamic AI Agents With Supervision Learning for Real-Time Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional large language models (LLMs) face challenges in performing complex tasks due to unpredictable output, lack of self-awareness, alignment with diverse human preferences, and inefficiencies in resource utilization, which hinder their use in dynamic and real-time environments.

Innovation Solution

Integrating generative AI models with adaptive machine learning processes and layered memory structures to align agents with context data, using observer agents and Bayesian-inspired approaches to regulate output and optimize resource use, while ensuring security and reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional large language models are used to perform complex tasks, then they can handle diverse tasks with generative capabilities, but their output is unpredictable and they lack alignment with diverse human preferences

Engineering Contradiction:
Improvetask handling capabilityVSAvoidoutput consistency
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system segments the LLM's operation into distinct phases: planning phase where multiple candidate plans are generated, execution phase where plans are carried out, and evaluation phase where outcomes are assessed. This segmentation allows for controlled exploration of different task approaches while maintaining reliability through systematic evaluation and selection of the best plan.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements feedback loops where the LLM's outputs are evaluated against human preferences and task outcomes. Evaluation metrics and human feedback are used to refine and adjust the planning and execution processes, ensuring that the LLM's generative capabilities remain aligned with human preferences while maintaining output consistency.

Inventive Principle:
Principle #23Feedback

2Extent of automation

If LLMs operate autonomously in dynamic environments, then they can perform user-level tasks without direct human instruction, but they lack self-awareness and alignment with diverse human preferences

Engineering Contradiction:
Improveautonomous task executionVSAvoidalignment with human preferences
Core Design Contradiction:
Extent of automationVSAdaptability or versatility

Solution Approach 1:

The system introduces an intermediary evaluation and regulation layer between the LLM's autonomous operations and human preferences. This intermediary component translates diverse human preferences into evaluable metrics and uses these metrics to guide and adjust the LLM's autonomous behavior, enabling alignment without requiring direct human instruction for each task.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements dynamic adjustment mechanisms where the LLM's autonomous behavior is continuously adapted based on real-time evaluation against human preferences. The planning, execution, and evaluation phases form a dynamic cycle that allows the LLM to learn and adjust its behavior to better align with diverse human preferences in changing environments.

Inventive Principle:
Principle #15Dynamics

3Productivity

If LLMs process complex tasks with high computational requirements, then they can achieve sophisticated output, but they exhibit inefficiencies in resource utilization

Engineering Contradiction:
Improvecomplex task processing capabilityVSAvoidcomputational resource consumption
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The system applies partial action by generating multiple candidate plans and executing only the most promising ones based on evaluation metrics. Rather than exhaustively exploring all possible task approaches or using maximum computational resources for every operation, the system selectively processes the most relevant plans, reducing overall resource consumption while maintaining productivity.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system changes operational parameters dynamically based on task complexity and resource availability. By adjusting the number of candidate plans generated, the depth of evaluation, and the computational resources allocated to different phases, the system optimizes the balance between processing capability and resource utilization, achieving sophisticated output with improved efficiency.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4657312A1Dynamic agents with real-time alignment
Publication Date: 2025.12.03 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP4657312A1 patent drawingFigure 1
  • EP4657312A1 patent drawingFigure 2
  • EP4657312A1 patent drawingFigure 3

AI summary

An example may receive at least one input via at least one device. An example may use the at least one input to determine an entity identity. An example may use the entity identity to create an automated agent and load context data associated with the entity identity into at least one layer of a multi-layer memory of the automated agent. An example may cause the automated agent to machine-learn a supervision level via the context data. The machine-learned supervision level may indicate a level of supervision of the automated agent by an entity associated with the entity identity. An example may configure the automated agent to execute a task on behalf of the entity and in accordance with the machine-learned supervision level.