LLM Capability Extraction Through Indirect Output Evaluation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional large language models with billions of parameters are complex and expensive, making it difficult to determine their actual capabilities effectively, as direct natural language queries are not sufficient for discovering their potential.

Innovation Solution

Employ indirect interaction stages with the model, where the output is not natural language or not semantically responsive, and evaluate the results to estimate the model's capabilities through task completion, code generation, or external references.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If direct natural language queries are used to determine model capabilities, then the interaction is simple and straightforward, but the capability discovery is insufficient and inaccurate

Engineering Contradiction:
Improvecapability discovery accuracyVSAvoidinteraction complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary evaluation system that mediates between the language model and the capability assessment. Instead of directly querying the model about its capabilities, the system uses indirect tasks and outputs as intermediaries to infer capabilities. This resolves the contradiction by maintaining simple direct queries while achieving accurate capability discovery through the intermediary evaluation layer.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent inverts the traditional capability assessment approach by not asking the model directly what it can do, but rather observing what the model produces and inferring capabilities from the outputs. This inversion allows for more accurate capability measurement while avoiding the limitations of direct self-reporting by the model.

Inventive Principle:
Principle #13The other way round (Inversion)

2Measurement precision

If indirect interaction stages are employed to discover model capabilities, then capability discovery accuracy improves, but the evaluation process becomes more complex

Engineering Contradiction:
Improvecapability estimation accuracyVSAvoidevaluation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements preliminary action by pre-defining a structured set of evaluation tasks and capability categories before conducting the assessment. The evaluation framework, task templates, and capability taxonomy are prepared in advance, which streamlines the actual evaluation process. This reduces the time loss during evaluation while maintaining high accuracy through the comprehensive pre-planned indirect interaction stages.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If comprehensive capability evaluation is performed through multiple indirect stages, then model potential is fully identified, but the computational resources and complexity increase

Engineering Contradiction:
Improvemodel capability identificationVSAvoidevaluation system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the capability evaluation into distinct stages and independent task categories. Each indirect interaction stage focuses on specific capability dimensions, and the evaluation system processes different aspects separately. This segmentation enables comprehensive capability identification while managing complexity through modular, organized evaluation components that can be independently configured and executed.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12475331B2Model capability extraction
Publication Date: 2025.11.18 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12475331B2 patent drawing
  • US12475331B2 patent drawing
  • US12475331B2 patent drawing

AI summary

The indirect querying of models to determine capabilities possessed by the model. Such indirect queries take the form of model input that potentially includes a natural language input user data. Such model input is structured such that the output of the model is either not natural language at all, or else is natural language that is not semantically responsive to the natural language input. Nevertheless, the output is evaluated to estimate or determine the capability possessed by the model. Thus, models may be more fully utilized to their better potential.