GAI Reasoning Evaluation Using Factual and Counterfactual Prompts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing generative artificial intelligence (GAI) models struggle with computational reasoning, particularly in handling counterfactual scenarios, leading to inaccuracies in decision-making and problem-solving tasks.

Innovation Solution

A method involving submitting factual and counterfactual prompts to multiple GAI models, computing probability of necessity (PN) and probability of sufficiency (PS) values, and selecting the model with superior reasoning performance based on these values for specific tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If GAI models are used for computational reasoning tasks, then decision-making capability is improved, but accuracy in counterfactual scenarios deteriorates

Engineering Contradiction:
Improvedecision-making capabilityVSAvoidaccuracy in counterfactual scenarios
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent applies universality by developing evaluation frameworks that assess multiple reasoning capabilities (factual reasoning, counterfactual reasoning, causal reasoning) within a single GAI model evaluation system. This allows comprehensive measurement of decision-making capabilities across different scenario types, enabling the selection of models that perform well universally across both factual and counterfactual domains.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent employs parameter changes by introducing new evaluation metrics (PN and PS values) that specifically measure counterfactual reasoning performance. By changing the parameters of evaluation from traditional accuracy metrics to these new probabilistic necessity and sufficiency metrics, the system can accurately assess and compare models' counterfactual reasoning capabilities, thereby resolving the accuracy deterioration issue.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If multiple GAI models are evaluated using PN and PS values, then model selection accuracy is improved, but computational complexity increases

Engineering Contradiction:
Improvemodel selection accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the evaluation process into distinct components: factual prompt evaluation, counterfactual prompt evaluation, PN value computation, and PS value computation. This segmented approach allows each component to be optimized independently and enables parallel processing, thereby managing computational complexity while maintaining high model selection accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces PN and PS values as intermediary metrics that mediate between raw model outputs and final model selection decisions. These intermediary values simplify the comparison process by providing standardized probabilistic measures of reasoning performance, reducing the overall computational complexity of model selection while improving accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If factual and counterfactual prompts are submitted to multiple models, then reasoning performance evaluation is improved, but time consumption increases

Engineering Contradiction:
Improvereasoning performance evaluationVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-processing and structuring the factual and counterfactual prompts before submitting them to multiple models. This includes preparing standardized prompt formats, pre-computing necessary baseline values, and organizing evaluation criteria in advance. Such preliminary preparations enable more efficient model evaluation, reducing time consumption while maintaining high measurement precision.

Inventive Principle:
Principle #10Preliminary action

4Measurement precision

If GAI models with superior reasoning performance are selected, then problem-solving accuracy is improved, but system complexity increases

Engineering Contradiction:
Improveproblem-solving accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements feedback mechanisms where PN and PS evaluation results are fed back into the model selection process. This feedback loop enables continuous optimization of model selection based on actual reasoning performance measurements, ensuring that models with superior problem-solving accuracy are selected while managing system complexity through iterative refinement rather than complex upfront design.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20260050792A1Evaluating computational reasoning performance of generative artificial intelligence models
Publication Date: 2026.02.19 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20260050792A1 patent drawing
  • US20260050792A1 patent drawing
  • US20260050792A1 patent drawing

AI summary

Systems and methods evaluate computational reasoning performance of generative artificial intelligence (GAI) models. Both a factual prompt and a counterfactual prompt are submitted to both first and second GAI models, thereby generating first factual and counterfactual outputs for the first GAI model and second factual and counterfactual outputs for the second GAI model. Probability of necessity (PN) and probability of sufficiency (PS) values are computed for both the first and second GAI models based on their associated factual output and counterfactual output. The computational reasoning performance of the first GAI model relative to the second GAI model are compared based on the PN and PS values. One of the first or the second GAI models is selected based on the comparison and submitted a target prompt using the selected one of the first and second GAI model.