Mixture of Factual Experts Framework for Controlling Hallucinations in Abstractive Summarization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Abstractive summarization models frequently hallucinate information, leading to inaccurate summaries due to extrinsic and intrinsic factual errors, which are not effectively controlled by existing methods despite high empirical performance on evaluation metrics.

Innovation Solution

The Mixture of Factual Experts (MoFE) framework, which ensembles factual expert models trained to minimize different types of hallucinations, using factual consistency metrics to filter training data and adjust weights for optimal factual quality, combines experts through logits or weighted averaging to generate summaries with controlled factual accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If neural abstractive text summarization systems are trained by maximizing the likelihood of reference summary given its source document, then plausible summaries are generated, but factual errors (hallucinations) occur at high frequency

Engineering Contradiction:
Improvesummary generationVSAvoidfactual accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent segments the summarization task into multiple independent factual expert models, each specialized in generating summaries with different types of factual quality (e.g., low extrinsic hallucinations, low intrinsic hallucinations, high informativeness). Each expert model is trained on filtered training data subsets corresponding to its specific factual quality goal, allowing them to operate independently and contribute different strengths to the final summary.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges multiple factual expert models into an ensemble system where each expert's output is combined through weighted averaging or logits ensembling. The final summary is generated by combining the outputs of multiple experts, each with different specialization, thereby aggregating their factual accuracy strengths while mitigating individual weaknesses.

Inventive Principle:
Principle #5Merging (Combining)

2Productivity

If higher empirical performance is achieved on standard evaluation metrics such as ROUGE score, then summarization quality improves, but faithfulness to the source document decreases

Engineering Contradiction:
Improveevaluation metric performanceVSAvoidfaithfulness to source
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies local quality by training different factual expert models on different subsets of training data filtered by specific factual consistency metrics. Each expert model focuses on optimizing for a particular aspect of factual quality (e.g., entity overlap precision, dependency arc entailment accuracy), allowing each model to have specialized local expertise in different types of factual accuracy.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements feedback mechanisms during training where factual consistency metrics (such as entity overlap precision and dependency arc entailment accuracy) are used to evaluate and filter training data. This feedback loop ensures that only training samples meeting certain factual consistency thresholds are used to train experts, continuously improving factual accuracy while maintaining evaluation metric performance.

Inventive Principle:
Principle #23Feedback

3Reliability

If multiple factual expert models are ensembled to control hallucinations, then factual accuracy improves, but system complexity increases

Engineering Contradiction:
Improvefactual accuracyVSAvoidmodel ensemble structure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent uses parameter changes by adjusting the weights of different factual expert models in the ensemble based on their performance on specific factual consistency metrics. The system dynamically determines which experts to include and their relative weights, allowing optimization of factual accuracy while managing computational resources. This parameter-based approach provides flexibility to balance accuracy and complexity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20230119109A1Systems and methods for controlling hallucinations in abstractive summarization
Publication Date: 2023.04.20 SALESFORCE INC
  • US20230119109A1 patent drawing
  • US20230119109A1 patent drawing
  • US20230119109A1 patent drawing

AI summary

Embodiments described herein provide a document summarization framework that controls different factual errors, referred to as “Mixture of Factual Experts (MoFE)” framework. MoFE applies an ensemble of factual expert models to control hallucination in summarization systems. Each factual expert model is trained to generate summaries with a unique type of factual quality. Factual consistency metrics may be used to filter training data in order to adjust the training inputs for each respective expert. The overall factual quality of MoFE may be achieved by controlling the relative weight of each factual expert. The experts may be ensembled (either through logits ensembling, or weighted average of parameters) in order to create a combined output that shares characteristics from each according to its relative weight.