Mixture of Factual Experts Framework for Controlling Hallucinations in Abstractive Summarization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Abstractive summarization models frequently hallucinate information, leading to inaccurate summaries due to extrinsic and intrinsic factual errors, which are not effectively controlled by existing methods despite high empirical performance on evaluation metrics.
Innovation Solution
The Mixture of Factual Experts (MoFE) framework, which ensembles factual expert models trained to minimize different types of hallucinations, using factual consistency metrics to filter training data and adjust weights for optimal factual quality, combines experts through logits or weighted averaging to generate summaries with controlled factual accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If neural abstractive text summarization systems are trained by maximizing the likelihood of reference summary given its source document, then plausible summaries are generated, but factual errors (hallucinations) occur at high frequency
Solution Approach 1:
The patent segments the summarization task into multiple independent factual expert models, each specialized in generating summaries with different types of factual quality (e.g., low extrinsic hallucinations, low intrinsic hallucinations, high informativeness). Each expert model is trained on filtered training data subsets corresponding to its specific factual quality goal, allowing them to operate independently and contribute different strengths to the final summary.
Solution Approach 2:
The patent merges multiple factual expert models into an ensemble system where each expert's output is combined through weighted averaging or logits ensembling. The final summary is generated by combining the outputs of multiple experts, each with different specialization, thereby aggregating their factual accuracy strengths while mitigating individual weaknesses.
2Productivity
If higher empirical performance is achieved on standard evaluation metrics such as ROUGE score, then summarization quality improves, but faithfulness to the source document decreases
Solution Approach 1:
The patent applies local quality by training different factual expert models on different subsets of training data filtered by specific factual consistency metrics. Each expert model focuses on optimizing for a particular aspect of factual quality (e.g., entity overlap precision, dependency arc entailment accuracy), allowing each model to have specialized local expertise in different types of factual accuracy.
Solution Approach 2:
The patent implements feedback mechanisms during training where factual consistency metrics (such as entity overlap precision and dependency arc entailment accuracy) are used to evaluate and filter training data. This feedback loop ensures that only training samples meeting certain factual consistency thresholds are used to train experts, continuously improving factual accuracy while maintaining evaluation metric performance.
3Reliability
If multiple factual expert models are ensembled to control hallucinations, then factual accuracy improves, but system complexity increases
Solution Approach 1:
The patent uses parameter changes by adjusting the weights of different factual expert models in the ensemble based on their performance on specific factual consistency metrics. The system dynamically determines which experts to include and their relative weights, allowing optimization of factual accuracy while managing computational resources. This parameter-based approach provides flexibility to balance accuracy and complexity.
Data Source
AI summary
Embodiments described herein provide a document summarization framework that controls different factual errors, referred to as “Mixture of Factual Experts (MoFE)” framework. MoFE applies an ensemble of factual expert models to control hallucination in summarization systems. Each factual expert model is trained to generate summaries with a unique type of factual quality. Factual consistency metrics may be used to filter training data in order to adjust the training inputs for each respective expert. The overall factual quality of MoFE may be achieved by controlling the relative weight of each factual expert. The experts may be ensembled (either through logits ensembling, or weighted average of parameters) in order to create a combined output that shares characteristics from each according to its relative weight.


