Document QA Data Generation Using Multi-Hop Reasoning Chains
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current document question-answering systems face challenges in generating high-quality, diverse, and complex question-answer pairs due to reliance on manual annotation methods that are labor-intensive and prone to biases, or automated synthesis methods that lack diversity and fail to cover multi-hop and set questions.
Innovation Solution
A method involving a preset question-answering data generation model that extracts descriptive information from page images, generates reasoning chains based on multi-hop and set questions, and produces question-answering data using large language models to enhance data quality, diversity, and complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If simple template generation or question reuse methods are used for automated data synthesis, then the process is efficient and automated, but the quality and diversity of synthesized data is low
Solution Approach 1:
The patent introduces an intermediary component (the question-answering data generation model with reasoning chain capability) between the automated template generation and the final question-answer pairs. This intermediary processes the template-generated content through multi-hop reasoning and set operation reasoning, transforming simple automated output into high-quality, diverse question-answer data that covers complex question types.
Solution Approach 2:
The patent replaces the mechanical template-filling approach with an intelligent system that uses reasoning chains and question-answering models. Instead of simply substituting values into predefined templates, the system generates questions through reasoning processes that understand document content, enabling both automation and high quality.
2Productivity
If simple template generation or question reuse methods are used for automated data synthesis, then the process is efficient and automated, but the diversity of synthesized data is low
Solution Approach 1:
The patent implements dynamics by making the question generation process adaptive rather than static. The system dynamically selects between different reasoning types (multi-hop reasoning, set operation reasoning) based on the document content and question requirements. This allows the automated process to generate diverse question types including complex multi-hop and set questions, rather than being constrained to fixed template patterns.
Solution Approach 2:
The patent changes the parameters of the generation system by introducing reasoning chain depth and reasoning type as variable parameters. Instead of using fixed template parameters, the system adjusts reasoning parameters (number of hops, set operations) to generate diverse question types while maintaining automated efficiency.
3Manufacturing precision
If manual annotation methods are used, then the quality of question-answer pairs can be high, but the process is labor-intensive and prone to biases
Solution Approach 1:
The patent implements self-service by enabling the system to automatically generate high-quality question-answer pairs without human intervention. The question-answering data generation model with reasoning chain capability performs the annotation task itself, eliminating labor-intensive manual processes while maintaining high quality through intelligent reasoning.
Solution Approach 2:
The patent replaces the mechanical manual annotation process with an intelligent automated system. Instead of human annotators manually creating question-answer pairs, the system uses reasoning chains and question-answering models to automatically generate high-quality data, eliminating labor intensity and human biases while maintaining or improving quality.
Data Source
AI summary
The present disclosure provides a method for generating document question-answering data, a training method, a generating apparatus, an electronic device, a computer-readable storage medium, and a computer program product. The method includes: extracting page content from page images in a document to obtain descriptive information corresponding to seed pages in the document; generating a reasoning chain corresponding to the seed pages by using a preset question-answering data generation model based on the descriptive information, question definitions of preset question types, and question-answering examples of the preset question types; and in response to the reasoning chain constituting a question-type reasoning chain corresponding to the preset question types, generating question-answering data corresponding to the preset question types by using the question-answering data generation model based on the question-type reasoning chain.


