Document QA Data Generation Using Multi-Hop Reasoning Chains

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current document question-answering systems face challenges in generating high-quality, diverse, and complex question-answer pairs due to reliance on manual annotation methods that are labor-intensive and prone to biases, or automated synthesis methods that lack diversity and fail to cover multi-hop and set questions.

Innovation Solution

A method involving a preset question-answering data generation model that extracts descriptive information from page images, generates reasoning chains based on multi-hop and set questions, and produces question-answering data using large language models to enhance data quality, diversity, and complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If simple template generation or question reuse methods are used for automated data synthesis, then the process is efficient and automated, but the quality and diversity of synthesized data is low

Engineering Contradiction:
Improveautomation of data synthesisVSAvoidquality of synthesized data
Core Design Contradiction:
Extent of automationVSManufacturing precision

Solution Approach 1:

The patent introduces an intermediary component (the question-answering data generation model with reasoning chain capability) between the automated template generation and the final question-answer pairs. This intermediary processes the template-generated content through multi-hop reasoning and set operation reasoning, transforming simple automated output into high-quality, diverse question-answer data that covers complex question types.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical template-filling approach with an intelligent system that uses reasoning chains and question-answering models. Instead of simply substituting values into predefined templates, the system generates questions through reasoning processes that understand document content, enabling both automation and high quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If simple template generation or question reuse methods are used for automated data synthesis, then the process is efficient and automated, but the diversity of synthesized data is low

Engineering Contradiction:
Improveefficiency of data synthesisVSAvoiddiversity of question types
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamics by making the question generation process adaptive rather than static. The system dynamically selects between different reasoning types (multi-hop reasoning, set operation reasoning) based on the document content and question requirements. This allows the automated process to generate diverse question types including complex multi-hop and set questions, rather than being constrained to fixed template patterns.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameters of the generation system by introducing reasoning chain depth and reasoning type as variable parameters. Instead of using fixed template parameters, the system adjusts reasoning parameters (number of hops, set operations) to generate diverse question types while maintaining automated efficiency.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If manual annotation methods are used, then the quality of question-answer pairs can be high, but the process is labor-intensive and prone to biases

Engineering Contradiction:
Improvequality of question-answer pairsVSAvoidcomplexity of annotation process
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent implements self-service by enabling the system to automatically generate high-quality question-answer pairs without human intervention. The question-answering data generation model with reasoning chain capability performs the annotation task itself, eliminating labor-intensive manual processes while maintaining high quality through intelligent reasoning.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual annotation process with an intelligent automated system. Instead of human annotators manually creating question-answer pairs, the system uses reasoning chains and question-answering models to automatically generate high-quality data, eliminating labor intensity and human biases while maintaining or improving quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20260079973A1Document question-answering data generation method, electronic device and storage medium
Publication Date: 2026.03.19 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US20260079973A1 patent drawing
  • US20260079973A1 patent drawing
  • US20260079973A1 patent drawing

AI summary

The present disclosure provides a method for generating document question-answering data, a training method, a generating apparatus, an electronic device, a computer-readable storage medium, and a computer program product. The method includes: extracting page content from page images in a document to obtain descriptive information corresponding to seed pages in the document; generating a reasoning chain corresponding to the seed pages by using a preset question-answering data generation model based on the descriptive information, question definitions of preset question types, and question-answering examples of the preset question types; and in response to the reasoning chain constituting a question-type reasoning chain corresponding to the preset question types, generating question-answering data corresponding to the preset question types by using the question-answering data generation model based on the question-type reasoning chain.