Multi-hop retrieval question and answer data set generation method based on large language model

By constructing a multi-hop retrieval question-answer dataset using a depth-first search framework and a controlled action state machine, we solve the problems of missing logical chains and fluctuating construction quality in existing question-answer datasets. This enables efficient, interpretable, and traceable multi-hop question-answer data generation, improving evidence coverage and search success rate for complex tasks.

CN122045374APending Publication Date: 2026-05-15JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610247190.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-02
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing question-answering datasets lack multi-step reasoning logic chains, have large fluctuations in construction quality, and are difficult to interpret and ensure the completeness of the evidence chain. Furthermore, there are redundant calculations and noisy data in the generation process.

Method used

Employing a depth-first search framework and a controlled action state machine, a multi-hop retrieval question-answering dataset is constructed by outputting controlled action sequences through a generative large language model. This dataset includes multi-step retrieval trajectories, evidence fragments, and structured process records. Combined with path evaluation and backtracking error correction mechanisms, multi-hop question-answering samples in JSONL format are generated.

Benefits of technology

It significantly improves the logical interpretability and end-to-end traceability of question-answering datasets, increases the evidence coverage and path search success rate for complex tasks, reduces the spread of irrelevant information and computational costs, has adaptive error correction capabilities, and maintains high throughput and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122045374A_ABST
    Figure CN122045374A_ABST
Patent Text Reader

Abstract

The invention provides a multi-hop retrieval question and answer data set generation method based on a large language model, and belongs to the field of artificial intelligence, natural language processing and deep learning and data engineering. The method is used for automatically converting a common question and answer sample only containing a question-answer tag into a multi-hop retrieval question and answer sample containing a multi-step retrieval track, an evidence fragment and a structured process record. Comprising the steps of reading an original question and answer sample and performing fragmentation processing; constructing an initial search state taking a search tree node as a carrier, and initializing a root node containing a problem, a path and an evidence pool; under a depth-first search framework supported by an explicit stack, outputting a search tree driven by a controlled action sequence by the generative large language model; performing evidence paragraph extraction and quality filtering on the retrieved webpage, and performing backtracking freezing and experience injection according to a path evaluation result; and outputting a JSONL-format multi-hop question and answer sample containing a complete retrieval track and a structured process record.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to artificial intelligence, natural language processing and deep learning, and data engineering technology, specifically to a method for generating multi-hop retrieval question-answering datasets based on a large language model. Background Technology

[0002] Large Language Models (LLMs) and search engines have developed methods for automatically constructing multi-hop retrieval question-answer (QA) datasets from ordinary question-answer samples, enabling them to possess interpretable retrieval trajectories, multi-dimensional evidence chains, and structured reasoning processes. With the widespread application of LLMs across various industries, retrieval enhancement generation techniques have become a core means of addressing model illusions and improving answer accuracy. However, most existing question-answer datasets only contain "question-answer" pairs, severely lacking intermediate process data that reflects the model's reasoning path.

[0003] The limitations of existing technologies are mainly reflected in the following aspects: (1) Lack of logical chain: Ordinary QA samples only provide the final result, which cannot support the model to learn the complex reasoning logic of "decomposition of problem - multi-step retrieval - information aggregation", which makes it easy for the model to break evidence when dealing with complex multi-hop problems. (2) Large fluctuation in construction quality: Traditional automated construction methods mostly rely on hard-coded rules or simple template splicing, lacking dynamic diagnosis and error correction mechanisms. When faced with typical failure forms such as "irrelevant retrieval" or "repeated retrieval", the system is difficult to automatically backtrack. (3) Lack of interpretability: Existing single-hop retrieval methods often only perform a single query, which is difficult to meet the needs of multi-stage information completion and the retrieval process is black box, which cannot provide a traceable structured evidence flow for model training. (4) Low scalability efficiency: In the process of constructing large-scale datasets, due to the lack of fine concurrency control and quality filtering strategies, a large amount of redundant calculation and noise data generation are often caused, which increases the computing power cost and reduces the purity of the dataset.

[0004] In summary, how to construct a method for generating multi-hop retrieval question-answer datasets that can simulate human thought processes, possess adaptive backtracking capabilities, and output highly structured results has become a critical problem urgently needing to be solved in the field of data engineering. The method proposed in this invention optimizes retrieval state modeling and controlled action decision-making mechanisms, fully leverages the synergistic advantages of generative large language models and explicit search algorithms, improves the logical interpretability and evidence chain completeness of multi-hop question-answer samples, and significantly enhances the automated construction and adaptive error correction capabilities of reasoning paths in complex retrieval scenarios. Summary of the Invention

[0005] This invention provides an achievable and scalable method for generating multi-hop retrieval question-answering datasets. Given only the input condition of "question + answer label", it automatically generates multi-step retrieval trajectories, evidence fragments, structured process records and candidate answers. Furthermore, it improves the quality and efficiency of dataset generation through evaluation feedback and backtracking error correction.

[0006] This invention provides a method for generating multi-hop retrieval question-answering datasets based on a large language model, including:

[0007] Step 1: Read the original question-and-answer samples and perform segmentation processing;

[0008] Step 2: Construct the initial search state using search tree nodes as carriers, and initialize the root node containing the question, path, and evidence pool;

[0009] Step 3: Under the depth-first search framework supported by the explicit stack, the search tree is driven by the controlled action sequence (retrieval, selection, termination) output by the generative large language model.

[0010] Step 4: Extract evidence paragraphs and perform quality filtering on the retrieved web pages, and perform backtracking freeze and experience injection based on the path evaluation results;

[0011] Step 5: Output a multi-hop question-and-answer sample in JSONL format containing the complete retrieval trajectory and structured process record.

[0012] Firstly, the specific steps of step 1 above are as follows:

[0013] Step 1.1: Open the QA sample set file and read the QA sample lines sequentially. Each line contains a question field and an answer label field.

[0014] Step 1.2: Divide the samples into pieces to support concurrent processing.

[0015] Secondly, the specific steps of step 2 above are as follows:

[0016] Step 2.1: Define the structure of the search tree nodes and construct the search state with action nodes as carriers. The node structure includes sub-problems to be solved, current search depth, historical action trajectory, evidence accumulation pool, structured process record and learning signal core fields, which are used to completely map the state transitions in the search process.

[0017] Step 2.2: Initialize the root node S0 by filling the original problem text and initial system meta-instructions into the corresponding fields and pushing the root node onto the explicit control stack.

[0018] Thirdly, the specific steps of step 3 above are as follows:

[0019] Step 3.1: Establish a depth-first search framework, perform a depth-first search with the support of an explicit stack, and pop the node from the top of the stack as the current state S. t ;

[0020] Step 3.2: Generate a controlled action sequence, which is generated by the generative large language model based on the current state S. t Output the next action, which is limited to a predefined set of controlled actions. One of them;

[0021] Step 3.3: When the action includes a retrieval action, the retrieval machine is called to obtain several web page texts. The length of the web page texts is truncated and evidence paragraphs are extracted to control the input scale and generate a candidate evidence list. The health of nodes with empty evidence selection is reduced and the entire path is frozen when the health is below the threshold to suppress the spread of irrelevant evidence. Then, the generative language model outputs the selected evidence number and adds the corresponding evidence to the evidence accumulation.

[0022] Step 3.4: When an action includes a termination action, an answer to the question is generated based on accumulated evidence and encapsulated as a candidate path report. The report is structured to include: the final solution, the aggregated evidence context, the complete retrieval and selection action path, and the entire process of prompt words, thereby realizing the transformation from discrete actions to a complete logical chain and providing standardized data objects for subsequent quality assessment.

[0023] Fourthly, the specific steps of step 4 above are as follows:

[0024] Step 4.1: If the maximum depth is reached but the process has not ended, trigger a depth overrun assessment and record the reason for stopping in the candidate path report;

[0025] Step 4.2: For the end path or depth-exceeding path, call the path evaluation module to generate improvement suggestions and redo step numbers, and perform backtracking freeze based on the redo step numbers. At the same time, inject the improvement suggestions as experience into sibling nodes at the same level to drive differentiated exploration.

[0026] Fifthly, the specific steps of step 5 above are as follows:

[0027] Step 5.1: Encapsulate the generated data according to the aggregated output structure, which includes the original question, standard answer labels, and a path report list consisting of multiple exploration paths;

[0028] Step 5.2: Convert the dynamic search trajectory into static structured fields and output a multi-hop question-and-answer sample in JSONL format containing complete search reasons, selection logic, and evidence numbers.

[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0030] 1. Significantly improves the logical interpretability and end-to-end traceability of question-and-answer datasets. Unlike traditional black-box datasets that only contain "question-answer," the samples output by this invention fully preserve the entire lifecycle of reasoning from the initial question to the final answer. By encapsulating the retrieval query sequence, evidence context, and structured process records containing retrieval reasons and selection logic in the output, each piece of generated training data possesses a clear thought chain. This not only helps downstream models learn complex reasoning logic but also provides precise traceability for manual verification and error analysis.

[0031] 2. This invention achieves high controllability and reproducibility of generative large language models in complex tasks. Addressing the issues of illusions or uncontrollable formatting in large language models, this invention uses a controlled action state machine to strictly constrain the model's output space to a predefined set of executable instructions (retrieval, selection, termination). The system forces the model to make decisions according to a fixed action format, shielding it from unstructured free text output. This mechanism ensures the stability of the automated process, transforming the construction process from relying on random model generation into a parsable, verifiable, and deterministic engineering process.

[0032] 3. Significantly improves evidence coverage and path search success rate under limited computational budget. This invention employs a depth-first search framework based on an explicit stack, combined with constraints on the maximum number of branches and the upper limit of the number of candidates. This design allows the system to explore multiple potential reasoning paths within a fixed budget, effectively avoiding overall generation interruption caused by the failure of a single path retrieval. Through parallel exploration with multiple solutions to a single problem, more diverse combinations of evidence can be discovered, significantly improving the ability to solve complex multi-hop problems and the richness of the dataset.

[0033] 4. Effectively suppresses the spread of irrelevant information, reducing noise interference and redundant computational costs. A dynamic node health assessment mechanism and path freezing pruning strategy are introduced. The system monitors search quality in real time. Once consecutive invalid searches (such as empty selections) are detected, causing the health value to fall below the threshold, a backtracking freeze is immediately triggered to forcibly stop the invalid exploration of that branch. This not only prevents irrelevant documents from polluting the context window and ensures the purity of the final generated answer, but also significantly reduces invalid calls to external search engine APIs, lowering the computational and economic costs of large-scale data construction.

[0034] 5. It possesses adaptive error correction capabilities, improving the convergence efficiency of the search process. This invention constructs a unique path evaluation feedback and experience injection closed loop. When a path exploration fails or the depth limit is exceeded, the system does not simply discard it, but uses a large model to generate improvement suggestions (experience signals) and dynamically injects these lessons learned into the prompts of sibling nodes at the same level. This mechanism enables subsequent search actions to avoid known error patterns, thereby quickly converging to the correct answer, significantly outperforming memoryless random attempts.

[0035] 6. Maintaining large-scale industrial-grade concurrent processing ensures high throughput and stability of task execution. For massive data construction scenarios, a semaphore-based sharding concurrency control architecture was designed. This architecture not only fully leverages the advantages of multi-threading to improve construction throughput but also effectively manages the access frequency limits of external APIs. Simultaneously, it supports file-based streaming output and breakpoint resume functionality, ensuring that generated data is not lost in the face of network fluctuations or unexpected interruptions, greatly improving the engineering robustness of large-scale dataset production tasks. Attached Figure Description

[0036] Figure 1 This is the overall logical flowchart of the method of the present invention. It details the entire lifecycle processing flow from inputting the original question-and-answer samples to the final output of the structured JSONL file, covering the logical connections between initialization, loop search, dynamic evaluation, and output modules.

[0037] Figure 2 This is a schematic diagram of the search tree expansion process based on an explicit stack. It intuitively illustrates the specific application of depth-first search in node expansion, including popping nodes, initialization, pushing child nodes onto the stack, and the control logic for search depth.

[0038] Figure 3 This is a state machine interaction sequence diagram of controlled actions between the generative language model and the system controller. It reveals the synchronous and collaborative mechanism between the model's generation of "retrieval-selection-end" actions and the system's execution of webpage crawling and numbering feedback.

[0039] Figure 4 This is a schematic diagram of the key data structures involved in this invention. It defines the class diagram and core fields of ActionNode (node ​​structure), PathReport (report structure), and OutputSample (aggregate output structure), reflecting the accumulation and transformation relationship of data during the search process. Detailed Implementation

[0040] The present invention will be further described below with reference to the accompanying drawings and embodiments. It should be noted that the described embodiments are only intended to facilitate the understanding of the present invention and do not constitute any limitation thereof.

[0041] This invention aims to address the problems of missing logical chains, large fluctuations in construction quality, and untraceable reasoning processes in existing question-and-answer datasets by proposing a multi-hop retrieval question-and-answer dataset generation method based on a large language model. This invention introduces a depth-first search framework and a controlled action state machine to achieve automatic discovery and dynamic error correction of retrieval trajectories, providing a complete data construction framework.

[0042] like Figure 1 Figure 2 As shown, the present invention proposes a method for generating multi-hop retrieval question-answering datasets based on a large language model, comprising:

[0043] Step 201: Read the original question and answer samples and perform fragmentation. Use semaphore mechanism to fragment the samples to achieve concurrent rate limiting between external search engine interface and LLM call.

[0044] Step 2011: Read the QA sample set file. The system loads the original data file containing "question-answer" pairs from the specified path and reads the QA sample lines sequentially. The system will perform format validation to ensure that each sample line contains the question field and the standard answer label.

[0045] Step 2012 uses a semaphore mechanism to slice the samples to support multi-threaded concurrency. By setting a concurrency threshold, the number of concurrent calls to the external search engine interface and the large language model is limited to effectively prevent the API access frequency limit from being triggered. At the same time, it supports a file output mode, writing the results of different slices to different files so that the task can be resumed from the breakpoint after interruption.

[0046] Step 202: Construct the initial search state with search tree nodes as carriers, model the multi-hop retrieval process as a variant of the Markov decision process, and complete the initialization of the data structure.

[0047] Step 2021: To fully map the state transitions during the search process, construct the search state using the ActionNode data structure, and define the search state as N. t The node state at time t is formally defined as a quadruple, corresponding to the core fields of ActionNode:

[0048] N t ={Q,Path t ,Refs t Prompts t Learning t};

[0049] Where Q is the original problem, and Path is the path. t For the sequence of actions that have been executed, Refs tFor the current evidence accumulation pool, Prompts t The tracking message for the current node, Learning t This serves as a learning signal based on feedback from the previous round of assessments.

[0050] Step 2022: Root node initialization and stack push. The system constructs an initial state S0 (i.e., root node N0) for each question, fills the corresponding fields with the original question text and initial system meta-instructions, sets the evidence pool and path to empty, and then pushes the root node onto the explicit control stack to prepare for starting depth-first traversal.

[0051] Step 203: Under the depth-first search framework supported by the explicit stack, the search tree is expanded by the controlled action sequence (retrieval, selection, termination) output by the generative large language model.

[0052] Step 2031: Establish a depth-first search framework, perform depth-first search with the support of an explicit stack, and pop the node from the top of the stack as the current state S. t And determine whether to continue expanding based on the current depth;

[0053] Step 2032: Generate a controlled action sequence, using the large language model as an intelligent decision-making agent, based on the current node state N. t The probability distribution generates the next action A t+1 The following probability generation formula is used to ensure process controllability:

[0054] ;

[0055] The action space is strictly constrained to a predefined set. Within this set, the model output must conform to the format defined to ensure the parsability of the process; where Prompt(·) is the prompt word state serialization function, responsible for serializing the current search tree node N. t The multidimensional internal states are concatenated according to preset prompt word templates and transformed into an input text stream that can be parsed by a large language model; Logits LLM (·) represents the original logarithmic output function of the large language model. After the input sequence passes through the deep neural network of the large language model for forward propagation, it is applied to the predefined controlled action space in the final output layer. The system generates an unnormalized prediction score for each candidate action; Softmax(·) is a normalization exponential function that exponentially normalizes and normalizes the multiple raw prediction scores output by the model, mapping these real scores to probability values ​​in the interval (0,1), and ensuring that the sum of the probabilities of all candidate controlled actions is strictly 1. Through the nested use of these functions, the system ultimately derives the prediction score for the current state N. t Below, the model selects a specific action A. t+1The mathematical probability P(A) t+1 |N t ).

[0056] Step 2033: Perform retrieval, evidence selection, and health assessment circuit breaking. When the model output is a retrieval action, perform the following sequential operations:

[0057] External retrieval and extraction: Call an external retrieval tool to obtain several web page texts, perform paragraph extraction based on length truncation, and generate a list of candidate evidence with numbers;

[0058] Health assessment: A scoring function is used before selection. Real-time monitoring of search quality to suppress the spread of irrelevant evidence:

[0059] ;

[0060] in, For initial health, This is an indicator function; if the evidence selected in the current step is empty, then... The penalty factor is 1, and λ is an empty selection factor. This represents the depth penalty coefficient.

[0061] Freeze and select: When When the path is in a state of emergency, the system executes a backtracking freeze to forcibly stop the path. If the path is healthy, the LLM outputs the selected evidence number, and the system parses it and adds the corresponding evidence to the evidence accumulation pool.

[0062] Step 2034 When the action includes a termination action, an answer to the question is generated based on accumulated evidence. At this time, the system creates a candidate path report (PathReport) object, which encapsulates the final solution, the aggregated evidence context, the complete retrieval and selection action path, and the entire process prompt word trajectory, thereby realizing the transformation from discrete ActionNode to a complete logical chain and providing standardized data objects for subsequent quality assessment.

[0063] Step 204: Perform evidence paragraph extraction and quality filtering on the retrieved web pages, and perform backtracking freeze and experience injection based on the path evaluation results.

[0064] Step 2041 If the current depth reaches the maximum preset depth but the end action is not output during the search process, the depth limit assessment is triggered. The system will record the reason for stopping (such as "depth_limit_reached") in the candidate path report and mark the path as failed or to be optimized.

[0065] Step 2042: Invoke the path evaluation module to generate improvement suggestions (experience vectors) for failed paths. through operators Inject experience into sibling nodes at the same level In the prompt sequence, implement prior avoidance across branches:

[0066] ;

[0067] Step 205: Structured output. Transform the data objects in memory into multi-hop question-and-answer samples in JSONL format, which contain complete retrieval trajectories and structured process records.

[0068] Step 2051 The system encapsulates the generated data in memory according to the aggregated output structure (OutputSample). An OutputSample object contains the original question, standard answer label, and a path_report_list (composed of multiple PathReports generated in step 2034).

[0069] Step 2052: Serialize the encapsulated object into JSONL format and output it to disk. The output sample explicitly includes a structured process report (struct_procedure_report), forming a traceable chain of evidence in the following form:

[0070] Reasoning steps In order to solve the subproblems Execute query From the five retrieved paragraphs, the one numbered 1 was selected. Evidence, based on the fact that it contains key entities. .

[0071] II. System Composition and Module Responsibilities Figure 3 )

[0072] Input processing and concurrent sharding: reading JSONL samples, sharding, concurrent scheduling, and batch writing.

[0073] Search tree management: The search state is represented by nodes, and DFS traversal is performed using an explicit stack.

[0074] Action decision and prompt word state machine: The generative language model produces the next action according to a fixed action format, and the system parses the action and feeds back the external search results in the form of numbers.

[0075] Search and webpage extraction: Perform searches, crawl webpage HTML, and extract plain text.

[0076] Evidence paragraph extraction and answer generation: Extract relevant paragraphs from web page text; generate the final answer based on accumulated evidence.

[0077] Evaluation and backtracking: Evaluate successful or out-of-limit paths, generate suggestions and redo steps, write them into the node learning signal, and freeze pruning.

[0078] III. Key Data Structures Figure 4 )

[0079] The specific components of the data structure are as follows:

[0080] 1. Node Action State Structure (ActionNode)

[0081] The subproblem to be solved (query_to_solve): The specific problem target that needs to be solved by retrieval or reasoning in the current search step.

[0082] Search depth: Records the level of the current node in the search tree, used to trigger depth truncation protection.

[0083] Action path: Records all historical actions (retrieval, selection, or termination) executed from the root node to the current node in chronological order.

[0084] Evidence Accumulator Pool (refs_accumulator): Stores all evidence fragments that have been acquired and confirmed as valid for the current path, along with their corresponding numbers.

[0085] Structured process report (struct_procedure_report): Records the metadata of the current step, including the search rationale for model generation, selection logic, and timestamps.

[0086] Learning signals: Empirical data fed back by the path evaluation module, used to guide the action correction of subsequent nodes.

[0087] Freeze flag: A Boolean variable that is set to true to stop the expansion of a branch when the path is determined to be invalid or duplicate.

[0088] Node health: A quantitative score calculated using a formula, used to dynamically measure the quality of the current search path.

[0089] 2. Candidate path report structure: PathReport (success / limit exceeded status)

[0090] For successful paths, the following fields are included:

[0091] Final solution: The final answer to the original problem generated based on accumulated evidence.

[0092] Assembled context: A block of text that rearranges scattered pieces of evidence according to reasoning logic.

[0093] Search path overview (search_path): The complete chain of retrieval and selection actions that the successful sample underwent.

[0094] Complete Explore Prompt: Records all raw Prompt streams that interacted with the LLM during the generation of this sample, which are used for subsequent model fine-tuning.

[0095] For paths exceeding depth limits, the following fields are included:

[0096] Stop reason (stop_reason): uniformly marked as "reached maximum search depth (depth_limit_reached)", used for subsequent analysis of the boundary conditions of retrieval failure.

[0097] 3. Problem-level aggregated output structure OutputSample

[0098] Original question (query): The initial complex question in the input dataset.

[0099] Standard answer label (ans_label): Used as the original answer to verify the accuracy of the generated results.

[0100] The path report list contains all the valid paths explored for this problem and their corresponding structured proofs, supporting data augmentation for "multiple paths for one problem".

[0101] 4. Key parameters and configurations

[0102] In this embodiment, the configurable parameters include at least: maximum depth, maximum branch, maximum number of candidates, maximum concurrency, number of fragments, and page truncation length.

[0103] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0104] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A method for generating multi-hop retrieval question-answering datasets based on a large language model, characterized in that, Includes the following steps: Step 1: Read the original question-and-answer samples and perform segmentation processing; Step 2: Construct the initial search state using search tree nodes as carriers, and initialize the root node containing the question, path, and evidence pool; Step 3: Under the depth-first search framework supported by the explicit stack, the generative large language model outputs a search tree driven by a sequence of controlled actions, including retrieval, selection, and termination. Step 4: Extract evidence paragraphs and perform quality filtering on the retrieved web pages, and perform backtracking freeze and experience injection based on the path evaluation results; Step 5: Output a multi-hop question-and-answer sample in JSONL format containing the complete retrieval trajectory and structured process record.

2. The method according to claim 1, characterized in that, The specific implementation of step 1 includes the following steps: Step 1.1: Open the question and answer sample set file. The question and answer sample is also called the QA sample. Read the QA sample lines sequentially. Each line contains a question field and an answer label field. Step 1.2: Divide the samples into pieces to support concurrent processing.

3. The method according to claim 1, characterized in that, The specific implementation of step 2 includes the following steps: Step 2.1: Define the structure of the search tree nodes and construct the search state with custom data structure action nodes as carriers to fully map the state transitions in the search process; Step 2.2, Root node initialization: Fill the original problem text and initial system meta-instructions into the corresponding fields.

4. The method according to claim 1, characterized in that, Step 3 includes the following steps: Step 3.1: Establish a depth-first search framework and execute depth-first search with the support of an explicit stack; Step 3.2: Generate a controlled action sequence, and output the next action from the generative large language model. The action is limited to one of the predefined controlled actions. Step 3.3: When the action includes a retrieval action, the retrieval machine is called to obtain several web page texts. The length of the web page text is truncated and evidence paragraphs are extracted to control the input scale and generate a candidate evidence list. The health of nodes with empty evidence selection is reduced and the entire path is frozen when the health is below the threshold to suppress the spread of irrelevant evidence. Then, the generative language model outputs the selected evidence number and adds the corresponding evidence to the evidence accumulation. Step 3.4: When the action includes an ending action, generate an answer to the question based on accumulated evidence and encapsulate it as a candidate path report with a custom data structure.

5. The method according to claim 1, characterized in that, Step 4 includes the following steps: Step 4.1: If the maximum depth is reached but the process has not ended, trigger a depth overrun assessment and record the reason for stopping in the candidate path report; Step 4.2: For the ending path or the path exceeding the depth limit, call the path evaluation module to generate improvement suggestions and redo step numbers, and perform backtracking freeze and experience injection based on the redo step numbers, thereby affecting the action generation of subsequent nodes.

6. The method according to claim 1, characterized in that, Step 5 includes the following steps: Step 5.1: Encapsulate the generated data by aggregating the output structure according to the custom data structure; Step 5.2: Convert the dynamic search trajectory into static structured fields and output multi-hop question-and-answer samples in JSONL format.