Question and answer data synthesis method and system based on plan driving
By generating synthesis specifications and decomposing them into independent execution plan items, and combining them with quality verification and correction mechanisms, the problems of target drift and quality instability in question-and-answer data synthesis are solved, and high-quality, controllable question-and-answer data generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI FEISHU INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2026-04-10
- Publication Date
- 2026-05-12
AI Technical Summary
Existing question-and-answer data synthesis methods based on large language models lack overall planning for macro synthesis goals and fine-grained constraints on micro generation processes. This results in unpredictable distribution of question types and difficulty in the data output, and fails to effectively correct the problem of substandard quality of individual question-and-answer data, leading to computational waste and unstable dataset quality.
By generating synthesis specifications, decomposing them into multiple independent execution plan items, and performing quality verification and correction on each plan item, a local and global governance closed loop is established to ensure that the generation process of question and answer records is controlled and of consistent quality.
It achieves the generation of high-quality structured question-and-answer data that strictly conforms to the initial multidimensional constraints, reduces computing power consumption, and improves the data qualification rate and controllability of the generation process.
Smart Images

Figure CN122019731A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a plan-driven question-and-answer data synthesis method and system. Background Technology
[0002] With the continuous expansion of large-scale language model application scenarios, the demand for high-quality, controllable, and large-scale question-answering datasets for vertical domain model training has increased dramatically.
[0003] In existing plan-driven question-answering data synthesis workflows, an end-to-end automatic generation approach based on large language models is typically used to expand the data scale. This approach mainly involves directly segmenting the original input document and inputting it into a large language model along with fixed, common prompt words. It relies entirely on the large language model's own free divergence and probability sampling capabilities to continuously and randomly output a large number of question-answer pairs in a pipeline batch processing manner.
[0004] This end-to-end probability sampling-based generation method lacks overall planning for the macro-level synthesis goals and fine-grained constraints on the micro-level generation process, leading to unpredictable drift in the distribution of question types and difficulty in the output data. Furthermore, when faced with substandard quality individual question-and-answer data points in the model output, this method can only rely on mechanical, blind resampling or outright post-processing filtering, failing to effectively correct specific errors locally. This results in significant computational waste and ultimately makes it difficult to stably deliver high-quality structured datasets that strictly adhere to the initial multidimensional constraints. Summary of the Invention
[0005] This invention provides a plan-driven question-and-answer data synthesis method and system to overcome the deficiencies in the prior art and achieve the generation of structured question-and-answer data that strictly conforms to the initial multidimensional constraints and has a high degree of consistency in quality.
[0006] This invention provides a plan-driven question-and-answer data synthesis method, comprising the following steps: The input user configuration is compiled and processed to generate a synthetic specification; Based on the synthesis specifications and the preset literature set, an execution plan set is generated, which contains multiple execution plan items with independent generation intentions; Generate the question and answer records corresponding to each of the aforementioned execution plan items; Perform quality verification on the question and answer records corresponding to any of the execution plan items; wherein, when it is determined that the quality of the question and answer records meets the preset acceptance conditions, the question and answer records are stored in the final record set; All question-and-answer records in the final record set are aggregated to generate the final dataset.
[0007] This invention also provides a plan-driven question-and-answer data synthesis system, comprising the following modules: The synthesis specification compilation module is used to compile the input user configuration and generate the synthesis specification; An execution plan generation module is used to generate an execution plan set based on the synthesis specification and a preset literature set, wherein the execution plan set contains multiple execution plan items; The question-and-answer synthesis module is used to generate question-and-answer records corresponding to each of the execution plan items. The quality verification and gating module is used to perform quality verification on the question and answer records corresponding to any of the execution plan items; wherein, when it is determined that the quality of the question and answer record meets the preset acceptance conditions, the question and answer record is stored in the final record set; The delivery aggregation module is used to aggregate all question and answer records in the final record set to generate the final dataset.
[0008] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the plan-driven question-and-answer data synthesis method as described above.
[0009] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the plan-driven question-and-answer data synthesis method as described above.
[0010] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the plan-driven question-and-answer data synthesis method as described above.
[0011] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: By compiling and processing the input user configuration to generate a synthesis specification, an unambiguous machine-readable standard is established, eliminating uncertainty at the data production source. Based on the synthesis specification and a pre-defined literature set, an execution plan set containing multiple execution plan items is generated, with each execution plan item corresponding to an independent question-and-answer generation intent. This precisely decomposes the macro-level data synthesis goal into micro-level atomic tasks and completely avoids data distribution drift during the generation process. By independently generating corresponding question-and-answer records for each execution plan item, the single data generation process is strictly controlled, and the source intent of each generated data is clear and explicit. Automated quality verification is performed on each question-and-answer record. If the record quality is substandard and the retry limit has not been reached, the generation and verification steps are repeated based on the correction suggestions generated by the verification. This establishes a micro-level closed-loop data quality governance mechanism, significantly improving the final data pass rate and reducing the computational consumption caused by blind retries. After processing all execution plan items, the question-and-answer records in the final record set are aggregated to generate the final dataset, ultimately delivering structured question-and-answer data that strictly conforms to the initial multi-dimensional constraints and has a high degree of consistency in quality. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0013] Figure 1 This is the overall architecture diagram of the plan-driven question-and-answer data synthesis system provided by the present invention.
[0014] Figure 2 This is the main flowchart of the plan-driven question-and-answer data synthesis method provided by the present invention.
[0015] Figure 3 This is a flowchart illustrating the plan-driven question-and-answer data synthesis method provided by the present invention.
[0016] Figure 4 This is a flowchart illustrating the synthetic specification compilation method provided by the present invention.
[0017] Figure 5 This is a flowchart illustrating the coverage optimization-driven execution plan generation method provided by the present invention.
[0018] Figure 6 This is a schematic diagram of the structure of the plan-driven question-and-answer data synthesis system provided by the present invention.
[0019] Figure 7This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0021] It should be noted that in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. The terms "upper," "lower," etc., indicating orientation or positional relationships according to the accompanying drawings, are only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the system or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0022] The terms "first," "second," etc., used in this invention are used to distinguish similar objects, not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0023] The following is combined Figures 1 to 7 This invention describes the plan-driven question-and-answer data synthesis method, system, electronic device, storage medium, and computer program product provided by this invention.
[0024] Reference Figure 1 , Figure 1 This is an overall architecture diagram of the plan-driven question-answering data synthesis system provided by the present invention. The embodiments of the present invention are combined with... Figure 1The overall architecture and data flow process of a plan-driven question-and-answer data synthesis system are explained.
[0025] Combination Figure 1 As shown, this system receives external task requests and simultaneously obtains the input document set and synthesis specification configuration. Figure 1 Solid arrows indicate the flow of data within the system, while dashed arrows indicate the flow of control signals or status reports. The overall architecture of this system is divided into three main parts: the core coordinator, the control plane, and the execution plane.
[0026] The core coordinator receives external task requests and injects global configurations into the control plane and execution plane. Through a global state interaction mechanism, the core coordinator continuously monitors and coordinates the operational status between the control plane and the execution plane.
[0027] The control plane is responsible for processing front-end input data and formulating the generation plan. The control plane includes a specification builder, a coverage optimizer, and an execution planner. The specification builder receives the composition specification configuration from external input, performs standardization and validation operations on the configuration, and outputs a machine-readable composition specification. The coverage optimizer receives a set of input documents, performs text segmentation and value assessment operations on the document set, and outputs a coverage optimized set containing high-value text segments. The execution planner receives the composition specification and the coverage optimized set, performs quota decomposition and task instantiation operations on the coverage optimized set according to the constraints set in the composition specification, and generates and outputs a list of execution plans. The execution plan list is then passed down to the execution plane.
[0028] The execution plane is responsible for data production and quality control according to the established plan. It includes a question-and-answer synthesizer, a quality verification and gating module, and a storage and caching module. The question-and-answer synthesizer reads each execution plan item from the execution plan list, driving the large language model to generate corresponding single-question question-and-answer records. The quality verification and gating module receives the single-question question-and-answer records and performs multi-dimensional quality assessment operations on them. When the quality assessment fails, the quality verification and gating module generates structured correction suggestions and feeds them back to the question-and-answer synthesizer to trigger a regeneration operation, thus constructing a local governance closed loop. When the quality assessment passes or the number of retries reaches the system limit, the quality verification and gating module outputs a final decision of acceptance or rejection.
[0029] When the quality verification and gating module still outputs a rejection decision after multiple retries, it reports a failure signal to the core coordinator. Upon receiving the failure signal, the core coordinator, in conjunction with the global state, sends a replanning signal to the execution planner on the control plane, thereby constructing a global governance closed loop.
[0030] The storage and caching module receives single-question answer records carrying acceptance decisions and performs a persistent storage operation on these records. After all execution plan items have been processed, the system performs an aggregation operation and outputs the final question-and-answer dataset and other auxiliary output files to the outside world.
[0031] Next, refer to Figure 2 , Figure 2 This is the main flowchart of the plan-driven question-and-answer data synthesis method provided by the present invention. The embodiments of the present invention are combined with... Figure 2 The main process of a plan-driven question-and-answer data synthesis system is described in detail. In the following embodiments, the processor is used as the execution entity for elaboration.
[0032] like Figure 2 As shown, the main system flow is logically divided into four parallel execution swimlanes: the core coordinator swimlane, the plan generation swimlane, the execution verification swimlane, and the storage swimlane. The processor advances the task sequentially according to the following steps: The processor receives an external task request, initiates the execution flow in the core coordinator swimlane, and proceeds to step S1. In step S1, the processor performs task initialization operations and synchronously initializes the global state of the task being generated.
[0033] After initialization, the processor executes step S2 in the plan generation side lane. In step S2, the processor performs a composition specification parsing operation on the input configuration information and a coverage optimization operation on the input document set, ultimately generating a machine-readable composition specification and a coverage optimization set that identifies high-value content. Subsequently, the processor executes step S3. In step S3, the processor generates an execution plan list based on the composition specification and the coverage optimization set. The execution plan list contains multiple specific execution plan items. After generating the execution plan list, the processor distributes the execution plan items downstream.
[0034] The processor receives the issued execution plan item in the verification swimlane and executes step S4. In step S4, the processor performs a task distribution operation and loads the corresponding document context information based on the content of the current execution plan item. Next, the processor executes step S5. In step S5, the processor performs an atomic question-answer synthesis operation, driving the large language model to generate a single question-answer record by combining the context information. Subsequently, the processor executes step S6. In step S6, the processor performs a multi-dimensional quality verification operation on the generated question-answer record.
[0035] After completing the quality verification, the processor proceeds to decision node D1. At decision node D1, the processor determines whether the verification result of the question-and-answer record is satisfactory.
[0036] When the processor determines that the verification result at decision node D1 is unsuccessful, the processor proceeds to decision node D2. At decision node D2, the processor checks whether the number of retries already performed for the current execution plan item is less than the preset maximum retry limit. If the number of retries is less than the maximum retry limit, the processor generates a corresponding correction suggestion and returns to step S5 along the local closed-loop path. In step S5, the processor re-executes the atomic question-and-answer synthesis operation based on the correction suggestion. When the number of retries reaches the maximum retry limit, the processor executes step S7. In step S7, the processor determines that an unrecoverable anomaly has occurred, marks the current execution plan item as rejected, and reports a failure signal carrying systemic failure information to the upstream core coordinator swimlane.
[0037] After receiving a failure signal in the core coordinator swimlane, the processor executes step S8. In step S8, the processor performs a strategic decision operation, conducting a comprehensive evaluation based on the failure signal and the current global state. Next, the processor proceeds to decision node D3. At decision node D3, the processor determines whether the current global state meets the conditions for triggering a replanning operation. If the conditions are met, the processor returns to step S3 along the global closed-loop path and regenerates a supplementary execution plan item. If the conditions are not met, the processor confirms the rejection status of the current execution plan item and forwards it downstream to decision node D4.
[0038] When the processor determines that the verification result at the preceding decision node D1 is passed, the processor directly proceeds to the storage-side swimlane and executes step S9. In step S9, the processor performs atomic storage operations on the question-and-answer records in the passed state. The atomic storage operations include performing normalization processing on the question-and-answer records, calculating the content hash value, and performing a secure disk write operation. After completing the storage write, the processor proceeds to decision node D4.
[0039] At decision node D4, the processor checks whether all execution plan items in the execution plan list have been processed. If there are any unprocessed execution plan items, the processor loops back to step S4 to extract and process the next execution plan item. When all execution plan items have been processed and the final state is reached, the processor executes step S10 in the storage-side swimlane. In step S10, the processor performs an aggregation and delivery operation on all successfully stored question-and-answer records, generating and outputting the final question-and-answer dataset, thus ending the entire external task request processing flow.
[0040] Within the framework of the main process of the system described above, the following section elaborates on the specific details of the plan-driven question-and-answer data synthesis performed by this system, combining specific methodological steps.
[0041] This invention provides a plan-driven question-and-answer data synthesis method. The execution entity of the method provided in this invention includes, but is not limited to, a server, a cloud computing platform, or a computer device equipped with a processor and memory. In the following embodiments, a processor is used as the execution entity for description.
[0042] In this embodiment of the invention, after receiving an external task request, the processor executes a plan-driven question-and-answer data synthesis method to automatically synthesize a high-quality question-and-answer dataset from a preset document set. The document set serves as the knowledge source for generating the question-and-answer data, and includes one or more structured long documents, such as technical white papers, textbook chapters, academic papers, or internal knowledge base documents. (See also...) Figure 3 , Figure 3 This is a flowchart illustrating the plan-driven question-answering data synthesis method provided by the present invention, as shown below. Figure 3 As shown, the method includes steps 110 to 150: Step 110: Compile the input user configuration to generate a synthesis specification.
[0043] The processor receives user configurations provided by the user or the upper-level system. The user configuration describes the macro-level objectives and constraints of this plan-driven question-answering data synthesis task. Information included in the user configuration covers the total number of questions, the target distribution ratio of question types and difficulties, the segment ratio, the lower limit for definitions, the upper limit for retry counts, and quality verification rules. The user configuration is in a machine-readable structured data format such as JSON or YAML.
[0044] The processor performs compilation processing on the user configuration. Compilation processing refers to transforming the user configuration from a high-level, potentially ambiguous or incomplete description of intent into a precise, machine-readable, and logically consistent internal specification. Compilation processing includes at least validating the validity of each parameter in the user configuration, detecting and handling parameter conflicts, and completing any missing parameters.
[0045] After compilation, the processor generates a composition specification. The composition specification is an immutable internal system object containing all validated and standardized constraint parameters. It serves as the direct input for generating the subsequent execution plan set. After generating the composition specification, the processor also performs normalization and serialization processing and calculates a hash value to generate a composition specification fingerprint. This fingerprint serves as a unique identifier for the composition specification and is used for subsequent version locking and lineage association.
[0046] It should be noted that in this step, the processor only performs compilation processing on the user-configured data; it does not perform segment selection operations on the literature set, nor does it generate an execution plan. The output of this step is the synthesis specification and the corresponding verification report, ensuring that subsequent planning steps are based on the defined constraint boundaries.
[0047] Step 120: Based on the synthesis specifications and the preset literature set, generate an execution plan set, which contains multiple execution plan items with independent generation intentions.
[0048] After obtaining the synthesis specification, the processor, in conjunction with a pre-defined set of literature, generates an execution plan set. The process of generating the execution plan set involves progressively decomposing the macroscopic objectives defined in the synthesis specification into a series of microscopic, specific atomic tasks.
[0049] In one embodiment, the process of generating a set of execution plans includes two sub-stages.
[0050] The first sub-stage is coverage optimization. The processor analyzes and trims the documents in the document collection, selecting paragraphs with high information density and question-setting value to form a coverage set. The coverage set is an ordered, cross-document list of paragraph indexes, with each element in the list accompanied by metadata such as the source chapter and estimated difficulty. The purpose of coverage optimization is to concentrate subsequent generation resources on the content areas that are "most worthy of question setting," avoiding wasting generation quotas on low-value or repetitive text content.
[0051] The second sub-stage is execution plan decomposition. The processor, in conjunction with the constraints defined in the synthesis specification, such as the proportion of question types, difficulty distribution, and total number of questions, allocates the macro-level statistical quotas to specific segments in the coverage set, and instantiates each allocated quota position as an execution plan item.
[0052] The execution plan set consists of multiple execution plan items. Each execution plan item is an atomic execution unit, corresponding to an independent single-question question and answer generation intent. Each execution plan item contains at least the following information: source document identifier, used to specify the document on which this generation is based; paragraph identifier, used to specify the specific paragraph in the document; target question type, used to specify the type of question to be generated; target difficulty, used to specify the difficulty level of this generation; and local random seed, used to control operations involving randomness in this generation process.
[0053] Once the execution plan set is generated, it is frozen. "Frozen" means that in subsequent execution phases, the processor consumes each execution plan item in the set in a read-only manner, without modifying any of the execution plan items in place. If adjustments to the plan are necessary in later execution phases, the processor generates new versions of the execution plan items, rather than modifying the frozen ones. This mechanism ensures that each subsequent question-and-answer generation action has a clear source and constraint, thus preventing target drift and unbounded exploration at the architectural level.
[0054] Step 130: Generate the Q&A records corresponding to each execution plan item.
[0055] The processor sequentially retrieves the next execution plan item to be processed from the frozen execution plan set, and loads the corresponding text paragraph and context information from the document set according to the source document identifier and paragraph identifier recorded in the execution plan item, providing complete input for the question-and-answer generation process.
[0056] The processor constructs prompt words based on the target question type, target difficulty, and constraints specified in the current execution plan, combined with the loaded text paragraphs. The processor inputs the prompt words into the large language model, driving the large language model to generate a question-and-answer record.
[0057] A question-and-answer record is a structured data record, containing at least a stem, an answer, and an evidence field. The stem is the question text generated by the large language model. The answer is the response to the stem generated by the large language model. The evidence field records the location information of the original text supporting the answer, pointing to specific sentences within the text paragraph. During the generation phase, the processor requires the large language model to bind the evidence field, establishing a traceable link between the answer and the original text content in the literature collection.
[0058] When executing this step, the processor strictly adheres to the principle of "generating only one question at a time." This means the processor performs an independent question-and-answer generation operation for each execution plan item, generating one question-and-answer record. This atomic generation method ensures that the behavior of each question-and-answer record can be independently located and diagnosed, facilitating subsequent quality verification and issue tracing.
[0059] Step 140: Perform quality verification on the Q&A records corresponding to any execution plan item.
[0060] After obtaining the question-and-answer records, the processor performs automated, multi-dimensional quality checks. The purpose of the quality checks is to objectively evaluate the quality of the question-and-answer records and make subsequent processing decisions based on the evaluation results. The quality checks produce a structured decision result, which includes the scores for each verification dimension and the corresponding reason labels.
[0061] Based on the decision, the processor makes a processing decision for the current execution plan item. The processing decision includes at least the following three scenarios: Scenario 1: When the quality of the question and answer records meets the preset acceptance conditions, the question and answer records will be stored in the final record set.
[0062] The preset acceptance criteria mean that all hard quality indicators defined in the synthesis specification have reached their corresponding thresholds, and the decision result does not contain fatal error labels. When the question and answer record meets the acceptance criteria, the processor marks the status of the question and answer record as accepted and passes the question and answer record to the storage stage to be stored in the final record set.
[0063] Scenario 2: When the quality of the question and answer record is determined to be unacceptable and the number of times the question and answer record is regenerated has not reached the preset retry limit, the steps of generating the question and answer record and performing quality verification are repeated based on the correction suggestions generated by the verification.
[0064] When the quality of the question and answer record does not meet the acceptance criteria, but the error label contained in the judgment result is of a locally repairable type, and the number of retries executed for the current execution plan item has not reached the preset retry limit, the processor triggers a local closed-loop correction process.
[0065] In the local closed-loop correction process, the processor generates structured correction suggestions based on the cause labels in the decision results. These suggestions contain targeted guidance for the large language model's next generation behavior, such as instructing the model to reread specific sentences to ensure the answer is consistent with the original facts, or instructing the model to adjust the output format to meet preset structured requirements. The processor appends these correction suggestions to the prompts when generating question-and-answer records next time, as additional guidance.
[0066] Subsequently, the processor repeats the operations of generating question-and-answer records in step 130 and performing quality verification in step 140 for the same execution plan item. This repeated execution is a targeted correction guided by the correction suggestions, rather than indiscriminate blind retries. The processor increments the retry count with each retry until the question-and-answer records meet the acceptance conditions and enter scenario one, or the retry count reaches a preset retry limit and enters scenario three.
[0067] Scenario 3: When the quality of the question and answer record is determined to be unacceptable and the number of times the question and answer record is regenerated reaches the preset retry limit.
[0068] When all retries for the current execution plan item have been exhausted, and the query-response record still does not meet the acceptance criteria, the processor marks the current execution plan item as rejected. The processor records the rejection status, its corresponding failure reason, and related metrics for subsequent statistical analysis and audit traceability. In one embodiment, after marking the execution plan item as rejected, the processor also reports a failure signal to the global scheduling layer, which then evaluates whether adjustments to the subsequent execution strategy are necessary.
[0069] After the processor completes processing any of the three scenarios described above for the current execution plan item, it continues to retrieve the next execution plan item to be processed from the execution plan set and repeats the operations of steps 130 and 140 until all execution plan items in the execution plan set have been processed, that is, all execution plan items have entered the accept state or the reject state.
[0070] Step 150: Aggregate all question and answer records in the final record set to generate the final dataset.
[0071] When the processor determines that all execution plan items in the execution plan set have reached the final state, the processor performs aggregation and delivery operations. The final state means that each execution plan item has been marked as either accepted or rejected.
[0072] The processor reads the question-and-answer records corresponding to all execution plan items marked as "acceptable" from the final record set. Following a predefined order in the synthesis specification, the processor arranges and aggregates all read question-and-answer records to generate a standard dataset file, which is the final dataset. Each question-and-answer record in the final dataset contains the question stem, answer, evidence citations, and source annotations. The content distribution of the final dataset strictly adheres to the constraints defined in the synthesis specification.
[0073] While generating the final dataset, the processor also calculates quantitative metrics for the entire task process and generates a task-level lineage record file. Quantitative metrics include at least the total number of planned items, the total number of accepted items, the total number of rejected items, and the pass rate. The lineage record file records complete correlation information for the task from initial configuration to final output, enabling traceability of the task's output source, auditability of the generation process, and replayability of synthesis experiments. The processor organizes the final dataset, quantitative metrics, and lineage record file together as the final deliverable of the task.
[0074] Through steps 110 to 150 above, this embodiment of the invention implements a plan-driven question-and-answer data synthesis method. The method transforms the originally highly random large language model generation process into a predictable and measurable engineered pipeline by first compiling the user's macro-intention into a synthesis specification before executing any generation operation, then decomposing the synthesis specification into a frozen set of execution plans, and finally executing generation and verification item by item using the execution plan set as read-only instructions. During pipeline execution, the method ensures output quality at the single-question level and effectively reduces invalid attempts by performing quality checks on each question-and-answer record and triggering retry based on correction suggestions when conditions are not met through a local closed-loop mechanism. After all execution plan items are processed, the method aggregates and produces the final dataset and simultaneously generates a lineage record file, enabling the final deliverable to have stable acceptance criteria and version management capabilities. Overall, the method achieves large-scale synthesis and production of question-and-answer data for large language model training while meeting controllable constraints and reliable quality.
[0075] In one embodiment, the step of compiling the input user configuration in step 110 to generate the synthetic specification is further explained.
[0076] The purpose of compilation is to transform the high-level user-provided configuration, which may contain ambiguities, potential conflicts, or incompleteness, into a precise, machine-readable, and logically consistent synthesis specification. Functionally, the compilation process is equivalent to a "configuration compiler," which does not perform segmentation operations on the literature set, nor does it generate an execution plan; it only outputs a valid, frozen synthesis specification and a verification report. The compilation process is formalized as a configuration compilation function: ; in, Configure for users, For the purpose of synthesis specifications, This is a verification report generated during the compilation process. The verification report contains warning or error blocking information generated during compilation.
[0077] Reference Figure 4 , Figure 4 This is a flowchart illustrating the synthetic specification compilation method provided by the present invention, as shown below. Figure 4 As shown, the compilation process includes at least the following steps: Step 210: Receive user configuration, which contains multiple first configuration parameters for describing the goals of plan-driven question-and-answer data synthesis.
[0078] The processor receives user configuration provided by the user or the upper-layer system. The user configuration is in JSON or YAML format. The user configuration contains multiple first configuration parameters, which collectively describe the macro-level objectives and constraints of this plan-driven question-answering data synthesis task. These first configuration parameters cover, but are not limited to, the following categories: target total number of questions, specifying the expected total number of question-answer records to be produced in this synthesis task; question type ratio, specifying the target proportion of different question types in the target total number of questions; difficulty distribution, specifying the target proportion of different difficulty levels in the target total number of questions; paragraph span ratio, specifying the proportion of questions requiring information synthesis across multiple paragraphs to answer; lower limit of interpretation, specifying the minimum degree to which the answer modifies the original text; uncertainty threshold, specifying the critical value of the uncertainty score that triggers strong model routing; maximum number of retries, specifying the maximum number of retries allowed for a single execution plan item; and quality verification rules, defining the scoring thresholds and gating conditions for each quality dimension.
[0079] Step 220: Construct a constraint dependency graph based on the constraint relationships between multiple first configuration parameters.
[0080] Since the multiple first configuration parameters are not independent but rather constrained by each other, the processor needs to explicitly model these constraints as a structured graph before performing conflict detection. The processor treats the multiple first configuration parameters as a set of nodes and the constraints between them as a set of edges, and constructs a constraint dependency graph based on the node set and edge set.
[0081] The constraint dependency graph is formally represented as ,in V It is a set of nodes, where each node corresponds to a first configuration parameter in the user configuration. E Let be the set of edges, where each edge represents a constraint relationship between two first configuration parameters. Constraint relationships include at least one of summation constraints and bounds constraints. A summation constraint is a constraint that the sum of the values of multiple first configuration parameters must satisfy a specific numerical condition, such as the sum of all question type ratio parameters should equal 1. A bounds constraint is a constraint that the value of a first configuration parameter must fall within a numerical range defined by one or more other first configuration parameters, such as the minimum token number parameter not exceeding the maximum token number parameter.
[0082] In one embodiment, the constraint dependency graph is a directed graph, and a path in the directed graph consists of nodes. u Pointing to node v The directed edge representation parameter v The constraint verification or value determination depends on the parameter. u .
[0083] By constructing a constraint dependency graph, the processor transforms the implicit constraint relationships between multiple first configuration parameters into an explicit graph structure, providing a structured analytical foundation for subsequent systematic conflict detection.
[0084] Step 230: Based on the constraint dependency graph, perform conflict detection on multiple first configuration parameters to determine the conflicting second configuration parameters and the conflict type corresponding to each second configuration parameter.
[0085] Based on the constraint dependency graph constructed in step 220, the processor performs systematic conflict detection on multiple first configuration parameters.
[0086] The processor performs a topological sort on the constraint dependency graph, obtaining a topological sort result. The topological sort result specifies a deterministic parameter verification order, ensuring that when verifying any first configuration parameter, all its preceding parameters have already been verified. The processor then performs constraint propagation and repair sequentially according to the topological sort result. Constraint propagation refers to passing the determined values of verified parameters along the directed edges of the constraint dependency graph to their downstream parameters, thereby gradually narrowing the valid value space of downstream parameters. Repair refers to, during propagation, when it is found that the current value of a first configuration parameter does not satisfy the valid value space defined by the propagated constraints, attempting to adjust the value of the first configuration parameter to within the valid range.
[0087] When strongly connected components exist in the constraint dependency graph, the parameters within these components form cyclic dependencies, making it impossible to determine the verification order through unidirectional topological sorting. Therefore, the processor performs a joint solution for the constraints within the strongly connected components. This joint solution is an iterative process where the processor repeatedly propagates constraints within the strongly connected components and attempts to repair them until all parameters in all components satisfy the constraints. If the joint solution fails to converge within a preset number of iterations, meaning there are still parameters that do not satisfy the constraints, the processor determines that there are irreparable conflicts among the parameters involved in the strongly connected components.
[0088] Through the above conflict detection process, the processor identifies all conflicting first configuration parameters from multiple first configuration parameters, marks the conflicting first configuration parameters as second configuration parameters, and determines the conflict type corresponding to each second configuration parameter.
[0089] Conflict types include at least the following four: The first type of conflict is an irreparable conflict. Irreversible conflicts include at least one of the following: mutually exclusive conflicts, illegal value ranges, or empty targets. Mutually exclusive conflicts refer to conflicts where the same question type is simultaneously required to have two mutually exclusive forms. Illegal value ranges refer to situations where the first configuration parameter has a negative proportional value, is not a numeric NaN, the minimum token number parameter has a value greater than the maximum token number parameter, or the threshold is not within the legal range of zero to one. Empty targets refer to situations where the question type set is empty or the proportion of all question types is zero.
[0090] The second type of conflict is resource budget conflict. Resource budget conflict refers to a situation where, based on the current initial configuration parameters, the estimated cost or number of tokens required for this task exceeds the user-set budget limit.
[0091] The third type of conflict is the proportion distribution conflict. Proportion distribution conflicts include at least one of the following: the sum of the proportions of the configuration parameters is not 1, some proportions are missing, or the configuration includes unsupported question types. A sum of proportions not equal to 1 means that the sum of the proportion parameters for all question types or all difficulty levels is not equal to 1. Missing proportions mean that the user has only specified proportions for some question types, omitting the remaining question types. Including unsupported question types means that the user's configuration includes question type identifiers that are not supported by the current version of the system.
[0092] The fourth type of conflict is soft constraint deviation. Soft constraint deviation includes at least one of the following: deviation from the recommended average difficulty or deviation from the recommended coverage. Soft constraint deviation refers to a situation where the value of the first configuration parameter, although logically valid, deviates from the optimal practice range recommended by the system.
[0093] By classifying conflicts according to their severity, the processor lays the foundation for adopting differentiated handling strategies for different conflict types, ensuring the rigor of the compilation process while avoiding unnecessary interruptions to the entire compilation process due to minor deviations.
[0094] Step 240: For any second configuration parameter, execute the corresponding processing strategy according to the conflict type.
[0095] For each second configuration parameter identified in step 230, the processor executes a differentiated processing strategy based on the conflict type corresponding to the second configuration parameter.
[0096] For unrecoverable conflicts, the processor outputs a blocking signal and terminates compilation. When the conflict type of the second configuration parameter is an unrecoverable conflict, the processor determines that there is a fundamental logical error in the user configuration that cannot be corrected by automated means. The processor directly outputs a blocking signal, terminates the current compilation process, and returns an error report containing conflict details to the user, instructing the user to correct the user configuration and resubmit.
[0097] For resource budget conflicts, the processor triggers a degradation scheme. When the conflict type of the second configuration parameter is a resource budget conflict, the processor executes a series of degradation operations in a preset priority order to reduce the estimated resource consumption until the estimated consumption falls back within the budget range. Degradation schemes include at least one of the following: reducing the target total number of problems, reducing the maximum number of retries, tightening strong model routing to reduce the expected call frequency of high-cost models, or rolling back a portion of the difficulty quota to reduce the proportion of high-difficulty problems. The processor records the actual degradation operations executed in the degradation scheme and the parameter changes before and after degradation in a verification report. If, after executing all preset degradation operations, the estimated resource consumption still exceeds the budget limit, the processor escalates the resource budget conflict to an unrecoverable conflict, outputs a blocking signal, and terminates compilation.
[0098] For proportionally distributed conflicts, the processor performs normalization correction or pruning. When the conflict type in the second configuration parameter is a proportionally distributed conflict, the processor selects the corresponding correction method based on the specific subtype of the conflict.
[0099] When the sum of the proportions of the conflicting subtype configuration parameters is not 1, the processor performs normalization correction. The processor uses a priority-based normalization formula to correct the second configuration parameter with the proportional conflict. The normalization formula is as follows: ; in, This is the original ratio value of the i-th second configuration parameter. Let be the corrected ratio value of the i-th second configuration parameter, and n be the total number of configuration parameters participating in the normalization. This is the original ratio value for the m-th configuration parameter. To prevent the removal of zero-adjustment terms. It is a very small positive value used to avoid the denominator being zero when the sum of all the original scale values is exactly zero. Using a normalization formula, the processor precisely corrects the sum of all scale parameters to 1 while maintaining the relative proportions of each parameter. The processor records both the original scale value before correction and the corrected scale value in the verification report.
[0100] When the conflicting subtype contains unsupported question types, the processor performs a pruning process. The processor removes the unsupported question types and their corresponding scaling parameters from the user configuration, and re-performs the normalization correction described above on the remaining scaling parameters. The processor records the removed question type identifiers and the pruning operation in the verification report.
[0101] When the subtype of the conflict is a partial proportion missing, the processor first assigns zero value to the question type with the missing proportion or assigns an initial value according to the recommended weight, and then performs the above normalization correction on all proportion parameters.
[0102] When a negative proportional value or an illegal value (not a numerical value) exists in the second configuration parameter, the processor will escalate the conflict corresponding to the illegal value to an unrecoverable conflict for processing. When the original proportional values of all configuration parameters participating in normalization are zero, the processor triggers the default allocation strategy, which includes distributing the proportion equally among all participating configuration parameters or allocating it according to the system-recommended default weight.
[0103] For soft constraint deviations, the processor records alarm information but does not block compilation. When the conflict type of the second configuration parameter is a soft constraint deviation, the processor only records the name of the deviated parameter, its current value, and the suggested value range in the verification report in the form of alarm information. It does not modify the value of the second configuration parameter, nor does it block the compilation process.
[0104] Through the above-mentioned hierarchical processing strategy, the processor preserves the user's original intent to the greatest extent possible while ensuring the logical closure of the synthesis specification, and performs minimal automatic corrections only when necessary.
[0105] Step 250: When the user configuration is missing a preset key parameter, the default value of the key parameter is estimated and injected based on the statistical characteristics of the input document using an adaptive parameter injection algorithm.
[0106] In practical applications, users may omit certain key parameters in their submitted user configurations. When the processor detects that a preset key parameter is missing from the user configuration, it does not simply fill it with a static constant. Instead, it dynamically estimates and injects default values for the key parameters based on the statistical characteristics of the input document. Key parameters include at least one of the target total number of questions and the maximum number of retries.
[0107] The processor first analyzes the documents in the document collection, extracting multi-dimensional statistical features. These statistical features combine at least three to five of the following characteristics: In terms of scale characteristics, statistical features include the length of valid tokens in the document. Number of documents and the number of candidate paragraphs Valid Token Length for Document This refers to the total number of tokens corresponding to the cleaned and valid content of all documents in the document collection. (Document count) This refers to the number of documents contained in the document collection. (Number of candidate paragraphs) This refers to the total number of candidate paragraphs obtained after the initial segmentation of the document collection.
[0108] In terms of information density features, statistical characteristics include entity density, term density, and key term coverage. Entity density refers to the number of named entities contained per unit text length. Term density refers to the number of specialized terms contained per unit text length. Key term coverage refers to the proportion of unique key terms covered in the document collection relative to the predefined domain terminology.
[0109] In terms of structural features, statistical characteristics include heading level depth, number of chapters, and the existence of structural anchors. Heading level depth refers to the maximum nested level of document headings in the document collection. Number of chapters refers to the total number of independent chapters contained in the document collection. The existence of structural anchors refers to whether there are structural chapters with specific functions, such as "definition," "summary," or "conclusion," in the document collection.
[0110] In terms of complexity features, statistical features include syntactic complexity. Syntactic complexity is a quantitative indicator of the text structure complexity calculated based on linguistic features such as dependency syntax tree depth and the number of clauses.
[0111] In terms of redundancy features, statistical features include paragraph similarity redundancy. Paragraph similarity redundancy refers to the degree of semantic similarity between candidate paragraphs, used to measure the degree of content repetition in a document set, so as to suppress excessive question generation on duplicate text when estimating the target total number of questions.
[0112] Based on the extracted multi-dimensional statistical features, the processor uses an adaptive parameter injection algorithm to estimate the default values of key parameters.
[0113] When the missing key parameter is the target total number of questions At that time, the processor uses the following formula to calculate the default value of the target total number of questions: ; in, This is the default value for the target total number of questions. The length of the valid token for the document. ρ Information entropy density, information entropy density ρ It is determined by a combination of one or more statistical features in the information density feature dimension. α The coefficient for questions per unit length. α This represents the expected number of question and answer records generated per unit of token length, and in one embodiment, it is set to 0.05 questions per token. β For density enhancement hyperparameters, density enhancement hyperparameters β A value greater than 1 is used to generate more target questions in texts with high information density. This is for floor function. This is a truncation function used to restrict the calculation result to a specific range. Within the range, This ensures that the total number of questions is at least 1 and does not exceed the system's allowed limit.
[0114] When the missing key parameter is the maximum number of retries In this case, the processor employs a similar adaptive strategy, estimating the maximum number of retries based on statistical characteristics such as the syntactic complexity of the document set. For document sets with higher syntactic complexity, the processor correspondingly relaxes the default value for the maximum number of retries, providing more opportunities for correction for more complex generation tasks.
[0115] The processor records the entire chain of "statistical feature values, intermediate estimated values derived from statistical features, and final injected default values" used in the adaptive parameter injection process in the verification report to support the auditability of the adaptive parameter injection process.
[0116] Through the adaptive parameter injection mechanism, even if the user starts the synthesis task with zero configuration, the processor can automatically generate a reasonable synthesis specification that matches the document collection based on its size and content characteristics, thus avoiding the mismatch problem caused by static constant injection.
[0117] Step 260: Combine all the second configuration parameters processed by the processing strategy with the first configuration parameters that do not conflict to obtain the composite specification.
[0118] After steps 240 and 250 are completed, the processor will summarize all the second configuration parameters after conflict resolution, all the first configuration parameters that have not been detected to have conflicts, and the default parameters supplemented by adaptive parameter injection to obtain a complete and logically consistent synthesis specification.
[0119] Step 270: Perform normalization and serialization processing on the synthesis specification, and perform hash calculation on the normalized and serialized synthesis specification to generate the synthesis specification fingerprint.
[0120] The processor performs normalization serialization processing on the synthesis specification obtained in step 260. The purpose of normalization serialization processing is to eliminate non-semantic differences between different representations of the synthesis specification, ensuring that content-equivalent synthesis specifications always produce the same serialization output, thereby guaranteeing that the generated fingerprint remains stable.
[0121] Normalized serialization processing includes at least two of the following operations: The nested structure of the synthesis specification is flattened into a one-dimensional key-value space. For example, the parameter "level under diff in section1 equals 0.8" in the original configuration, which is in a hierarchical nested form, is transformed into a one-dimensional key-value pair form "section1.diff.level equals 0.8" after flattening.
[0122] Sorting by key name means arranging all key-value pairs in the composition specification according to the lexicographical order of the key names.
[0123] Unified precision means unifying all floating-point values to the same precision representation.
[0124] Remove whitespace characters, that is, remove all unnecessary spaces, newlines, and tabs from the serialized text of the composition specification.
[0125] Perform type normalization, which means unifying values of different data types into a standard format.
[0126] The execution order rules are fixed, meaning that the values of the list type are arranged according to preset rules to ensure that the order of the elements inside the list is determined.
[0127] The processor denotes the above normalization and serialization process as a normalization function. .
[0128] In one embodiment, the normalization function The input only contains the static configuration fields in the synthesis specification and does not include runtime statistics, so as to ensure that the same static configuration always produces stable serialized output in different runtime instances.
[0129] The processor performs normalization on the function The processed serialized byte stream undergoes hash calculation to generate a synthetic canonical fingerprint. The formula for calculating the synthetic canonical fingerprint is: ; in, For deterministic hash functions, This is a synthesized, normalized byte stream after normalization and serialization processing. To synthesize standardized fingerprints.
[0130] The composition specification fingerprint serves as a unique identifier for the composition specification and is used throughout all subsequent stages, including execution plan generation, Q&A log production, and task-level lineage record file generation. When any static configuration field in the composition specification changes, the composition specification fingerprint changes accordingly, providing accurate version traceability for subsequent processes and avoiding data confusion caused by configuration changes.
[0131] Through steps 210 to 270, this embodiment of the invention implements a synthesis specification compilation method. The method constructs a constraint dependency graph and performs systematic conflict detection and hierarchical processing on the graph, orderly resolving various constraint conflicts in the user configuration, ensuring the logical closure and parameter validity of the output synthesis specification. By introducing an adaptive parameter injection mechanism based on document statistical features, the method enables the system to autonomously generate reasonable parameters matching the size of the document set even when user configuration information is incomplete, reducing the user's configuration burden and improving system usability. The method performs normalized serialization processing on the final synthesis specification and generates a deterministic synthesis specification fingerprint, providing a fundamental anchoring identifier for subsequent end-to-end version traceability and auditability. Overall, the method ensures that the input target is compilable, auditable, and traceable, reduces uncontrollable behavior caused by configuration ambiguity or conflicts, and provides definite constraint boundaries for the generation of subsequent execution plans.
[0132] In one embodiment, before the processor compiles the input user configuration in step 110, the processor also performs a global deterministic context construction operation. The purpose of the global deterministic context construction operation is to establish a unified randomness management foundation for this plan-driven question-answering data synthesis task, so that all operations involving randomness in this task are subject to deterministic control.
[0133] The processor first generates a global random seed. The global random seed is a fixed value, either from external input or generated autonomously by the processor, and serves as the sole entropy source for all random behavior in this task. Under identical input conditions, the value of the global random seed remains consistent.
[0134] After generating a global random seed, the processor derives corresponding local random seeds for each submodule and / or each execution plan item based on the global random seed using a deterministic hash function. A submodule refers to the various functional components invoked by the processor during the execution of this method, including coverage optimization components, question-answering synthesis components, and quality verification components. A deterministic hash function is a hash function that consistently produces the same output for the same input; deterministic hash functions include pseudo-random number generation functions based on HKDF or hash functions based on SHA-256. The processor combines the global random seed with the identification information of the target submodule and / or the target execution plan item, using the combined information as input to the deterministic hash function to calculate the corresponding local random seed.
[0135] The local random seed is independent of execution order and concurrency. This means that for the same target submodule or the same target execution plan item, as long as the global random seed and its corresponding identifier remain unchanged, the value of the local random seed derived from the deterministic hash function remains constant regardless of the order in which the processor executes the tasks or the number of concurrent threads. All subsequent operations involving randomness, including sampling in coverage optimization, sorting in question-answer synthesis, and perturbation in the retry process, use the corresponding local random seed as the initialization state of the random number generator, thus ensuring the deterministic nature of the output results of these operations.
[0136] Through the aforementioned global deterministic context construction operations, the processor establishes a centrally managed, hierarchically derived stochastic control system. This stochastic control system enables the reproducibility of results at the engineering level across the entire process, from coverage optimization and execution plan decomposition to question-and-answer generation and final dataset aggregation, given the same literature set, user configuration, and global random seed. This reproducibility supports rigorous comparative experiments with controlled variables on different versions of the synthesis specification or different model configurations, and provides a fundamental guarantee for problem localization and regression testing.
[0137] In one embodiment, the steps of the processor deriving corresponding local random seeds for each submodule and / or each execution plan item using a deterministic hash function in the above embodiments are further described.
[0138] The specific process of the processor deriving a local random seed is as follows: The processor concatenates the global random seed, the target module identifier, the target task identifier, and the retry count, and performs a deterministic hash operation on the concatenated result to obtain the local random seed.
[0139] The target module identifier is a unique identifier for the submodule currently being called by the processor. For example, the coverage optimization component corresponds to a fixed module identifier string, and the question-answering synthesis component corresponds to another fixed module identifier string. The target task identifier is a unique identifier for the execution plan item currently being processed. The retry count is the cumulative value of the number of retries performed for the current execution plan item. The retry count is zero on the first execution and increments by one with each subsequent retry.
[0140] The concatenation operation refers to appending the global random seed, target module identifier, target task identifier, and retry count in a predetermined order to form a complete input string. Deterministic hashing refers to the process of calculating a deterministic hash function on the input string.
[0141] The above derivation process can be expressed by the following formula: ; in, Represented as the first i The submodule processes the first... j The local random seed derived for each task; Represents a deterministic hash function; Represents a global random seed; Indicates the first i The target module identifier for each submodule; Indicates the first j The target task identifier for each task; Indicates the current retry count; This indicates a string concatenation operation.
[0142] By incorporating the retry count into a deterministic hash function, the processor uses a new, yet still reproducible, local random seed for each retry. This new local random seed differs from the one used in the previous attempt, allowing the large language model to explore different generation paths during retries and avoid repeatedly producing the same failure results within the same sampling space. Simultaneously, because the retry count itself is a deterministic and recordable value, the new local random seed still satisfies the reproducibility requirement; that is, during subsequent auditing or playback, the processor can accurately reconstruct the random state used in each retry.
[0143] By incorporating the global random seed, target module identifier, target task identifier, and retry count into deterministic hash operations, the processor achieves refined derivation of the local random seed. This refined derivation ensures that each submodule possesses an independent, deterministic, and traceable random state for each attempt at processing each task, thus maintaining the determinism and reproducibility of the entire process at the algorithm level even in scenarios supporting concurrent execution and multiple retries.
[0144] In one embodiment, the step of generating an execution plan set based on the synthesis specification and a preset literature set in step 120 is further explained. (Refer to...) Figure 5 , Figure 5 This is a flowchart illustrating the coverage optimization-driven execution plan generation method provided by the present invention, as shown below. Figure 5 As shown, generating an execution plan set specifically includes the following steps: Step 310: Compile the input user configuration to generate a synthesis specification.
[0145] The processor receives user configuration provided by the user or the upper-level system. The user configuration describes the macro-level goals and constraints of this plan-driven question-answering data synthesis task. The information included in the user configuration covers the total number of questions, the target distribution ratio of question types and difficulties, the segment ratio, the lower limit of the definition, the upper limit of the number of retries, and the quality verification rules, etc.
[0146] The processor performs compilation processing on the user configuration. This compilation process includes at least validating the validity of each parameter in the user configuration, detecting and handling parameter conflicts, and completing any missing parameters. After compilation, the processor generates a composite specification. The composite specification is an immutable, machine-readable, and logically self-consistent internal system specification object. It contains all validated and standardized constraint parameters, serving as direct input for subsequently generating the execution plan set.
[0147] Step 320: Segment the documents in the preset document set to obtain a candidate paragraph set.
[0148] After cleaning each document in the document collection, the processor segments the document into multiple semantically complete and appropriately sized logical paragraph units. The purpose of this segmentation operation is to transform unstructured long documents into atomic text units that can be independently evaluated and cited later.
[0149] When performing segmentation, the processor does not forcibly truncate at a fixed length. Instead, it searches for the position with the minimum semantic pause cost within a flexible range around the target length as the actual segmentation point. Semantic pause cost is a cost metric based on syntactic features at candidate positions. Low values are assigned to semantic pause costs at natural semantic boundaries, such as paragraph newlines or sentence termination characters; high values are assigned at non-natural boundaries, such as ordinary text characters. By segmenting at the position with the minimum semantic pause cost, the semantic integrity of each segmented paragraph is ensured.
[0150] After the segmentation operation, the processor obtains a candidate paragraph set. The candidate paragraph set contains all candidate paragraphs segmented from all documents in the document set. Each candidate paragraph is explicitly labeled with a source document identifier and a paragraph identifier. The source document identifier is used to uniquely identify the document to which the candidate paragraph belongs, and the paragraph identifier is used to uniquely identify the position of the candidate paragraph within its document.
[0151] Step 330: Calculate the value score for each candidate paragraph in the candidate paragraph set.
[0152] The processor calculates a value score for each candidate paragraph in the candidate paragraph set. The value score quantifies the richness of testable facts contained in the candidate paragraph and is an objective numerical indicator of whether the candidate paragraph has the value of being used in exam questions.
[0153] When calculating value scores, the processor uses objective statistical features for quantitative evaluation, rather than relying on subjective judgment. The processor extracts feature values across multiple dimensions for each candidate paragraph, including at least named entity density, syntactic complexity, and structural position gain. Named entity density is the normalized density of named entities in the text of the candidate paragraph. Named entities include personal names, place names, and technical terms, and are used to reflect the knowledge content of the candidate paragraph. Syntactic complexity is a score calculated based on the syntactic structure of the candidate paragraph. Methods for measuring syntactic structure include dependency tree depth and the number of clauses. Syntactic complexity is used to distinguish between paragraphs with substantial content and those that are too simple or fragmented. Structural position gain is a gain value determined based on the structural position of the candidate paragraph in the document. When a candidate paragraph is located in a key section such as the document's summary, conclusion, or definition, the structural position gain is assigned a gain coefficient higher than the default value.
[0154] The processor further performs a weighted summation of the feature values across multiple dimensions to obtain the value score of the candidate paragraph.
[0155] Step 340: Vectorize each candidate paragraph in the candidate paragraph set to obtain the paragraph vector corresponding to each candidate paragraph.
[0156] The processor invokes the vector embedding model to map each candidate paragraph in the candidate paragraph set into a high-dimensional dense vector, which is the paragraph vector corresponding to the candidate paragraph. The vector embedding model is a pre-trained semantic representation model that transforms text input into a fixed-dimensional numerical vector output, so that semantically similar texts are closer in the vector space, and semantically dissimilar texts are farther apart in the vector space.
[0157] Paragraph vectors serve as the core basis for calculating the semantic similarity between any two candidate paragraphs in subsequent steps. By converting text into paragraph vectors, the processor can quantitatively measure the degree of semantic overlap between candidate paragraphs through mathematical operations, thus providing a computational foundation for the redundancy penalty mechanism in the subsequent coverage set selection.
[0158] Step 350: Based on the value score and paragraph vector corresponding to each candidate paragraph, select the coverage set from the candidate paragraph set.
[0159] The processor iteratively selects from the candidate paragraph set based on the value score and paragraph vector of each candidate paragraph, thus filtering out the coverage set. The coverage set is a subset of the candidate paragraph set, containing paragraphs with high information value and distinct semantic content. The filtering process prioritizes high-value paragraphs while suppressing semantic redundancy, achieving a balance between breadth and depth of knowledge in the coverage set.
[0160] In one embodiment, the processor employs a greedy selection strategy with a redundancy penalty term to perform the screening. The greedy selection strategy embodies the idea of maximizing marginal relevance, that is, in each round of selection, it comprehensively considers the information value of the candidate paragraph itself and the degree of semantic redundancy between the candidate paragraph and the selected paragraph, and selects the paragraph with the best overall utility to add to the coverage set.
[0161] Step 360: Generate an execution plan set based on the synthesis specification and coverage set.
[0162] The processor receives the synthesis specification and coverage set as input, precisely decomposes the macro-statistical quotas defined in the synthesis specification, and assigns them to specific paragraphs in the coverage set, generating an execution plan set. The execution plan set consists of multiple execution plan items. Each execution plan item is an atomic execution unit, corresponding to an independent question-and-answer generation intent, and each execution plan item is associated with at least one target paragraph and one target question type.
[0163] In one embodiment, when generating the execution plan set, the processor first allocates the target proportion in the synthesis specification to each question type to determine the quota for each question type. Then, the quota for each question type is further allocated to each specific segment in the coverage set. Finally, each allocated non-zero quota position is instantiated as an execution plan item. The execution plan set is frozen after generation and serves as a read-only instruction sequence for subsequent execution phases.
[0164] Through steps 310 to 360, the processor, starting from the macro-level user intent, gradually refines the abstract synthesis objective into a set of execution plans consisting of multiple specific, independently executable atomic execution plan items through six stages: standardized compilation, text segmentation, value assessment, vectorization, coverage optimization, and plan decomposition. Each execution plan item in the set is explicitly bound to the source paragraph and the target question type, ensuring that the subsequent question-and-answer generation process is based on solid evidence. By performing coverage optimization before generating execution plans, the method concentrates limited generation resources on paragraphs with high information density, improving the depth of knowledge coverage and the density of effective output, while reducing resource waste on low-value or repetitive text.
[0165] In one embodiment, the step of segmenting documents in a preset document set in step 320 is further explained.
[0166] When performing document segmentation on the document collection, the processor employs a semantically constrained sliding window method. This method aims to address the problem that traditional fixed-length segmentation methods result in semantically interrupted paragraphs and incomplete content.
[0167] The processor first sets the target length of the sliding window and the semantic search radius. The target length is the expected length of each segment after partitioning, measured in tokens or characters. The semantic search radius is a flexible offset around the target length, used to define the range for searching the optimal partitioning point. Let the target length be... The semantic search radius is .
[0168] The processor starts at the beginning of the document and moves the sliding window along the text direction in preset steps. Each time a split point needs to be determined, the processor offsets the current window's starting position by the target length. Based on the location, within the semantic search radius Determined elasticity range Within each interval, the semantic pause cost at each candidate position is evaluated one by one.
[0169] The cost of semantic pauses is determined based on the syntactic features at the candidate positions. The processor predefines a set of cost functions based on syntactic features. The cost function assigns different values to different types of syntactic features. Specifically, when the character at the candidate position is a paragraph newline character or a sentence terminator (including periods and question marks), the processor assigns a very low value; when the character at the candidate position is a clause separator (including commas), the processor assigns a medium value; and when the character at the candidate position is a plain text character, the processor assigns a very high value.
[0170] The processor searches for the position with the minimum semantic pause cost within the elastic interval and uses the found position as the actual split point. The search process is expressed as the following optimization objective: ; In the above formula, Here, is the index of the actual split point, and is the index of each candidate position traversed within the elastic interval. Let i be the semantic pause cost at candidate position i. The processor chooses to set the semantic pause cost. The candidate position i that yields the minimum value is used as the final actual split point. .
[0171] The processor segments the text at the actual segmentation point, then updates the starting position of the sliding window to after the actual segmentation point, and continues to repeat the above search and segmentation operation until the end of the document.
[0172] By employing the semantically constrained sliding window segmentation method described above, the processor ensures that the segmentation points of each candidate paragraph always fall on natural semantic boundaries, rather than being forcibly truncated in the middle of words or sentences. Compared to traditional fixed-length segmentation methods, this method maintains roughly uniform paragraph length while guaranteeing the semantic integrity and independence of each candidate paragraph, thus providing high-quality atomic text units for subsequent value assessment and question generation.
[0173] In one embodiment, the step of calculating the value score of each candidate paragraph in the candidate paragraph set in step 330 is further described.
[0174] For any candidate paragraph, the processor extracts feature values across multiple dimensions. These dimensions include named entity density, syntactic complexity, and structural position gain.
[0175] The processor extracts the named entity density of candidate paragraphs. The processor performs named entity recognition on the candidate paragraphs, identifying all named entities contained within them. Named entities include names of people, places, and technical terms. The processor calculates the total number of identified named entities and divides this total by the total number of words or tokens in the candidate paragraph to obtain the normalized named entity density. Named entity density reflects the knowledge content of the candidate paragraph; a higher density indicates a richer amount of information with assessment value.
[0176] The processor extracts the syntactic complexity of candidate paragraphs. The processor performs syntactic analysis on the candidate paragraphs, including constructing a dependency syntax tree and counting the number of clauses. The processor calculates a syntactic complexity score based on the depth of the dependency syntax tree and the number of clauses. Syntactic complexity is used to distinguish between paragraphs with rich structure and abundant information content and those that are too simple or fragmented. Paragraphs with higher syntactic complexity have more dimensions available for assessment in test questions.
[0177] The processor extracts the structural position gain of candidate paragraphs. The processor identifies the structural position of a candidate paragraph within its document, including whether it belongs to a specific section with concise information characteristics, such as a summary, conclusion, or definition section. When a candidate paragraph is located within such a section, the processor assigns it a gain coefficient higher than the default value, for example, multiplying the default value by 1.2; when the candidate paragraph is not located within such a section, the processor assigns it the default gain coefficient. The structural position gain reflects a priori indication of the importance of the document structure to the paragraph.
[0178] After extracting feature values from the three dimensions mentioned above, the processor performs a weighted summation of the feature values across multiple dimensions to obtain the value score of the candidate paragraph. The weighted summation calculation method is described as follows: ; in, Indicates the i-th candidate paragraph Value rating; Indicates the i-th candidate paragraph Named entity density; Indicates the i-th candidate paragraph The syntactic complexity score; Indicates the i-th candidate paragraph Structural position gain; , and For the corresponding preset weights, satisfy The preset weights are set based on offline annotation data or engineering experience.
[0179] Through the aforementioned multi-dimensional feature extraction and weighted summation methods, the processor provides an objective and calculable value quantification index for each candidate paragraph in the candidate paragraph set. The value score integrates information from three dimensions: knowledge content, syntactic richness, and structural importance. This enables the subsequent coverage set selection process to identify knowledge regions with high information density and question-setting value without relying on subjective judgment.
[0180] In one embodiment, the step of filtering out the coverage set from the candidate paragraph set based on the value score and paragraph vector corresponding to each candidate paragraph in step 350 is further explained.
[0181] The processor executes an iterative greedy selection process based on diminishing marginal returns, progressively filtering the coverage set from the candidate paragraph set. The core idea of this iterative greedy selection process is that, in each iteration, not only is the informational value of the candidate paragraphs themselves considered, but also candidate paragraphs that are semantically highly similar to the selected paragraphs are penalized, thereby suppressing semantic redundancy while ensuring the informational value of the coverage set. This idea embodies a strategy of maximizing marginal relevance.
[0182] The iterative greedy selection process includes the following steps: The processor first initializes an empty selected set, denoted as . When t=0 This is an empty set. The processor simultaneously acquires a preset target quantity K, which is determined by the total quota or chapter quota in the synthesis specification.
[0183] The processor enters an iterative loop. In each iteration, let the selected set of the current iteration be denoted as . The processor processes the candidate paragraph set excluding the already selected set. For each target candidate segment p other than the target candidate segment p, calculate the marginal gain value of the target candidate segment p.
[0184] The marginal gain value is calculated as follows: the marginal gain value is the value score of the target candidate paragraph p. The positive weighted value is the same as the target candidate paragraph p in the selected set. The marginal gain is the sum of the negative weighted values of the maximum semantic similarity between existing paragraphs. The formula for calculating the marginal gain is: ; in, The preset balance coefficient, The value range is [0,1]; Indicates candidate paragraphs p Value rating; Indicates candidate paragraphs p With paragraphs already in the selected set q Semantic similarity between them; Indicates candidate paragraphs p With the selected set The maximum semantic similarity among all existing paragraphs in the text.
[0185] semantic similarity Based on the paragraph vectors corresponding to the target candidate paragraph p and the existing paragraph q respectively and The calculation is as follows. In one embodiment, semantic similarity is measured using vector cosine similarity, which is calculated as follows: ; In the above formula, Let p be the paragraph vector corresponding to the target candidate paragraph. Let q be the paragraph vector corresponding to the existing paragraph. The dot product of two paragraph vectors. and These are the L2 norms of the two paragraph vectors, respectively.
[0186] In another embodiment, semantic similarity is measured using Jaccard similarity. The processor transforms the target candidate paragraph p and the existing paragraph q into k-gram sets Sp and Sq, respectively. The calculation method of Jaccard similarity is described as follows: ; In the above formula, Let p be the set of k-grams for the target candidate paragraph. Given a set of k-grams for an existing segment q, The number of elements in the intersection of the two sets. The number of elements in the union of the two sets. To reduce the computational overhead for large-scale candidate paragraph sets, the processor also uses MinHash signatures to approximate the Jaccard similarity.
[0187] In the formula for marginal gain, the balance coefficient This controls the trade-off between information value and redundancy penalty. When the balance coefficient... When the value is larger, the positive weighting of the value score accounts for a higher proportion in the marginal gain value, and the screening results tend to select candidate paragraphs with the highest information value; when the balance coefficient is larger... When the value is small, the negative weighting of the maximum semantic similarity in the marginal gain value has a higher proportion, and the screening result tends to select candidate paragraphs that are semantically most different from the existing paragraphs in the selected set. In engineering practice, the balance coefficient... The value range is from 0.6 to 0.8.
[0188] After calculating the marginal gain values of all target candidate paragraphs in the current round, the processor selects the candidate paragraph that maximizes the marginal gain value. Candidate paragraphs Add to selected collection This updates the selected set to The select and add operations are described as follows: ; ; in, This represents the candidate paragraph with the largest marginal gain value in the current iteration round; P Represents the set of candidate paragraphs; This represents the set of selected items up to the current iteration round; This indicates the remaining paragraphs in the candidate paragraph set after excluding the already selected set; For covering sets.
[0189] The processor repeats the above iterative steps, that is, it repeatedly calculates the marginal gain value and adds the candidate paragraph with the largest marginal gain value to the selected set, until the number of paragraphs in the selected set reaches the preset target number K. When the number of paragraphs in the selected set reaches K, the processor outputs the selected set as the final coverage set.
[0190] Through the aforementioned iterative greedy selection process based on diminishing marginal returns, the processor achieves a balance between "mining high-value knowledge regions" and "suppressing semantic redundancy between knowledge points." The coverage set is the set of question materials with the optimal overall utility under the current generation resource budget, enabling subsequent execution plans to concentrate generation resources on content regions with high information value and distinct characteristics, avoiding quota waste caused by repeatedly generating questions from similar paragraphs.
[0191] In one embodiment, the step of generating an execution plan set based on the synthesis specification and the coverage set in step 360 is further explained.
[0192] The processor receives the synthesis specification and coverage set, performs quota calculation, integer allocation, and plan item instantiation operations, and generates an execution plan set. The process specifically includes the following steps: The first step is to calculate the expected quota for each question type based on the target total number of questions and the target proportion of each question type in the synthesis specification.
[0193] Let the target total number of questions set in the synthesis specification be... The set of question types defined in the synthesis specification is as follows: The question set contains K question types, and the target proportion of each question type is defined in the synthesis specification. ,satisfy .
[0194] The processor is based on the target total number of questions. Target proportion of each question type Calculate the product of the products for each question type. Expected quota The expected quota is calculated as follows: ; Due to the target total number of questions Target ratio as an integer It is usually a non-integer, therefore the expected quota They are usually floating-point numbers and cannot be directly used as the actual allocated quota.
[0195] The second step is to take the integer part of each expected quota to obtain the basic quota for each question type.
[0196] The processor is for each question type Expected quota Perform a floor operation to obtain the basic quota for each question type. The round-down operation is expressed as: ; In the above formula, This is a floor function. After floor operation, the sum of the base quotas for all question types is usually less than the target total number of questions. The difference between the two is the quota difference that needs to be made up in subsequent steps.
[0197] The third step is to calculate the decimal remainder of each expected quota. Then, in descending order of the decimal remainder, add one to the basic quota of the question type that is ranked first, until the sum of the basic quotas of all question types equals the target total number of questions, thus obtaining the final quota for each question type.
[0198] The processor calculates each question type. Expected quota With basic quota The difference between them yields the decimal remainder for each question type. The method for calculating decimal remainders is expressed as follows: ; The processor calculates the quota difference M that needs to be made up. The calculation method for the quota difference M is expressed as follows: ; The processor sorts all question types by their respective decimal remainders. Sort the questions in descending order. Based on the sorting result, the processor allocates the basic quotas to the top M question types in sequence. Add one. After adding one, the sum of the quotas for all question types is exactly equal to the target total number of questions. The processor will use the quota for each question type after incrementing by one as the final quota for each question type.
[0199] The integerization method used in the above steps is the maximum remainder method. The maximum remainder method minimizes the deviation between the actual quota for each question type and the target ratio by allocating the remaining quota to the question type with the largest decimal remainder. Compared to simple rounding or probability sampling methods, the maximum remainder method ensures that the total number of questions strictly meets the target, and that the deviation between the actual ratio of each question type and the target ratio is minimized. This avoids long-tail drift caused by probability sampling and achieves engineering stability in distribution fitting.
[0200] The fourth step is to allocate each final quota to each target paragraph in the coverage set, and to instantiate each non-zero quota position as an execution plan item. Each execution plan item is associated with at least one target paragraph and one target question type.
[0201] After obtaining the final quota for each question type, the processor constructs an allocation matrix. The allocation matrix uses the target paragraphs in the coverage set as its first dimension and the question types as its second dimension. Specifically, the allocation matrix is represented as a two-dimensional allocation matrix consisting of the number of paragraphs multiplied by the number of question types. , where N is the number of paragraphs in the covered set, and K is the number of question types.
[0202] The processor retrieves preset allocation rules, including paragraph weight constraints, chapter quota constraints, or a combination of paragraph weight constraints and chapter quota constraints. Based on these rules, the processor fills in the final quota for each question type into the corresponding target position in the allocation matrix. By filling in the final quota, the processor determines the number of questions generated for each target paragraph. Each element in allocation matrix A... The value represents the number of questions of type j assigned to the i-th paragraph.
[0203] The processor iterates through all target positions in the allocation matrix, instantiating non-zero target positions as execution plan items. For each instantiated execution plan item, it is forcibly bound to a unique target paragraph, a unique target question type, and a corresponding generation quota index. The generation quota index distinguishes the generation quotas of multiple identical target question types allocated to the same target paragraph. Each execution plan item is an independent atomic plan unit, containing a source document identifier, paragraph identifier, target question type, target difficulty, and local constraints inherited from the synthesis specification.
[0204] The processor gathers all instantiated execution plan items and generates an execution plan set based on them. Once generated, the execution plan set is frozen and processed sequentially as a read-only instruction sequence in subsequent generation phases.
[0205] The fifth step is to generate an execution plan set based on all execution plan items.
[0206] The processor aggregates all instantiated execution plan entries to generate an execution plan set. Once generated, the execution plan set is frozen and executed sequentially as a read-only instruction sequence in subsequent execution phases.
[0207] Through the steps described above, the processor decomposes the macroscopic statistical objective defined in the synthesis specification into a series of specific, independently executable atomic execution plan items by precisely allocating integer quotas and using two-dimensional allocation. This method ensures that the aggregated results of all execution plan items have the highest degree of fit to the macroscopic objective curve set by the synthesis specification in terms of question type and difficulty distribution, making the distribution characteristics of the final dataset controllable and predictable.
[0208] In one embodiment, after the processor completes the instantiation of non-zero quota locations into execution plan items, the processor also performs reachability checks and type downgrade operations on each execution plan item in the execution plan set.
[0209] The reachability check is based on the understanding that not every paragraph in the coverage set is suitable for generating questions and answers for all question types. For example, a paragraph containing only a single factual statement does not meet the necessary conditions for generating questions requiring multi-hop reasoning. Therefore, before freezing and issuing execution plan items to the execution plane, the processor pre-checks each execution plan item to determine whether the combination of its target paragraph and target question type is feasible, thereby avoiding assigning unachievable tasks to subsequent generation stages.
[0210] For any execution plan item, the processor performs a reachability check. The reachability check includes determining whether the target paragraph corresponding to the execution plan item meets the preset generation conditions of the corresponding target question type.
[0211] The preset generation conditions vary depending on the question type. Taking reasoning questions as an example, the preset generation conditions for reasoning questions are that the target paragraph contains at least two named entities and at least one entity relationship. The processor performs entity recognition and relationship extraction on the target paragraph, determining whether the number of named entities in the target paragraph is greater than or equal to two, and whether the number of entity relationships in the target paragraph is greater than or equal to one. The reachability check determination method is described as follows: In the above formula, p represents the target segment corresponding to the execution plan item. It is a reasoning question. The set of named entities extracted from the target paragraph p. This is the set of entity relationships extracted from the target paragraph p. This is an operation to retrieve the number of elements in a set.
[0212] When the reachability check is passed, the processor retains the execution plan without modification, and the execution plan is generated according to the target question type in subsequent execution stages.
[0213] When the reachability check fails, the processor does not directly discard the quota corresponding to the execution plan item, but instead performs target question type replacement. Target question type replacement involves replacing the target question type in the execution plan item with candidate question types according to a preset degradation order. The preset degradation order is an ordered chain from high-difficulty question types to low-difficulty question types. The degradation order starting with reasoning question types is: Reasoning Question Types The questions are downgraded to information extraction questions. It was then downgraded to a true / false question type. The demotion order is described as follows: .
[0214] After replacing the target question type with the next candidate question type in the downgrade order, the processor re-executes the reachability check for the replaced candidate question type. The processor repeats the reachability check and target question type replacement steps until one of the following two termination conditions is met: The first termination condition is that the reachability check passes. The processor uses the currently replaced candidate question type as the final execution question type for the execution plan item, and the execution plan item is generated according to the final execution question type in subsequent execution phases.
[0215] The second termination condition is when all candidate question types have been replaced and none have passed the reachability check. When all candidate question types in the degradation order have been tried and none meet the preset generation conditions, the processor marks the execution plan item as unreachable. Unreachable state indicates that the target paragraph is not suitable for generating any supported question types. After marking the execution plan item as unreachable, the processor does not delete the execution plan item from the execution plan set, but retains it as an explicit placeholder in the execution plan set. The purpose of explicit placeholder is to maintain a one-to-one correspondence between the paragraph index in the coverage set and the execution plan item sequence number in the execution plan set, which facilitates subsequent audit tracing and replay reconstruction. At the same time, when calculating the indicators, the processor counts the number and distribution of execution plan items marked as unreachable as an independent metric, and does not include execution plan items marked as unreachable in the denominator calculation of the effective plan distribution, thereby ensuring that the distribution fitting evaluation and replay process are not misaligned due to the absence of plan items.
[0216] Through the aforementioned accessibility checks and question type degradation mechanisms, the processor filters out "impossible tasks" before the execution plan is frozen, reducing the ineffective consumption of infeasible samples in subsequent generation and verification stages. Simultaneously, the question type degradation mechanism maintains the overall robustness of the execution plan set by replacing unreachable, high-difficulty questions with reachable, low-difficulty questions, ensuring total capacity while maintaining overall robustness. The explicit placeholder strategy for unreachable final states maintains the comparability of evaluation criteria and the integrity of the audit chain.
[0217] In one embodiment, after the processor has completed instantiating each non-zero quota location into an execution plan item, the processor also generates a stability identifier for each execution plan item.
[0218] A stability identifier is a unique identifier determined by the "plan intent" of an execution plan item. The "plan intent" refers to a combination of information describing "which paragraph, which original target quota position, and what original question type" an execution plan item will execute. The core characteristic of a stability identifier is that it remains unchanged even after the execution plan item undergoes question type downgrading, retries, and final state changes.
[0219] The stability identifier is calculated based on the following five elements: synthetic specification fingerprint, source document identifier, paragraph identifier, original target question type, and quota index.
[0220] The synthesized specification fingerprint is a unique identifier generated after normalizing and serializing the synthesized specification and calculating its hash in step 310. The source document identifier is the unique identifier of the document to which the target paragraph corresponding to the execution plan item belongs. The paragraph identifier is the unique identifier of the target paragraph corresponding to the execution plan item. The original target question type is the question type determined during the quota allocation phase, i.e., the initial question type assigned by the allocation matrix in the above steps, not the question type actually executed after reachability checks and question type downgrading. The quota index is the quota sequence number under the same source document identifier, the same paragraph identifier, and the same original target question type combination, used to distinguish multiple quotas of the same question type allocated to the same location.
[0221] The processor concatenates the above five elements and performs a hash operation to generate a stable identifier. The calculation method for the stable identifier is described as follows: ; In the above formula, For stable identification, For cryptographic hash functions, This is a string concatenation operation. To synthesize standardized fingerprints, For source document identification, For paragraph marking, The original target question type is represented by j, which is the quota subscript.
[0222] Because the stability flag is calculated based on only five elements that are determined during the quota allocation phase and remain unchanged thereafter, the stability flag does not change with subsequent task downgrades, retry operations during the execution phase, or the eventual acceptance, rejection, or unreachable state of the execution plan. The stability flag remains unique and unchanged throughout the entire lifecycle of the execution plan.
[0223] By generating a stable identifier for each execution plan item, the processor establishes a persistent identity anchor for each atomic plan unit in the execution plan set. The stable identifier serves as the basis for aggregation and grouping in subsequent audit statistics, as an index for exact hits in cache reuse, and as a correlation key for restoring the complete lifecycle trajectory of the plan item in task replay, thereby supporting the reliability of statistical consistency, cache reuse, and lineage replay.
[0224] In one embodiment, the step of the processor generating question-and-answer records corresponding to each execution plan item in step 130 is further explained.
[0225] For an execution plan item currently being processed, the processor performs the following operations: The first step is to load the corresponding original text paragraphs based on the source document identifier and paragraph identifier specified in the execution plan.
[0226] The processor reads the source document identifier and paragraph identifier recorded in the current execution plan. The source document identifier is used to locate the target document in the document set, and the paragraph identifier is used to locate the target paragraph in the target document. Based on the source document identifier and paragraph identifier, the processor retrieves and loads the corresponding original text paragraph from the document set. The original text paragraph serves as the context input for this question-and-answer session.
[0227] The second step is to split the original text paragraphs into an ordered sequence of sentences and add a position label to each sentence.
[0228] The processor performs sentence-level segmentation on the loaded raw text paragraph, splitting it into an ordered sequence of sentences. This ordered sequence contains multiple independent sentences arranged in the original text order. The processor adds an explicit position label to each sentence in the ordered sequence. The position label is a unique integer number, assigned in ascending order according to the sentence's position in the ordered sequence.
[0229] For example, given an original text paragraph containing three sentences, after splitting and adding position labels, the ordered sentence sequence takes the following form: [1] The TCP / IP protocol suite is divided into four layers.
[0230] [2] These are the application layer, transport layer, network layer, and network interface layer, respectively.
[0231] [3] The application layer is responsible for handling specific application details.
[0232] The numbers within the square brackets are the positional tags. The purpose of these positional tags is to provide a precise sentence-level anchor point for the large language model and processor during subsequent question-answer generation and quality verification processes.
[0233] The third step is to select the corresponding instruction template based on the target question type and constraints specified in the execution plan.
[0234] The processor reads the target question type and constraints recorded in the current execution plan. The target question type specifies the type of question to be generated, such as a multi-hop reasoning question, information extraction question, or judgment question. Constraints include target difficulty, maximum output length, and whether cross-sentence quotations are allowed. Based on the target question type, the processor selects the corresponding instruction template from a pre-defined instruction template library. An instruction template is a predefined structured text framework containing the roles of the large language model, descriptions of the generation task, and strict constraints on the output format. The instruction template mandates the inclusion of thought chain guidance constraints and structured output constraints. Thought chain guidance constraints require the large language model to generate a parsing or reasoning process before generating the answer. Structured output constraints require the large language model to output the result in a predefined JSON format, which must contain at least the question stem field, answer field, parsing field, and evidence field.
[0235] The fourth step is to assemble the ordered sentence sequence with location tags and the instruction template into a prompt word.
[0236] The processor uses the instruction template as both the system instruction and task instruction parts of the prompt, and embeds the ordered sentence sequence with position labels as the user input part of the prompt, i.e., as reference text embedded in the prompt. The processor also injects constraint information from the current execution plan into the corresponding positions in the instruction template. After these assembly operations, the processor obtains a complete prompt. At the input layer, the prompt fixes all requirements regarding the target question type, target difficulty, evidence constraints, and output format.
[0237] The fifth step involves the processor inputting the prompt words into the large language model to generate question-and-answer records. These records contain at least the question stem, the answer, the reasoning process, and evidence fields including the reference location label.
[0238] The processor sends the assembled prompt words to the application programming interface (API) of the large language model, driving the large language model to perform a generation operation. When calling the API of the large language model, the processor injects the local random seed corresponding to the current execution plan item as the sampling seed and sets conservative sampling parameters to ensure that the output of the large language model has high consistency at the engineering level.
[0239] The large language model generates a raw output based on the instruction constraints in the prompt words and the reference text. The processor performs JSON parsing and format validation on the raw output. When the parsing and validation operations are successful, the processor maps the raw output to a structured question-and-answer record. The question-and-answer record contains at least the stem, answer, parsing reasoning process, and evidence fields including reference position tags. The stem is the question text generated by the large language model. The answer is the response to the stem generated by the large language model. The parsing reasoning process records the reasoning process of the large language model before generating the answer. The evidence field is a list of integers, where each integer corresponds to a position tag of a sentence in an ordered sentence sequence, indicating the specific position of the original sentence supporting the answer. When the parsing and validation operations fail, the processor returns an explicit parsing error status, which is handled uniformly by subsequent quality verification steps; the processor does not initiate a retry in this step.
[0240] Through the five steps described above, the processor transforms an abstract execution plan into a structured question-and-answer record containing fine-grained evidence anchoring information. During processing, by adding positional tags to each sentence in the original text paragraph and requiring the large language model to explicitly cite these tags during the generation phase, each answer can be precisely traced back to a specific sentence in the original document, solving the problem of untraceable answer sources in traditional generation schemes. Simultaneously, through the strong constraint of "think before you answer and must cite evidence," free text generation is reshaped into a semi-structured "evidence plus reasoning" task, reducing the probability of the large language model generating illusions and providing usable structured signals for subsequent quality verification steps.
[0241] In one embodiment, the step of the processor performing quality verification on the question-and-answer record corresponding to any execution plan item in step 140 is further described. The quality verification performed by the processor on the question-and-answer record includes a structure format verification step, a semantic consistency verification step, and a logical specification verification step, which are performed sequentially.
[0242] Structure format verification steps: The processor first performs structural format validation on the question-and-answer records at the syntax level. Structural format validation includes the following three checks: The first check is structural integrity verification. The processor performs integrity and field type checks on the structured JSON object of the question-and-answer record. The processor verifies whether the option field in the question-and-answer record is a non-empty list, whether the evidence field is a list of integers, and whether the answer field is a valid element in the list contained in the option field. If any of these checks fails, the processor marks the question-and-answer record with a format type error label.
[0243] The second check is a content constraint check. The processor checks the value range and logical constraints of each field in the question-and-answer record, including verifying whether the question stem field is not empty and whether the number of options is within a preset reasonable range. If any of the above checks fails, the processor marks the question-and-answer record with a format constraint violation tag.
[0244] The third check is the evidence range check. The processor performs out-of-bounds detection on the evidence fields in the question-and-answer record. An evidence field is a list of integers, where each integer represents the position label of a sentence in an ordered sentence sequence. The processor verifies whether each integer in the evidence field falls within the numbering range of the ordered sentence sequence formed after sentence-level segmentation of the original text paragraph. When any integer in the evidence field exceeds the numbering range—for example, if the original text paragraph contains only five sentences but the evidence field references sentence number eight—the processor marks the question-and-answer record with an evidence out-of-bounds error label.
[0245] Semantic consistency verification steps: Provided that the structural format verification step is passed, the processor further verifies the question and answer records at the semantic level to determine whether the question and answer records are well-founded.
[0246] The processor performs a fidelity assessment of the answer. It concatenates the set of original sentences pointed to by the evidence fields in the question-and-answer record into premise text. It then combines the answers from the question-and-answer record with the content of the parsing fields into hypothesis text. The processor introduces a natural language inference model (NLP) by inputting the text pair consisting of the premise text and hypothesis text. The NLP performs inference on the text pair and outputs one of three results: implication, irrelevance, or contradiction. When the NLP's result is contradictory or irrelevance, the processor considers the fidelity of the question-and-answer record to have failed the verification and adds a disloyalty or illusion label to the reason tag of the question-and-answer record. The processor also records the confidence score output by the NLP, which is used in subsequent uncertainty calculations.
[0247] Logical specification verification steps: After the semantic consistency verification step is completed, the processor further verifies the question and answer records at the logical level to determine whether the question and answer records conform to the specification.
[0248] The processor performs a difficulty matching check. It scores the question stems in the question-and-answer records using a difficulty scoring model or a pre-defined combination of rules to obtain a predicted difficulty coefficient. The difficulty score considers not only surface features (including the number of options and sentence length) but also structured features (including the number of entities involved, the length of the inference chain, and the similarity between distractors). The processor maps the predicted difficulty coefficient to a range of zero to one and compares it with the target difficulty range specified by the synthesis specification for the current execution plan item. When the predicted difficulty coefficient exceeds the tolerance range allowed by the target difficulty range, the processor labels the question-and-answer record with a difficulty mismatch tag.
[0249] After completing the three levels of verification steps described above, the processor summarizes the scores and reason labels for each verification step, generating a structured judgment result. The judgment result includes at least the format score, fidelity score, difficulty matching score, reason label list, and a comprehensive judgment on whether to recommend passing.
[0250] Through the sequential execution of the structure format verification step, semantic consistency verification step, and logical specification verification step, the processor achieves multi-level automated quality verification of question-and-answer records, from surface format to deep semantics and then to logical specification. This multi-level verification transforms question-and-answer quality from subjective judgment into a calculable and statistically significant structured evaluation report composed of quantifiable indicators and cause labels. This forms a measurable, comparable, and regressible quality control closed loop, reducing the risk of rework and delivery due to inconsistencies between perceived and actual results.
[0251] In one embodiment, the semantic consistency verification steps in the above embodiments are further described. In addition to the aforementioned answer fidelity determination operation, the semantic consistency verification steps also include a question backtracking verification operation.
[0252] The purpose of the processor performing the backtracking verification operation is not only to verify whether the answer is derived from the evidence, but also to check whether the question itself is anchored to the evidence. Although the question stem in the question-and-answer record is an interrogative sentence, the core entities and keywords appearing in the question stem should mainly come from the original text fragments pointed to by the evidence field.
[0253] The specific process of the processor performing the problem backtracking verification operation is as follows: The processor performs entity and keyword extraction operations on the question stems of the question-and-answer records to obtain the first entity set. The first entity set contains all the core entities and keywords identified from the question stem text.
[0254] The processor performs entity and keyword extraction operations on the original sentence pointed to by the evidence field, resulting in a second entity set. The second entity set contains all the core entities and keywords identified from the original sentence text corresponding to the evidence field.
[0255] Entity matching in entity and keyword extraction is achieved through methods such as lemmatization, alias dictionary, or embedding similarity.
[0256] The processor calculates the intersection of the first entity set and the second entity set, and then calculates the ratio of the number of elements in the intersection to the number of elements in the first entity set to obtain the question anchoring score. The formula for calculating the question anchoring score is: ; in, This indicates the question-anchored score; Represents the first set of entities; Represents the second set of entities; This represents the intersection of the first set of entities and the second set of entities; This indicates the number of elements to be retrieved from the set.
[0257] When the question anchoring score is below a preset threshold, the processor determines that most of the key concepts in the question stem do not appear or correspond to in the original text pointed to by the evidence field. The processor marks the question-and-answer record as a question illusion. The preset threshold is set to 0.8 by default and can be adjusted according to the application scenario. The processor records the question illusion mark in the reason tag of the question-and-answer record for subsequent gating decisions and statistical analysis.
[0258] Through the aforementioned question backtracking verification operation, the processor verifies whether the question stem itself in the question-and-answer record has factual basis. The question backtracking verification operation detects illusions from the "question end," complementing the aforementioned answer fidelity judgment operation which detects illusions from the "answer end." Together, they constitute a two-way verification of the semantic consistency of the question-and-answer record, further reducing the risk of the final dataset containing unfounded content.
[0259] In one embodiment, the processing steps in step 140 when the processor determines that the quality of the question-and-answer record does not meet the acceptance criteria and the number of times the question-and-answer record is regenerated has not reached the preset retry limit are further explained.
[0260] In this scenario, the processor also performs a model routing decision operation before triggering the local closed-loop correction process. The model routing decision operation is used to determine whether, in the next retry, the generation task should be routed to a model level with a higher capability than the currently used large language model.
[0261] The process by which the processor executes the model routing decision operation is as follows: First, the processor calculates the uncertainty score of the question-and-answer record. The uncertainty score is a value normalized to the interval between zero and one, used to quantify the "hesitation" or "reliability" of the current large language model in generating this question-and-answer record. A higher uncertainty score indicates a lower reliability of the result generated by the current large language model for this question-and-answer record.
[0262] The processor then compares the uncertainty score with a preset uncertainty threshold.
[0263] When the uncertainty score exceeds a preset uncertainty threshold, the processor determines that the current large language model has an insufficient probability of success for this question-and-answer record. The processor will then route the model for the next regeneration of the question-and-answer record to a higher-level model than the current model. A higher-level model refers to a large language model version with a larger parameter scale or stronger inference capabilities. The processor will call the higher-level model to perform the question-and-answer generation operation in the next retry.
[0264] When the uncertainty score does not exceed the preset uncertainty threshold, the processor determines that the currently used large language model still has a high success rate for this question-and-answer record. The processor keeps the current model level unchanged and only makes targeted adjustments for the next retry based on the aforementioned correction suggestions.
[0265] In this embodiment, model routing decision is not an independent, unconditionally triggered action, but rather an optional upgrade strategy within a local closed-loop correction process. The processor only performs a model upgrade when all three conditions are met simultaneously: the question-and-answer record is determined to have a fixable defect, the number of retries has not been exhausted, and the uncertainty score exceeds a preset uncertainty threshold.
[0266] Through the aforementioned model routing decision operation based on uncertainty scores, the processor achieves precise limitation on the scope of high-cost model calls while ensuring pass rates. This model routing decision operation avoids the resource waste caused by indiscriminately calling high-cost models for all retry tasks, upgrading the model tier only in specific situations where the current low-cost model lacks confidence, thus achieving a dynamic balance between data quality and generation costs.
[0267] In one embodiment, the steps of the processor calculating the uncertainty score of the question-and-answer record in the above embodiment are further explained.
[0268] The uncertainty score is not determined by a single signal source, but is obtained by fusing multiple heterogeneous uncertainty signal sources to improve the robustness and accuracy of the uncertainty score.
[0269] The processor acquires component values from at least two uncertainty signal sources. The uncertainty signal sources include at least two of the following: model self-assessment confidence, generation probability entropy, self-consistency divergence, and external verification confidence.
[0270] Model self-assessment confidence refers to the confidence level that the large language model is required to output in the question-answer synthesis stage, which assesses whether the answer is entirely derived from evidence. Model self-assessment confidence reflects the large language model's subjective confidence level in its own output.
[0271] The generation probability entropy value refers to the negative logarithm of the information entropy or average probability of the generation probability of words corresponding to key fields such as answer fields and parsing fields in the question-and-answer record, when the application programming interface of the large language model supports returning the generation probability of each word during the generation process. The higher the information entropy, the greater the "hesitation" of the large language model when generating key fields.
[0272] Self-consistency divergence refers to the degree of divergence between several candidate results generated by a processor in a low-cost manner for the same execution plan item. The degree of divergence can be measured, for example, by the answer consistency ratio. The greater the divergence among the candidate results, the higher the uncertainty of the large language model representing that execution plan item.
[0273] External verification confidence refers to the processor directly using the output probabilities of each sub-model in the quality verification step as an uncertainty signal. The output probabilities of each sub-model include the confidence probability of the natural language inference model in judging the "implication" relationship, and the prediction variance output by the difficulty scoring model, etc.
[0274] After acquiring component values from at least two uncertain signal sources, the processor normalizes each component value and then performs weighted fusion to obtain an uncertainty score. The formula for weighted fusion is as follows: ; in, Indicates uncertainty fraction; Indicates the first m The component values of a type of uncertain signal source; Indicates the first m Preset weights corresponding to various uncertain signal sources; clip (·,[0,1]) represents a truncation function that restricts the calculation result to the interval between zero and one.
[0275] Through the multi-source signal fusion method described above, the processor simultaneously acquires uncertainty information from both the internal model perspective and the external verification perspective, and fuses this uncertainty information into a unified quantification score. Compared to relying on a single signal source, this fusion method has higher robustness and can more accurately identify the reliability of the current large language model on specific question-and-answer records, thereby providing a more precise basis for subsequent model routing decisions.
[0276] In one embodiment, the step of the processor processing the execution plan item in step 140 is further described, specifically involving the subsequent global governance process when the quality of the question and answer record does not meet the acceptance conditions and the number of times the question and answer record is regenerated reaches a preset retry limit.
[0277] When the processor determines that all retries for the current execution plan have been exhausted and the question-and-answer record still does not meet the acceptance criteria, the processor marks the current execution plan as rejected and reports a failure signal to the core coordinator. The failure signal includes the reason for the failure of the current execution plan and relevant indicators, such as the question type, chapter, and pass rate of previous retries.
[0278] Upon receiving a failure signal, the core coordinator executes a governance strategy based on the current global state. The core coordinator is the scheduling component responsible for maintaining the global lifecycle of the task in this method. It does not directly generate question-and-answer content; instead, it drives the collaborative work of various functional components by maintaining the global state and issuing control signals.
[0279] The global state includes at least the cumulative failure rate, the number of remaining unexecuted execution plans, and the cost incurred. The cumulative failure rate is the proportion of all processed execution plans marked as rejected as of the current moment. The number of remaining unexecuted execution plans is the number of execution plans in the execution plan set that have not yet been processed. The cost incurred is the number of tokens or the total cost incurred in this task as of the current moment.
[0280] The global state is represented in the form of a state vector as follows: ; in, Indicates the current time t The global state vector; This indicates the current cumulative failure rate; This indicates the number of remaining unexecuted execution plan items; This indicates the cost already incurred.
[0281] The core coordinator executes governance strategies and makes global-level governance decisions based on the global state vector and the received failure signals. These governance decisions differ from local retries for individual execution plan items within the execution plane; instead, they involve strategic adjustments to subsequent execution strategies from a global perspective.
[0282] Through the aforementioned global governance process, the processor constructs a global closed-loop governance mechanism built upon local closed loops. This global closed-loop governance mechanism enables the core coordinator to make strategic-level interventions and adjustments based on the global state when the local synthesis and verification closed loops fail to resolve the failure of a certain execution plan item. This avoids the problem of the system repeatedly consuming resources on the same type of failure mode without taking any macro-level corrective measures, ensuring that the overall task maintains a controllable balance between quality, cost, and capacity.
[0283] In one embodiment, the steps of the core coordinator executing the governance strategy based on the current global state in the above embodiments are further explained.
[0284] The core coordinator executes the governance strategy by comparing each indicator in the global state vector with its corresponding preset threshold and making one of the following three governance decisions based on the comparison results: The first governance decision: When the cumulative failure rate exceeds the preset failure rate threshold and the number of remaining unexecuted execution plan items is greater than zero, the core coordinator triggers a replanning operation.
[0285] The preset failure rate threshold is the maximum allowed cumulative failure rate. When the cumulative failure rate exceeds the preset failure rate threshold, it indicates that there is a systemic, non-incidental problem in the current execution plan set. At the same time, a number of remaining unexecuted execution plan items greater than zero indicates that the task is not yet finished and there is still room and necessity for adjustment.
[0286] During the replanning operation, the core coordinator discards execution plan items marked as rejected and instructs the coverage optimization-driven execution plan generation component to retrieve replacement resources from the coverage set and generate new execution plan items to supplement the execution plan set. While performing the replanning operation, the core coordinator, while maintaining the main constraints of the synthesis specification, allows for moderate relaxation of local thresholds or adjustment of problem type selection to improve the accessibility of subsequent execution plan items. Newly generated execution plan items are assigned new plan fingerprints to distinguish them from discarded execution plan items.
[0287] The second governance decision: When the cumulative failure rate does not exceed the preset failure rate threshold, the core coordinator confirms the rejection status without triggering global intervention.
[0288] When the cumulative failure rate does not exceed the preset failure rate threshold, it indicates that the current failure is an occasional event within the normal noise range and does not constitute a systemic deviation requiring global intervention. The core coordinator confirms the final status of the execution plan item currently marked as rejected, includes the execution plan item and its failure reason in the statistics and lineage records, and the process continues to the next pending execution plan item.
[0289] The third governance decision: When the cost incurred exceeds a preset cost threshold, the core coordinator terminates the current task.
[0290] The preset cost threshold is the maximum number of tokens or the maximum cost allowed for this task. When the cost consumed exceeds the preset cost threshold, it indicates that continuing execution will result in unacceptable resource waste. The core coordinator immediately terminates the entire execution flow of the current task to prevent further resource consumption. After terminating the task, the processor aggregates and delivers all question-and-answer records of the completed acceptance states up to the termination time.
[0291] Governance decisions can be expressed as a formula: ; in, π represents governance decision-making; π represents the governance strategy function. This represents the global state vector at the current time t; This indicates a received failure signal; This indicates the preset failure rate threshold; This indicates a preset cost threshold; REPLAN corresponds to the first governance decision, i.e., replanning; IGNORE corresponds to the second governance decision, i.e., confirming the rejection status without triggering global intervention; ABORT corresponds to the third governance decision, i.e., terminating the current task.
[0292] Through the aforementioned three hierarchical judgment mechanisms for governance decisions, the core coordinator transforms individual failure events at the micro level into macro-level, evidence-based strategic adjustments. This hierarchical judgment mechanism enables the system to take differentiated responses to anomalies of varying severity: proactively correcting systemic deviations through replanning, tolerating occasional noise to maintain overall throughput, and providing a safety net through circuit breakers to protect against resource exhaustion. This mechanism avoids the efficiency losses or resource waste that result from a one-size-fits-all approach to all failure events.
[0293] In one embodiment, the step 140 in which the processor stores the question-and-answer records into the final record set is further described. The storage step is not simply writing the question-and-answer records into a list, but includes operations such as normalization, fingerprint generation, and atomic writing to ensure data integrity and traceability.
[0294] The specific process of the processor performing the store step is as follows: First, the processor performs normalized serialization processing on the question-and-answer records to generate record fingerprints.
[0295] The processor performs normalized serialization on the question-and-answer records. The purpose of normalized serialization is to eliminate non-semantic differences in field order, character encoding, and whitespace formatting, ensuring that two question-and-answer records with identical semantic content produce identical byte streams after normalization. The processor then performs a hash operation on the normalized byte stream to generate a record fingerprint. The record fingerprint is a unique identifier calculated based on the semantic content of the question-and-answer records; question-and-answer records with identical semantic content share the same record fingerprint.
[0296] Secondly, the processor determines the storage path based on the recorded fingerprint and performs an existence check.
[0297] The processor uses the recorded fingerprint as the primary key to plan the storage path of the question-and-answer record in the content-addressable storage system. In one embodiment, the storage path consists of a prefix directory consisting of the first few characters of the recorded fingerprint and a filename named after the recorded fingerprint. The processor checks whether a corresponding file already exists in the storage path. If a corresponding file already exists in the storage path, it indicates that a record with semantically identical content to the current question-and-answer record already exists in the content-addressable storage system. At this point, the processor does not need to perform a write operation again; it only needs to update the cache index and task lineage information to complete the storage.
[0298] Finally, when the corresponding file does not exist in the storage path, the processor writes the normalized question and answer records to a temporary file, performs integrity verification on the temporary file, and then converts the temporary file into a formal file through an atomic renaming operation.
[0299] The processor performs a three-stage atomic write operation: "write-verify-rename".
[0300] During the writing phase, the processor writes the normalized and serialized question-and-answer records to a temporary file located in the same directory as the storage path. The temporary file's filename is different from the official file's filename; for example, a temporary tag suffix is appended to the end of the official filename.
[0301] During the verification phase, the processor reads back the contents of the temporary file and re-executes normalization serialization and hash operations on the read-back content. The processor compares the recalculated hash value with the recorded fingerprint to verify whether the contents of the temporary file are complete and consistent with expectations.
[0302] During the renaming phase, when verification passes, the processor uses the atomic renaming operation provided by the operating system to rename the temporary file to a formal file named according to the recorded fingerprint. The characteristic of the atomic renaming operation is that it either succeeds completely or no changes occur; there are no intermediate states.
[0303] Through the aforementioned normalized serialization, content-based addressing existence checks, and three-phase atomic write operations, the processor achieves secure and idempotent writing of question-and-answer records. The write operation ensures that no corrupted or partially written record files exist in the storage system, and that repeated writing to the same record fingerprint will not generate duplicate data or corrupt existing files. The record fingerprint, as a content-based unique identifier, provides a solid foundation for subsequent data deduplication, cache reuse, and lineage auditing.
[0304] In one embodiment, the steps of the processor performing normalized serialization processing on the question-and-answer records and generating record fingerprints in the above embodiment are further explained.
[0305] The first step is for the processor to filter fields related to the semantic results from the question-and-answer records, excluding volatile fields.
[0306] The processor scrutinizes all fields in the question-and-answer records, retaining only those fields strongly correlated with the semantic results of the records for subsequent hash calculations. Fields strongly correlated with semantic results include plan identifier, question stem, options, answer, parsing, and evidence fields. The processor excludes volatile fields from the question-and-answer records. Volatile fields are those whose values change under different operating environments or execution times but do not affect the semantic content of the records. Volatile fields include generation time and execution time. The purpose of excluding volatile fields is to ensure that for two question-and-answer records with completely identical semantic content, even if their generation times or execution times differ, the information participating in the hash calculation after this step remains identical, thus guaranteeing that the final generated record fingerprints of the two question-and-answer records are consistent.
[0307] The second step is for the processor to sort the filtered fields by key name in lexicographical order and perform encoding normalization on the strings.
[0308] The processor sorts the JSON object according to the lexicographical order of the keys for the fields retained after the first filtering step. This sorting operation eliminates the problem of different byte streams for the same content due to different field order. The processor also performs encoding normalization on the string values in the fields. Encoding normalization includes converting the strings to the Unicode standard form C, and standardizing newline and whitespace characters. Encoding normalization ensures that the strings are standardized in both encoding and format.
[0309] The third step is for the processor to perform a hash operation on the processed byte stream to obtain the recorded fingerprint.
[0310] The processor serializes the sorted and encoded normalized JSON object into a byte stream. The processor performs a SHA-256 hash operation on the byte stream to obtain a hash digest. The processor extracts the first few bytes of the hash digest as a fingerprint. The formula for calculating the fingerprint is: ; in, This indicates that fingerprints have been recorded; This represents a question-and-answer record; This represents the combination of field filtering, key name sorting, and encoding normalization operations in the first and second steps above. This indicates a SHA-256 hash operation; [0:16] indicates extracting the first sixteen bytes of the hash digest.
[0311] Through the three steps described above, the processor achieves format-independent, semantically content-based uniqueness calculation for question-and-answer records. The fingerprint generation process ensures that question-and-answer records with the same semantic content generate completely consistent fingerprints regardless of the operating environment, the original order of fields, or changes in volatile information such as generation timestamps. As the primary key in the content-addressable storage system, the fingerprint provides a deterministic foundation for subsequent idempotent writes, data deduplication, and lineage tracing.
[0312] In one embodiment, prior to step 130, the processor also performs a cache query operation on the current execution plan item. The purpose of the cache query operation is to check whether the current execution plan item to be processed already has a successfully generated result in the historical task that can be reused. When a reusable result exists, the processor directly reuses the historical result, skipping the entire process of question-answer synthesis and quality verification, thereby avoiding redundant calculations.
[0313] The specific process of the processor performing a cache lookup operation is as follows: The first step is for the processor to perform a hash operation on the current execution plan item and the static configuration after normalizing and serializing them to generate a plan fingerprint.
[0314] The processor retrieves the contents of the current execution plan and the static configuration frozen at the start of this task. The static configuration only includes static control variables related to the generated results, such as the synthesis specification fingerprint, global random seed, local random seed, model-side temperature parameters and kernel sampling probability thresholds and maximum output token count, the version identifier of the selected prompt template, the hard thresholds and maximum retry limit of the quality control gating, and the uncertainty threshold for strong model routing triggering. The static configuration does not include runtime dynamic statistics, such as cumulative failure rate, remaining plan count, consumed cost, real-time throughput latency, timestamp, concurrency, and machine identifier. Excluding dynamic statistics ensures that the same execution plan item is still identified as "the same thing" in different runtime environments, avoiding cache misses due to runtime differences.
[0315] The processor performs normalized serialization on the current execution plan and static configuration, then performs a hash operation to generate a plan fingerprint. The formula for calculating the plan fingerprint is: ; in, Indicates the planned fingerprint; Indicates the current execution plan item; This indicates a static configuration; This indicates a normalized serialization operation; This indicates a SHA-256 hash operation; [0:16] indicates extracting the first sixteen bytes of the hash digest.
[0316] The second step is for the processor to query the planned fingerprint input to the L1 cache; the L1 cache is a Bloom filter.
[0317] The processor will schedule fingerprint input into the L1 cache. The L1 cache is a Bloom filter. The Bloom filter consists of a bit array and multiple independent hash functions. The lookup function for the L1 cache is: ; in, This represents the output of the first-level cache lookup function; h represents the plan fingerprint; B represents the bit array; This represents the i-th hash function; k represents the total number of hash functions. This represents a logical AND operation. The output space of the first-level cache lookup function is either yes or no.
[0318] The third step is that when the first-level cache returns a non-existent result, the processor determines that the cache has been missed and performs the step of generating a question-and-answer record for the current execution plan item.
[0319] When the output of the L1 cache lookup function is negative, it means that the plan fingerprint absolutely does not exist in the Bloom filter. In this case, the processor directly determines that a cache miss has occurred, without accessing the L2 cache, and immediately performs the normal question-and-answer record generation and quality verification process for the current execution plan item.
[0320] Fourth, when the L1 cache returns a possible result, the processor will plan to input the fingerprint into the L2 cache for an exact lookup; the L2 cache is a key-value mapping table.
[0321] When the L1 cache lookup function outputs "yes," it indicates that the plan fingerprint may exist in the Bloom filter, requiring further access to the L2 cache for precise confirmation. The L2 cache is a persistent key-value map. The processor performs an exact lookup in the L2 cache using the plan fingerprint as the key. The L2 cache lookup function is: ; in, This represents the output of the second-level cache lookup function; Represents the set of keys in the key-value map; E represents the cached entry returned upon a hit. Cache entry E is a tuple containing metadata, including at least the record fingerprint, final state, update timestamp, and model configuration summary.
[0322] Fifth, when the second-level cache is hit and the status of the corresponding entry is accepted, the processor reads the existing question-and-answer record associated with the entry for reuse, skipping the steps of generating question-and-answer records and performing execution quality verification for the current execution plan item.
[0323] When the second-level cache lookup function returns a non-empty cache entry, and the final state of the cache entry is "accepted," the processor determines a cache hit. The processor extracts the record fingerprint from the cache entry and reads the corresponding physical file from the content-addressed storage system based on the record fingerprint. It then deserializes the physical file into a question-and-answer record and returns it. At this point, the processor skips the entire process of question-and-answer synthesis and quality verification for the current execution plan item, only recording the "cache hit" event in the log.
[0324] When the question-and-answer synthesis and quality verification process is completed and the gating decision is in the accept state, the processor writes the plan fingerprint and the corresponding record fingerprint into the secondary cache, and inserts the plan fingerprint into the Bloom filter of the primary cache to complete the cache update operation so that it can be hit when subsequent tasks query.
[0325] Through the two-level caching query operation based on Bloom filters and key-value mapping tables, the processor achieves rapid retrieval and reuse of historical results before question-and-answer generation. The first-level cache eliminates the vast majority of "absolutely non-existent" query requests with extremely low computational overhead, avoiding invalid access to the second-level cache; the second-level cache performs precise matching after the initial screening by the first-level cache, ensuring the accuracy of the reused results. This two-level caching mechanism allows the processor to directly reuse existing, quality-verified question-and-answer records when the same synthesis specification and execution plan reappear in subsequent tasks, skipping the costly large language model calls and quality verification processes. This significantly reduces the lexical consumption and time cost caused by repeated computation, making the system increasingly economical over time.
[0326] In one embodiment, after the processor generates the final dataset in step 150, the processor also performs the operation of generating a task-level lineage record file.
[0327] The task-level lineage record file is a structured metadata file that describes the complete association information of this plan-driven question-and-answer data synthesis task from the initial input to the final output. The task-level lineage record file includes at least the following fields: Task Identifier. The task identifier is a unique number for this synthesized task, used to distinguish and index multiple tasks.
[0328] Synthetic specification fingerprint. The synthetic specification fingerprint is a unique identifier generated after performing normalization and serialization processing on the synthetic specification and calculating its hash value in step 110. The synthetic specification fingerprint records the version information of the synthetic specification on which this task is based.
[0329] Global random seed. The global random seed is the unique entropy source value used to control all random behaviors in this task. Recording the global random seed allows for accurate reconstruction of the random states used in all random operations during subsequent auditing.
[0330] The plan fingerprint set contains the plan fingerprints of all execution plans involved in the execution of this task. When a replanning operation occurs during the execution of this task, the plan fingerprint set contains the plan fingerprints of the initial plan and the plan fingerprints of the new plan generated after the replanning, to fully record the evolution history of the plan.
[0331] A fingerprint set is recorded. This set contains the fingerprints of all question-and-answer records in the final dataset delivered for this task. By using this fingerprint set, auditors can trace the physical storage location of each question-and-answer record in the content-addressed storage system and verify the integrity of the question-and-answer records.
[0332] Metric Summary. The metric summary provides quantitative statistical information for the entire task process. It includes at least the following values: total number of planned items, total number of accepted items, total number of rejected items, total number of discarded items, pass rate, discard rate, percentage of strong model usage, average number of retries, and average cost per item. The metric summary also includes a task-level quality score summary, which includes at least the aggregated average of format score, fidelity score, issue anchoring score, and difficulty matching score.
[0333] In one embodiment, the task-level lineage record file also includes a delivery information field and a log reference field. The delivery information field records the storage location and checksum of the final dataset. The log reference field records the storage location of detailed adjudication logs and plan logs. The adjudication log records detailed data such as the verification decision, reason label, number of retries, and model level for each execution plan item. The plan log records information such as unreachable final states, degradation chains, and triggering reasons. The adjudication log and plan log are referenced by the task-level lineage record file and are not directly embedded in the body of the task-level lineage record file.
[0334] The processor organizes the task-level lineage record file together with the final dataset and metric summary into the final deliverables of this task.
[0335] By generating a task-level lineage record file, the processor centrally records the synthesis specification version information, random state information, plan evolution history, content fingerprints of output data, and key statistical indicators of this task in a structured file. Using the synthesis specification fingerprint, global random seed, plan fingerprint set, and record fingerprint set recorded in the task-level lineage record file, along with code version information, auditors can reconstruct or audit the complete behavior of this task at any subsequent time, thus achieving task-level traceability, quality auditability, and experiment replayability.
[0336] Reference Figure 6 , Figure 6 This is a schematic diagram of the structure of the plan-driven question-answering data synthesis system provided by the present invention. The system includes: The synthesis specification compilation module is used to compile the input user configuration and generate the synthesis specification; The execution plan generation module is used to generate an execution plan set based on the synthesis specifications and a preset literature set. The execution plan set contains multiple execution plan items with independent generation intentions. The question-and-answer synthesis module is used to generate question-and-answer records corresponding to each execution plan item; The quality verification and gating module is used to perform quality verification on the question and answer records corresponding to any execution plan item; when the quality of the question and answer records is determined to meet the preset acceptance conditions, the question and answer records are stored in the final record set. The delivery aggregation module is used to aggregate all question and answer records in the final record set to generate the final dataset. The storage and caching module is used to perform standardized serialization processing on question-and-answer records and generate record fingerprints; The storage and caching module is also used to determine the storage path based on the recorded fingerprint and perform an existence check; The storage and caching module is also used to write the standardized question and answer records to a temporary file when the corresponding file does not exist in the storage path. After verifying the integrity of the temporary file, the temporary file is converted into a formal file through an atomic renaming operation.
[0337] The core coordinator is used to trigger a replanning operation when the cumulative failure rate exceeds a preset failure rate threshold and the number of remaining unexecuted execution plan items is greater than zero. This operation discards execution plan items marked as rejected and generates new execution plan items to supplement the execution plan set. The core coordinator is also used to confirm the rejection status without triggering global intervention when the cumulative failure rate does not exceed the failure rate threshold. The core coordinator is also used to terminate the current task when the cost consumed exceeds a preset cost threshold.
[0338] It should be noted that the plan-driven question-and-answer data synthesis system provided by the present invention can execute the plan-driven question-and-answer data synthesis method of any of the above embodiments during specific operation, which will not be elaborated in this embodiment.
[0339] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 7 As shown, the electronic device may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communications interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute the plan-driven question-and-answer data synthesis method provided in the above embodiments.
[0340] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0341] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer is able to execute the plan-driven question-and-answer data synthesis method provided in the above embodiments.
[0342] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the plan-driven question-and-answer data synthesis method provided in the above embodiments.
[0343] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0344] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0345] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A plan-driven question-and-answer data synthesis method, characterized in that, include: The input user configuration is compiled and processed to generate a synthesis specification; Based on the synthesis specifications and the preset literature set, an execution plan set is generated, which contains multiple execution plan items with independent generation intentions; Generate the question and answer records corresponding to each of the aforementioned execution plan items; Perform quality verification on the question and answer records corresponding to any of the execution plan items; wherein, when it is determined that the quality of the question and answer records meets the preset acceptance conditions, the question and answer records are stored in the final record set; All question-and-answer records in the final record set are aggregated to generate the final dataset.
2. The plan-driven question-answering data synthesis method according to claim 1, characterized in that, Before the compilation processing of the input user configuration, the following is also included: Generate a global random seed; Based on the global random seed, a corresponding local random seed is derived for each submodule and / or each execution plan item using a deterministic hash function; wherein, the local random seed is independent of the execution order and concurrency.
3. The plan-driven question-answering data synthesis method according to claim 2, characterized in that, The step of deriving a corresponding local random seed for each submodule and / or each execution plan item using a deterministic hash function includes: The global random seed, target module identifier, target task identifier, and retry count are concatenated, and a deterministic hash operation is performed on the concatenated result to obtain the local random seed.
4. The plan-driven question-and-answer data synthesis method according to claim 1, characterized in that, The generation of question-and-answer records corresponding to each of the execution plan items includes: Load the corresponding original text paragraphs based on the source document identifier and paragraph identifier specified in the execution plan item; The original text paragraph is split into an ordered sequence of sentences, and a position label is added to each sentence; Select the corresponding instruction template based on the target question type and constraints specified in the execution plan item; The ordered sentence sequence with the location tag is assembled with the instruction template into a prompt word; The prompt words are input into a large language model to generate the question-and-answer record; wherein the question-and-answer record includes at least the question stem, the answer, the analytical reasoning process, and the evidence field that references the location tag.
5. The plan-driven question-answering data synthesis method according to claim 1, characterized in that, The quality verification of the question-and-answer records corresponding to any of the execution plan items includes the following verification steps: Structure format verification steps: Verify the completeness of the fields in the question and answer record, the field types, and whether the position tags referenced in the evidence fields are within the sentence number range of the original text paragraph; Semantic consistency verification steps: Take the set of original sentences pointed to by the evidence fields in the question and answer records as the premise text, and take the answers and explanations in the question and answer records as the hypothetical text. Perform semantic implication judgment on the premise text and the hypothetical text to obtain the fidelity judgment result. Logical specification verification steps: Assess the difficulty of the questions in the question-and-answer records, and compare the difficulty scores with the target difficulty range specified in the execution plan to obtain the difficulty matching result.
6. The plan-driven question-answering data synthesis method according to claim 5, characterized in that, The semantic consistency verification step also includes: Entity and keyword extraction is performed on the question stems of the question and answer records to obtain the first entity set; Entity and keyword extraction is performed on the original sentence pointed to by the evidence field to obtain a second entity set; The problem anchoring score is obtained by calculating the ratio of the intersection of the first entity set and the second entity set to the first entity set. When the question anchor score is lower than a preset threshold, the question and answer record is marked as a question illusion.
7. The plan-driven question-and-answer data synthesis method according to claim 1, characterized in that, When it is determined that the quality of the question-and-answer record does not meet the acceptance condition and the number of times the question-and-answer record is regenerated has not reached the preset retry limit, the method further includes: Calculate the uncertainty score of the question-and-answer records; When the uncertainty score exceeds the preset uncertainty threshold, a model with a higher capability than the current model will be called when the question and answer records are regenerated next time. When the uncertainty score does not exceed the uncertainty threshold, the current model level remains unchanged.
8. The plan-driven question-answering data synthesis method according to claim 7, characterized in that, The calculation of the uncertainty score of the question-and-answer record includes: Obtain component values of at least two uncertainty signal sources; wherein, the uncertainty signal sources include at least two of the following: model self-assessment confidence, generation probability entropy value, self-consistency divergence degree, and external verification confidence. The uncertainty score is obtained by normalizing each component value and then performing a weighted fusion.
9. The plan-driven question-answering data synthesis method according to claim 1, characterized in that, Also includes: When it is determined that the quality of the question and answer record does not meet the acceptance conditions and the number of times the question and answer record is regenerated has not reached the preset retry limit, the question and answer record corresponding to the execution plan item is regenerated based on the correction suggestions generated by the verification, and the quality verification is performed on the regenerated question and answer record until the regenerated question and answer record meets the acceptance conditions, or the number of times the question and answer record is regenerated reaches the retry limit. When the quality of the question-and-answer record does not meet the acceptance condition and the number of times the question-and-answer record is regenerated reaches the retry limit, the execution plan item is marked as rejected and a failure signal is reported to the core coordinator. The core coordinator executes governance strategies based on the current global state; wherein the global state includes at least the cumulative failure rate, the number of remaining unexecuted execution plan items, and the costs incurred.
10. The plan-driven question-answering data synthesis method according to claim 9, characterized in that, The implementation governance strategy includes: When the cumulative failure rate exceeds a preset failure rate threshold and the number of remaining unexecuted execution plan items is greater than zero, a replanning operation is triggered, the execution plan items marked as rejected are discarded, and new execution plan items are generated to supplement the execution plan set. When the cumulative failure rate does not exceed the failure rate threshold, the rejection status is confirmed without triggering global intervention. When the consumed cost exceeds a preset cost threshold, the current task is terminated.
11. The plan-driven question-answering data synthesis method according to claim 1, characterized in that, The step of storing the question-and-answer records into the final record set includes: The question-and-answer records are then processed into a standardized serialization sequence to generate record fingerprints; The storage path is determined based on the recorded fingerprint, and an existence check is performed. When the corresponding file does not exist in the storage path, the normalized question and answer record is written to a temporary file. After the integrity of the temporary file is verified, the temporary file is converted into a formal file through an atomic renaming operation.
12. The plan-driven question-answering data synthesis method according to claim 11, characterized in that, The step of standardizing and serializing the question-and-answer records to generate record fingerprints includes: Filter the fields related to the semantic results from the question and answer records, excluding volatile fields; Sort the filtered fields by key name in lexicographical order, and perform encoding normalization on the strings; A hash operation is performed on the processed byte stream to obtain the record fingerprint.
13. The plan-driven question-answering data synthesis method according to claim 1, characterized in that, Before generating the question-and-answer records corresponding to each execution plan item, a cache query is also performed on the current execution plan item: After normalizing and serializing the current execution plan item with the static configuration, a hash operation is performed to generate a plan fingerprint; The plan fingerprint is input into the first-level cache for querying; the first-level cache is a Bloom filter. When the first-level cache returns a non-existent result, it is determined that a cache miss has occurred, and the step of generating a question-and-answer record is executed for the current execution plan item; When the first-level cache returns a possible result, the plan fingerprint is input into the second-level cache for precise lookup; the second-level cache is a key-value mapping table. When the second-level cache is hit and the status of the corresponding entry is accepted, the existing question and answer record associated with the entry is read and reused, skipping the steps of generating question and answer records and verifying execution quality for the current execution plan item.
14. The plan-driven question-and-answer data synthesis method according to claim 1, characterized in that, Following the generation of the final dataset, the following is also included: Generate a task-level lineage record file; the task-level lineage record file includes at least a task identifier, a synthesis specification fingerprint, a global random seed, a plan fingerprint set, a record fingerprint set, and an indicator summary.
15. A plan-driven question-and-answer data synthesis system, characterized in that, include: The synthesis specification compilation module is used to compile and process the input user configuration to generate the synthesis specification; An execution plan generation module is used to generate an execution plan set based on the synthesis specification and a preset literature set, wherein the execution plan set contains multiple execution plan items; The question-and-answer synthesis module is used to generate question-and-answer records corresponding to each of the execution plan items. The quality verification and gating module is used to perform quality verification on the question and answer records corresponding to any of the execution plan items; wherein, when it is determined that the quality of the question and answer record meets the preset acceptance conditions, the question and answer record is stored in the final record set; The delivery aggregation module is used to aggregate all question and answer records in the final record set to generate the final dataset.
16. The plan-driven question-and-answer data synthesis system according to claim 15, characterized in that, Also includes: The storage and caching module is used to perform standardized serialization processing on the question-and-answer records and generate record fingerprints; The storage and caching module is also used to determine the storage path based on the recorded fingerprint and perform an existence check; The storage and caching module is further configured to write the standardized question and answer records into a temporary file when the corresponding file does not exist under the storage path, perform integrity verification on the temporary file, and then convert the temporary file into a formal file through an atomic renaming operation.
17. The plan-driven question-and-answer data synthesis system according to claim 15, characterized in that, Also includes: The core coordinator is used to trigger a replanning operation when the cumulative failure rate exceeds a preset failure rate threshold and the number of remaining unexecuted execution plan items is greater than zero. This operation discards execution plan items marked as rejected and generates new execution plan items to supplement the execution plan set. The core coordinator is also used to confirm the rejection state without triggering global intervention when the cumulative failure rate does not exceed the failure rate threshold. The core coordinator is also used to terminate the current task when the cost consumed exceeds a preset cost threshold.
18. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the plan-driven question-and-answer data synthesis method as described in any one of claims 1 to 14.
19. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the plan-driven question-and-answer data synthesis method as described in any one of claims 1 to 14.
20. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the plan-driven question-and-answer data synthesis method as described in any one of claims 1 to 14.