A semi-automatic corpus acquisition method and system based on human-machine collaboration

By employing a semi-automated corpus collection method that combines human-machine collaboration with large models and human experts, high-quality question-answer sample pairs are generated. This solves the problems of high cost and low efficiency in existing technologies, realizes the generation of large-scale, high-quality data and alignment with complex human preferences, and improves the safety and reliability of the model.

CN122133715APending Publication Date: 2026-06-02STATE GRID HUNAN ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST +2

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID HUNAN ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST
Filing Date
2026-01-23
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies suffer from high costs, low efficiency, and difficulty in scaling up large-scale model training data through manual construction. Furthermore, the data generated by traditional automated methods is of uncontrollable quality and difficult to align with complex human preferences.

Method used

We employ a semi-automated corpus collection method based on human-machine collaboration, combining the batch generation capabilities of large models with the accurate evaluation capabilities of human experts. We optimize and generate high-quality question-answer sample pairs through reinforcement learning algorithms, including data cleaning, knowledge base screening, expert scoring, and multiple rounds of iterative optimization.

Benefits of technology

It effectively solves the problems of high cost and low efficiency of traditional methods, generates large-scale, high-quality question-answer sample pairs, ensures the alignment of data with human values, and improves the safety and reliability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122133715A_ABST
    Figure CN122133715A_ABST
Patent Text Reader

Abstract

This invention discloses a semi-automated corpus acquisition method and system based on human-machine collaboration. The method constructs a three-stage iterative framework of "data generation – feedback acquisition – reinforcement learning optimization": First, power industry experts define the target subdomain and formulate multi-dimensional quality standards. A large-scale model guided by domain knowledge generates diverse candidate questions and answers in batches and performs automated pre-screening. Subsequently, experts conduct multi-dimensional evaluation and optimization to form high-quality positive and negative example samples. Finally, a dedicated reward model is trained based on paired ranking loss and power industry domain constraint regularization terms, and algorithms such as group relative strategy optimization are used to optimize the generated model through reinforcement learning, forming a continuously optimized data production closed loop. This method achieves a deep integration of large-scale model batch generation capabilities and accurate human expert evaluation, enabling the efficient construction of large-scale, high-quality training corpora that conform to power safety regulations and professional standards at a low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to artificial intelligence large model technology, specifically to a semi-automatic corpus acquisition method and system based on human-machine collaboration. Background Technology

[0002] Large-scale, high-quality training data is the core foundation for building high-performance, highly reliable large language models. Currently, the construction of large model training data mainly relies on two methods: manual annotation and fully automated generation. While manual annotation methods offer controllable quality, they suffer from inherent drawbacks such as high cost, low efficiency, and difficulty in scaling. Traditional automated generation methods, while rapidly expanding data volume, suffer from uncontrollable generation quality and difficulty in aligning with complex human preferences and value orientations. With the rapid development of artificial intelligence technology, the demand for massive, high-quality training data aligned with human values ​​is increasingly urgent, and existing data construction methods can no longer meet the training requirements of next-generation intelligent agent models. Therefore, how to efficiently and sustainably produce large-scale, ultra-high-quality question-and-answer sample pairs with minimal human cost has become a key technical bottleneck restricting further improvements in the performance of large models, directly impacting the reliability, security, and usability of artificial intelligence systems. Summary of the Invention

[0003] The technical problem to be solved by this invention is that in the prior art, the cost of manually constructing large model training data is high, the efficiency is low, and it is difficult to scale up. In addition, the data generated by traditional automated methods is of uncontrollable quality and is difficult to align with complex human preferences.

[0004] To address the aforementioned problems in existing technologies, a semi-automatic corpus acquisition method and system based on human-machine collaboration is provided. This method deeply integrates the batch generation capability of large models with the ability to align human preferences through reinforcement learning fine-tuning mechanisms based on human feedback. The aim is to efficiently and sustainably produce large-scale, high-quality question-answer sample pairs that conform to human values ​​with minimal human cost.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A semi-automated corpus acquisition method based on human-computer collaboration includes the following steps: The base model is used to generate an initial set of candidate questions based on specified data generation criteria. ; For the initial candidate problem set Data cleaning was performed to obtain a set of candidate questions. ; For candidate problem set Each question uses a different large model to generate corresponding candidate answers, resulting in an initial candidate answer set for each question. The candidate answer set is obtained by filtering the data in the initial candidate answer set based on the knowledge base. ; For candidate answer set The candidate answers are scored by experts to form positive and negative samples, and a special reward model for the power industry is trained based on the positive and negative samples. Based on the score given to candidate answers by the reward model, a reinforcement learning algorithm is used to optimize and generate a large base model for candidate questions. After final processing of the high-quality question-answer pairs generated from multiple iterations, they are combined with the corresponding questions to form the final optimized question-answer pairs and stored in the corpus.

[0006] Furthermore, the specified data generation criteria include specified technical fields, core quality dimensions, and output format specifications. The core quality dimensions include technical accuracy, processing procedure standardization, completeness of security measures, and terminology standardization. Based on the specified data generation criteria, the base model generates an initial set of candidate questions. Specifically, a large-scale foundational model for question generation is used, employing high-temperature sampling and topic control techniques. Based on specified data generation domains, core quality dimensions, and output format specifications, a first batch of candidate questions is generated to obtain an initial candidate question set. .

[0007] Furthermore, regarding the initial candidate problem set During data cleaning, specific techniques employed include semantic deduplication, perplexity filtering, and topic relevance classifiers that embed specialized word vectors from the aforementioned technical field, from the initial candidate question set. The candidate question set is obtained by selecting the second number of candidate questions. .

[0008] Furthermore, when filtering data in the initial candidate answer set based on the knowledge base, the following steps are included: The initial candidate answer set is scored using a knowledge base, and candidate answers with strong factual basis are selected to obtain the candidate answer set. ; Calculate the answer set The relevance scores between the data and the knowledge base are used to select the candidate answers that best fit the procedures, thus obtaining a candidate answer set. .

[0009] Furthermore, when using the knowledge base to score the data in the initial candidate answer set, the mathematical expression is as follows:

[0010] in, This represents a given document knowledge base. As candidate answers, Indicates in the given document knowledge base and the generated query prefix Under the condition of generating the next word The probability of.

[0011] Furthermore, the first candidate answer set is calculated. When calculating the relevance score between the data and the knowledge base, the mathematical expression is as follows:

[0012] in, This represents a given document knowledge base. As candidate answers, Indicates the number of terms in the candidate answer. Indicates the first candidate answer 1 term, Indicates inverse document frequency. Indicate word frequency, and This represents the freely adjustable hyperparameter. Indicates the document length.

[0013] Furthermore, when selecting candidate answers with strong factual basis, and when selecting candidate answers that best conform to the procedures, candidate answers are selected in descending order of scores.

[0014] Furthermore, regarding the candidate response set When evaluating candidate responses, the specified data generation criteria are used as multidimensional evaluation criteria, and the candidate response set is evaluated based on these criteria. The candidate answers are scored using multi-dimensional absolute scoring to obtain scores for each candidate answer across different dimensions; alternatively, the candidate answer set is evaluated based on multi-dimensional evaluation criteria. The candidate answers are evaluated by relative ranking to obtain a ranking for each candidate answer, forming positive examples (answers with higher expert scores), negative examples (answers with lower expert scores), and comparison pairs (preference labels for positive and negative examples). The following loss is used to train a reward model specifically for the power industry:

[0015] Where L_ranking is the pairwise ranking loss, L_regularization is the power domain constraint regularization term, and λ is the weight of the power domain constraint regularization term. The pairwise ranking loss is used to enable the reward model to learn the relative preferences of power experts, while the power domain constraint regularization term enables the reward model to comply with power professional rules, ensuring compliance with power safety regulations.

[0016] Furthermore, in each round of optimization of the base model for generating candidate questions and answers using reinforcement learning algorithms, specifically, a group relative policy optimization algorithm is used to fine-tune the generated model before regenerating the data, forming an optimization closed loop of "generation-feedback-fine-tuning".

[0017] in It is the probability ratio of the new strategy to the old strategy. It is the normalized advantage value. For the clipping function, for KL Divergence penalty term: The group relative policy optimization algorithm replaces absolute value estimation with relative rewards within the group. Combined with a value-free network design and a stable optimization mechanism, it significantly improves computational efficiency and training stability compared to the traditional PPO.

[0018] Furthermore, when processing the high-quality question-answer pairs generated from multiple iterations, the first step is to deduplicate them, convert them into a standard training format, and classify and store them according to the power sub-domains (transmission and transformation, distribution, consumption, new energy, etc.). The final optimized question-answer pairs are then used to construct high-quality training sample pairs and stored in the corpus.

[0019] The present invention also proposes a semi-automatic corpus acquisition system based on human-computer collaboration, including a processor and a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the semi-automatic corpus acquisition method based on human-computer collaboration.

[0020] Compared with the prior art, the advantages of the present invention are as follows: (1) By integrating the batch generation capability of large models with the accurate evaluation capability of human experts through the human-machine collaboration mechanism, it can efficiently construct large-scale, high-quality question-answer sample pairs with extremely low manual cost. This effectively solves the problems of low efficiency, high cost and uncontrollable quality of traditional manual annotation, and provides a solid data foundation for training high-performance and high-reliability large language models.

[0021] (2) By constructing a semi-automated corpus collection workflow, human experts are freed from the heavy work of data creation, allowing them to focus on high-level decision-making tasks such as quality assessment, criterion formulation, and key sample optimization. This not only significantly improves the efficiency and consistency of sample data generation, but more importantly, by integrating the professional preferences and values ​​of human experts (such as "safety first, procedures above all") and combining them with the constraints of power safety regulations (i.e., L_regularization), it also ensures that the generated content is strictly aligned with complex human preferences and values, thereby enhancing the security and reliability of the model output. Attached Figure Description

[0022] Figure 1 This is a flowchart of an embodiment of the present invention. Detailed Implementation

[0023] The present invention will be further described below with reference to the accompanying drawings and specific preferred embodiments, but this does not limit the scope of protection of the present invention.

[0024] To address the core issues of high cost, low efficiency, and difficulty in scaling large-scale training data through manual construction in existing technologies, as well as the uncontrollable quality and difficulty in aligning with complex human preferences of data generated by traditional automated methods, this embodiment proposes a semi-automated corpus acquisition method based on human-machine collaboration. This method can serve as a new technology for generating and optimizing large-scale training data, effectively supplementing and significantly expanding existing manual annotation and automated generation methods. It can provide a scalable, efficient, and highly controllable general solution for the data preparation stage in the field of artificial intelligence, and has significant engineering application value and promising prospects for widespread application.

[0025] like Figure 1 As shown, the method includes the following steps: S101) Data generation stage: First, the data objectives and data generation criteria are manually determined, including: S11: Human experts define the domain boundaries, core quality dimensions, and output format specifications for data generation based on the target application scenario. For example, a power grid expert team defines the target sub-domain and formulates multi-dimensional quality standards, specifying the data generation domain as "single-phase grounding fault handling in distribution networks." Core quality dimensions include: technical accuracy (compliance with the "Power Safety Work Regulations"), standardized handling procedures (compliance with dispatch operation procedures), completeness of safety measures (requirements for voltage testing, grounding wire installation, etc.), and standardized terminology (use of standard power terminology). The output format is standardized as a JSON structure, containing fields such as fault phenomenon, handling steps, hazard analysis, and technical basis.

[0026] Secondly, candidate problems for agent generation include: S12: Use the base model to generate an initial set of candidate questions based on the specified data generation criteria. Specifically, it uses a large-scale foundational model for problem generation, employs high-temperature sampling and topic control techniques, and automatically generates a large number of diverse candidate problems in batches based on specified data generation domains, core quality dimensions, and output format specifications to obtain an initial candidate problem set. ; For example, based on the criteria specified by the power grid expert team in the preceding steps, the large model GPT-4 is invoked, employing high-temperature sampling (T=0.95) and thematic control techniques such as "distribution network faults," "grounding handling," and "relay protection" to automatically generate an initial set of 5,000 candidate questions. ,like: "A metallic grounding fault has occurred on a 10kV line. The insulation monitoring device shows that the voltage of phase A is 0, while the voltage of phases B and C has risen to 10kV. How should this be handled?" "The distribution transformer is making abnormal noises and the oil temperature is rising. The initial judgment is that it is an internal fault. What emergency measures should be taken?"

[0027] S13: For the initial candidate problem set Data cleaning was performed to obtain a set of candidate questions. Specifically, it employs semantic deduplication, perplexity filtering, and topic relevance classifiers that embed specialized word vectors from the aforementioned technical field to analyze the initial candidate question set. After cleaning and deduplication, a second set of candidate questions is selected to obtain a refined set of high-quality candidate questions. ; For example, semantic deduplication based on power industry-related word vectors (cosine similarity < 0.8), perplexity filtering (eliminating logically confused questions), and BERT-based power industry-related topic classification (accuracy > 92%) were used to select 1000 high-quality candidate questions from 3000 questions, forming a candidate question set. .

[0028] Finally, the agent generates candidate answers, including: S14: On the candidate problem set For each question, a different large model is used to generate corresponding candidate answers, resulting in an initial candidate answer set for each question; specifically, for the question set... For each question in the process, call Several different large models, each generated using a kernel sampling strategy. indivual( This results in a diverse range of candidate answers. Therefore, a candidate answer is generated for each question. An initial set of candidate answers; For example, for each problem (such as the ground fault problem mentioned above), three different domain knowledge-guided large models (GPT-4, ChatGLM3, PowerBERT) are invoked. Each model uses kernel sampling (top-p=0.9) to generate four candidate answers, resulting in a total of 12 candidate processing schemes as the initial candidate answer set for the corresponding problem, forming a total of 12,000 initial candidate answers.

[0029] S15: Use a knowledge base to score the data in the initial candidate answer set and filter out candidate answers with strong factual basis to obtain the candidate answer set. Specifically, a sequence-to-sequence reordering model is used to perform refined relevance and quality scoring on candidate answers based on a knowledge base, thereby automating the pre-screening of the initial candidate answer set. In this embodiment, a T5-based sequence-to-sequence reordering model is used for refined scoring, and the mathematical expression is as follows:

[0030] in, This represents a given document knowledge base. As candidate answers, Indicates in the given document knowledge base and the generated query prefix Under the condition of generating the next word The probability of.

[0031] For example, the technical document "Electric Power Safety Regulations" can be used as a given document knowledge base. Substituting into the above formula, we can obtain the scores of each candidate answer in the initial candidate answer set. Then, by sorting the answers from highest to lowest score, we can select the candidate answers with stronger factual basis to obtain the answer set. .

[0032] S16: Calculate the candidate response set The relevance scores between the data and the knowledge base are used to select the candidate answers that best fit the procedures, thus obtaining a candidate answer set. Specifically, the BM25 algorithm is used to calculate the relevance score between candidate answers and a high-quality reference knowledge base. The mathematical expression is as follows:

[0033] in, This represents a given document knowledge base. As candidate answers, Indicates the number of terms in the candidate answer. Indicates the first candidate answer 1 term, Indicates inverse document frequency. Indicate word frequency, and This represents the freely adjustable hyperparameter. Indicates the document length.

[0034] The candidate answer set is calculated using the above formula to obtain a relevance score. Perform matching, sort by score from highest to lowest, and select the top score. indivual( The candidate answers are used as a set of high-quality candidate answers.

[0035] For example, using the formula above in the BM25 algorithm, candidate answers are matched with professional documents such as "Power System Fault Analysis" and "Distribution Network Operation Regulations". Based on the relevance scores, the top four candidate answers that best fit the regulations are selected to form a high-quality candidate answer set. .

[0036] Through steps S15 and S16, the data in the initial candidate answer set is filtered based on the knowledge base to obtain the candidate answer set. This allows for the selection of a high-quality subset from the initial candidate responses, significantly improving the efficiency of subsequent manual evaluation.

[0037] S102) Manual optimization stage: First, human experts evaluate and rank the candidate responses, including: S17: On the candidate response set Experts score candidate answers to the same question. In this embodiment, human experts, based on the multidimensional evaluation criteria defined in S11, evaluate the same problem... The candidate responses are evaluated using multi-dimensional absolute scoring or relative ranking. Specifically, multi-dimensional evaluation criteria are designed based on specified data generation criteria, and the candidate response set is then evaluated according to these criteria. After multi-dimensional absolute scoring of the candidate answers, a weighted sum is calculated to obtain a percentage score for each candidate answer. After obtaining the percentage score for each candidate answer, a relative ranking evaluation can be performed to obtain the ranking of each candidate answer.

[0038] For example, three experienced distribution network dispatching experts, based on the multi-dimensional criteria established in S11, independently evaluated four candidate answers to the same question from multiple dimensions: An absolute scoring mechanism is used, with scores based on technical accuracy (40%), safety and compliance (30%), practicality (20%), and clarity of expression (10%), on a 100-point scale. For example: Answer A (85 points), Answer B (90 points), Answer C (60 points), Answer D (50 points).

[0039] Alternatively, a relative ranking mechanism can be used to rank the four answers based on preference (e.g., answer B). Answer A Answer C Answer D), and indicate the main advantages and disadvantages.

[0040] Then, training and optimization are performed, including: S18: A training dataset is constructed based on expert feedback, forming positive examples (approximately 600 answers with expert scores ≥ 85), negative examples (approximately 400 answers with expert scores ≤ 60), and comparison pairs (preference labels for positive and negative examples under the same question). The following loss function is used to train a reward model specific to the power industry: λ = 0.2, using the AdamW optimizer with a learning rate of 1e-5, trained for 3 epochs:

[0041] In the formula, For the paired ranking loss, this embodiment constructs a comparison pair from the corresponding positive and negative examples for each problem, and calculates the probability that the score of the positive example should be higher than that of the negative example to obtain the paired ranking loss. Let λ be the power sector constraint regularization term, and λ be the weight of the power sector constraint regularization term. A paired ranking loss is used to enable the reward model to learn the relative preferences of power experts, while the power sector constraint regularization term ensures that the reward model adheres to power industry rules, guaranteeing compliance with power safety regulations.

[0042] S19: Based on the reward model's scoring of candidate answers, a reinforcement learning algorithm is used to optimize the base model for generating candidate questions. Specifically, in each round, a group relative policy optimization algorithm is used to fine-tune the generated model, forming an optimization closed loop of "generation-feedback-fine-tuning". A clipping threshold is set. =0.2, KL divergence penalty coefficient β=0.01, the objective function mathematical expression of the group relative policy optimization algorithm is as follows:

[0043] In the formula, This indicates the number of question groups in each training batch, where each group contains multiple answer samples for the same question. This represents the ratio of the probabilities of the new and old strategies. , The normalized advantage value is calculated using the reward model's rating of the i-th answer and the average score of the G answers to the same question. This represents the length of the i-th answer sequence. For the clipping function, for KL Divergence penalty term. The group relative policy optimization algorithm replaces absolute value estimation with relative rewards within groups, and combined with a value-free network design and a stable optimization mechanism, it significantly improves computational efficiency and training stability compared to the traditional PPO.

[0044] After fine-tuning, the optimized generative model weights are saved as the new base model and redeployed to the beginning of the data generation process, replacing the previous generation pre-trained base model used in S12. Subsequently, based on the updated model, the entire process from S12 to S19 is iterated until the required number of iterations is met.

[0045] S20: Perform final processing on the high-quality candidate questions and answers generated from multiple iterations. Specifically, remove redundant cases generated from multiple iterations based on the semantic similarity of electricity (threshold 0.88) to achieve deduplication. Then, convert them into a unified JSON-L format, including metadata annotation, to achieve standard training format conversion. Store them according to the power distribution field (transmission and transformation, power distribution, power consumption, new energy, etc.). Combine the final processed candidate questions and answers with the corresponding questions to form the final optimized question-answer pairs. Construct the final optimized question-answer pairs into high-quality training sample pairs and store them in the corpus.

[0046] Furthermore, this embodiment also proposes a semi-automatic corpus acquisition system based on human-computer collaboration, including a processor and a computer-readable storage medium. The computer-readable storage medium stores a computer program, which is executed by the processor to implement the steps of the semi-automatic corpus acquisition method based on human-computer collaboration described in this embodiment.

[0047] In summary, this invention addresses the problems of high cost, low efficiency, and difficulty in scaling manually constructed training data in existing technologies, as well as the uncontrollable quality and difficulty in aligning with complex human preferences in data generated by traditional automated methods. This invention proposes an agent training method capable of generating high-quality electricity question-and-answer datasets at a lower manual cost. To maximize the efficiency of manual scoring, classic screening mechanisms are employed to remove duplicate and irrelevant samples. These measures are relatively simple and mature, thus reusable and effectively reducing the workload of manual scoring. Regarding the question set generation process, this invention adds agent reinforcement learning algorithm training after expert scoring, giving the questioning agent the ability to autonomously optimize and generate a high-quality question set. Unlike human feedback reinforcement learning, RLHF optimizes the agent that generates candidate answers. While most agents currently have relatively mature answering capabilities, this invention optimizes the agent that generates candidate questions, and then calls publicly available mature agents to generate a high-quality question-and-answer set, fundamentally generating a high-quality question set that conforms to electricity regulations.

[0048] Furthermore, current methods of direct scoring by intelligent agents generally suffer from limitations such as bias and uncertainty, and their accuracy is particularly insufficient in specialized fields such as power. Therefore, this invention still employs human scoring. For multi-dimensional scoring, this invention uses weighted proportions to calculate scores and assign percentages to candidate answers. Unlike current methods that directly screen and optimize answers, the advantage of this invention lies in its ability to iteratively optimize and continuously generate higher-quality question sets. Through screening mechanisms and human scoring, it obtains an effective question-and-answer set, significantly reducing the human cost of involving power experts and efficiently generating large-scale, high-quality training corpora that conform to power regulations.

[0049] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0050] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A semi-automated corpus acquisition method based on human-machine collaboration, characterized in that, Includes the following steps: The base model is used to generate an initial set of candidate questions based on specified data generation criteria. ; For the initial candidate problem set Data cleaning was performed to obtain a set of candidate questions. ; For candidate problem set Each question uses a different large model to generate corresponding candidate answers, resulting in an initial candidate answer set for each question. The candidate answer set is obtained by filtering the data in the initial candidate answer set based on the knowledge base. ; For candidate answer set For the same question, candidate answers are scored by experts. Based on the expert scores, the candidate answers are divided into positive and negative samples and a reward model is trained. Then, the base model is optimized through reinforcement learning. The steps of using the base model to generate data according to the specified criteria are executed again to start a new round of iterations until the number of iterations meets the requirements. After obtaining candidate answers generated in multiple iterations and performing final processing, these answers are combined with the corresponding questions to form the final optimized question-answer pairs, which are then stored in the corpus.

2. The semi-automatic corpus acquisition method based on human-machine collaboration according to claim 1, characterized in that, The specified data generation criteria include specified technical fields, core quality dimensions, and output format specifications. The core quality dimensions include technical accuracy, standardization of processing procedures, completeness of security measures, and standardization of terminology. The base model is used to generate an initial set of candidate questions based on specified data generation criteria. Specifically, a large-scale foundational model for question generation is used, employing high-temperature sampling and topic control techniques. Based on specified data generation domains, core quality dimensions, and output format specifications, a first batch of candidate questions is generated to obtain an initial candidate question set. .

3. The semi-automatic corpus acquisition method based on human-machine collaboration according to claim 2, characterized in that, For the initial candidate problem set During data cleaning, specific techniques employed include semantic deduplication, perplexity filtering, and topic relevance classifiers that embed specialized word vectors from the aforementioned technical field, from the initial candidate question set. The candidate question set is obtained by selecting the second number of candidate questions. .

4. The semi-automatic corpus acquisition method based on human-machine collaboration according to claim 1, characterized in that, When filtering data from the initial candidate answer set based on a knowledge base, the following is included: The initial candidate answer set is scored using a knowledge base, and candidate answers with strong factual basis are selected to obtain the candidate answer set. ; Calculate the candidate answer set The relevance scores between the data and the knowledge base are used to select the candidate answers that best fit the procedures, thus obtaining a candidate answer set. .

5. The semi-automated corpus acquisition method based on human-machine collaboration according to claim 4, characterized in that, When using a knowledge base to score data in the initial candidate answer set, the mathematical expression is as follows: in, This represents a given document knowledge base. As candidate answers, Indicates in the given document knowledge base and the generated query prefix Under the condition of generating the next word The probability of.

6. The semi-automatic corpus acquisition method based on human-machine collaboration according to claim 4, characterized in that, Calculate the first candidate answer set When calculating the relevance score between the data and the knowledge base, the mathematical expression is as follows: in, This represents a given document knowledge base. As candidate answers, Indicates the number of terms in the candidate answer. Indicates the first candidate answer 1 term, Indicates inverse document frequency. Indicate word frequency, and This represents the freely adjustable hyperparameter. Indicates the document length.

7. The semi-automatic corpus acquisition method based on human-machine collaboration according to claim 4, characterized in that, When selecting candidate answers that are more factual and when selecting candidate answers that best conform to the procedures, candidate answers are selected in descending order of scores.

8. The semi-automatic corpus acquisition method based on human-machine collaboration according to claim 1, characterized in that, For candidate answer set When evaluating candidate responses, experts design multidimensional evaluation criteria based on specified data generation criteria, and then evaluate the candidate response set according to these criteria. After multi-dimensional absolute scoring of the candidate answers, a weighted sum is obtained to obtain a percentage score for each candidate answer. Based on the expert scoring results, when dividing the candidate answers into positive and negative samples, specifically, candidate answers with expert scores greater than a first specified value are classified as positive samples, and candidate answers with expert scores less than a second specified value are classified as negative samples. Preference labels are then applied to the positive and negative samples under the same question to form a comparison pair.

9. The semi-automatic corpus acquisition method based on human-machine collaboration according to claim 8, characterized in that, When optimizing the base model using reinforcement learning after training the reward model, the specific steps include: The pair ranking loss is calculated based on positive samples, negative samples, and sample pairs. Then, the reward model is trained using a loss function that includes the pair ranking loss. The mathematical expression of the loss function is as follows: in, For pairing sorting loss, Here, λ represents the power sector constraint regularization term, and λ is the weight of the power sector constraint regularization term. After obtaining the reward model's scores for candidate answers, the group relative policy optimization algorithm is used to fine-tune the generator model. After fine-tuning, the optimized generator model weights are saved as the new base model. The mathematical expression of the group relative policy optimization algorithm is as follows: In the formula, This represents the ratio of the probabilities of the new and old strategies. It is a normalized advantage value, calculated by the reward model's rating of the i-th answer and the average score of G answers to the same question. This represents the length of the i-th answer sequence. For the clipping function, for KL Divergence penalty term.

10. A semi-automatic corpus acquisition system based on human-machine collaboration, characterized in that, The method includes a processor and a computer-readable storage medium storing a computer program, which is executed by the processor to implement the steps of the semi-automatic corpus acquisition method based on human-computer collaboration as described in any one of claims 1 to 9.