Large language model evaluation set automatic generation method and device, equipment and medium
By acquiring keywords and using retrieval-enhanced generation technology to determine background knowledge, assembling complete prompts and performing quality assessments, we address the accuracy issues associated with large language models in the scheduling field, achieve automated evaluation and diverse generation, and improve the model's application efficiency and applicability in professional fields.
Patent Information
- Application Number
- CN202510546457.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies make it difficult to accurately and objectively evaluate large language models in the scheduling field, and the lack of a unified benchmark makes model testing difficult.
By obtaining keywords, using retrieval-enhanced generation technology to determine background knowledge, assembling complete prompts, generating question data and performing quality assessment, generating an evaluation set, and optimizing prompts to regenerate question data when the evaluation results are unsatisfactory.
It achieves automated and accurate evaluation of large language models in the scheduling field, reduces labor costs, improves the diversity and coverage of evaluation sets, provides a benchmark for model training and fine-tuning, and promotes the application of large language models in the scheduling field.
Smart Images

Figure CN120632385A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing technology, and in particular to a method, apparatus, device, and medium for automatically generating a large language model evaluation set. Background Art
[0002] Through continued pre-training, fine-tuning, and alignment of the base model, the large language model for dispatching uses a large amount of unlabeled data, such as work regulations, technical specifications, and dispatching logs, along with a small amount of instruction data. It can understand professional grid dispatching questions and generate relevant content. However, due to the complexity of dispatching scenarios, labeling questions and answers requires specialized personnel and is time-consuming. This results in a lack of a unified benchmark for testing large dispatching models, making accurate and objective model evaluation difficult. Summary of the Invention
[0003] In view of this, the purpose of this application is to propose a method, device, equipment and medium for automatically generating a large language model evaluation set, which is used to accurately evaluate the actual application performance of the large language model in the scheduling field.
[0004] Based on the above objectives, this application provides a method for automatically generating a large language model evaluation set, including:
[0005] Obtain keywords, determine background knowledge based on the keywords and search enhancement generation technology, determine examination points based on the initial prompt and the background knowledge, and determine the question requirements corresponding to the examination points;
[0006] Based on a preset role template, assembling a complete prompt according to the background knowledge, the examination key points, and the question requirements, generating question data according to the complete prompt, and performing a question quality assessment on the question data according to the background knowledge, the examination key points, and the question requirements to obtain an assessment result;
[0007] In response to the evaluation result that the question is qualified, the question data is stored in a question bank to obtain an evaluation set, and new keywords are obtained; wherein the evaluation set is used to evaluate the application performance of the large language model in the data scheduling scenario;
[0008] In response to the evaluation result that the question is unqualified, the number of prompt optimization times is determined, and the question data is regenerated according to the number of prompt optimization times, and the step of returning to execute the question quality evaluation of the question data based on the background knowledge, the examination points and the question requirements to obtain the evaluation result is returned.
[0009] Optionally, determining background knowledge based on the keywords and the search enhancement generation technology includes:
[0010] Determine the technical field corresponding to the keyword, and import professional documents in the technical field into a vector database;
[0011] At least one text data corresponding to the keyword is retrieved from the vector database based on vector similarity, and the at least one text data is integrated to obtain the background knowledge.
[0012] Optionally, determining examination key points based on the initial prompt and the background knowledge includes:
[0013] Performing content analysis on the background knowledge according to a preset background analysis model to obtain usable examination content;
[0014] The available examination contents are screened according to the initial prompt to obtain at least one examination key point.
[0015] Optionally, the step of assembling a complete prompt based on a preset role template and according to the background knowledge, the examination key points, and the question requirements includes:
[0016] Determining a question type and a question difficulty according to the question requirements, and determining a question sample in the assessment set according to the question type and the question difficulty;
[0017] The background knowledge and the examination points are introduced into the role template, and the complete prompt containing the examination points is determined according to the question sample.
[0018] Optionally, the question data includes a question, an answer to the question, and an explanation of the answer; and generating the question data based on the complete prompt includes:
[0019] The complete prompt is input into a trained question generation model, so that the question generation model determines the question according to the question requirements and examination points in the complete prompt, and determines the answer and the explanation corresponding to the question according to the question requirements and the background knowledge in the complete prompt, and integrates and outputs the question, the answer, and the explanation.
[0020] Optionally, the question quality assessment is performed on the question data based on the background knowledge, the examination points and the question requirements to obtain an assessment result, including:
[0021] Determine the assessment items corresponding to the question data according to the background knowledge, the examination points and the question requirements;
[0022] The problem data packet is evaluated item by item according to the evaluation items to obtain multiple item scores and score reasons corresponding to the item scores, and the sum of the item scores is determined as the total score. The item scores, the score reasons and the total score are integrated to obtain the evaluation result.
[0023] Optionally, the question data includes a question, an answer to the question, and an explanation of the answer; and determining the assessment items corresponding to the question data based on the background knowledge, the examination points, and the question requirements includes:
[0024] Determining the consistency between the question and the inspection points as the first evaluation item;
[0025] determining the similarity between the answer, the explanation, and the background knowledge as a second evaluation item;
[0026] evaluating the correctness of the answer and the explanation based on the background knowledge, and determining the correctness as a third scoring item;
[0027] Determining a question type according to the question requirement, evaluating the type similarity of the question according to the question type, and determining the type similarity as a fourth scoring item;
[0028] Determine the sample questions and the difficulty of the questions according to the question requirements, evaluate the correctness of the difficulty of the questions according to the sample questions and the difficulty of the questions, and determine the correctness of the difficulty as the fifth scoring item.
[0029] Optionally, regenerating question data according to the number of optimization times of the prompt includes:
[0030] In response to the prompt optimization number being less than a preset number threshold, optimizing the complete prompt according to the question data and the evaluation result based on a preset prompt optimization model to obtain an optimized prompt, regenerating question data according to the optimized prompt, adding one to the number of prompt optimization times to obtain a new prompt optimization number, and regenerating question data according to the new prompt optimization number when a new evaluation result indicates that the question is unqualified;
[0031] In response to the prompt being optimized a number of times greater than or equal to a preset number threshold, determining a new complete prompt by modifying the precondition data, and regenerating question data according to the new complete prompt;
[0032] The prerequisite data includes the background knowledge, the examination points and the question requirements.
[0033] Based on the same inventive concept, the present disclosure also provides a device for automatically generating a large language model evaluation set, comprising:
[0034] The precondition determination module is configured to: obtain keywords, determine background knowledge based on the keywords and the search enhancement generation technology, determine examination points based on the initial prompt and the background knowledge, and determine the question requirements corresponding to the examination points;
[0035] The question generation and evaluation module is configured to: assemble a complete prompt based on a preset role template according to the background knowledge, the examination key points, and the question requirements, generate question data based on the complete prompt, and perform a quality evaluation on the question data based on the background knowledge, the examination key points, and the question requirements to obtain an evaluation result;
[0036] a qualified question processing module configured to: in response to the evaluation result that the question is qualified, store the question data in an evaluation set and obtain new keywords; wherein the evaluation set is used to evaluate the application performance of the large language model in the data scheduling scenario;
[0037] The unqualified question processing module is configured to: in response to the evaluation result that the question is unqualified, determine the number of prompt optimization times, regenerate question data according to the number of prompt optimization times, and return to execute the question quality evaluation of the question data based on the background knowledge, the examination points and the question requirements to obtain the evaluation result.
[0038] Optionally, the precondition determination module includes:
[0039] A document importing unit is configured to: determine a technical field corresponding to the keyword, and import professional documents in the technical field into a vector database;
[0040] The data retrieval unit is configured as: a document importing unit is configured as: retrieving at least one text data corresponding to the keyword based on vector similarity in the vector database, integrating the at least one text data, and obtaining the background knowledge.
[0041] Optionally, the precondition determination module further includes:
[0042] The inspection content determination unit is configured to: perform content analysis on the background knowledge according to a preset background analysis model to obtain available inspection content;
[0043] The examination key point determination unit is configured to: filter the available examination content according to the initial prompt to obtain at least one examination key point.
[0044] Optionally, the question generation and evaluation module includes:
[0045] a question requirement parsing unit, configured to: determine a question type and a question difficulty according to the question requirement, and determine a question sample in the assessment set according to the question type and the question difficulty;
[0046] The prompt assembly unit is configured to: bring the background knowledge and the examination points into the role template, and determine the complete prompt containing the examination points based on the question sample.
[0047] Optionally, the question data includes a question, an answer to the question, and an explanation of the answer; the question generation and evaluation module further includes:
[0048] The question data generation unit is configured to: input the complete prompt into a trained question generation model, so that the question generation model determines the question according to the question requirements and examination points in the complete prompt, and determines the answer and the explanation corresponding to the question according to the question requirements and the background knowledge in the complete prompt, and integrates and outputs the question, the answer, and the explanation.
[0049] Optionally, the question generation and evaluation module further includes:
[0050] An assessment item determination unit is configured to: determine an assessment item corresponding to the question data according to the background knowledge, the examination points and the question requirements;
[0051] The evaluation score determination unit is configured to: perform item-by-item evaluation on the problem data packet according to the evaluation items, obtain multiple item scores and score reasons corresponding to the item scores, and determine the sum of the item scores as the total score, integrate the item scores, the score reasons and the total score to obtain the evaluation result.
[0052] Optionally, the question data includes a question, an answer to the question, and an explanation of the answer; and the assessment item determination unit further includes:
[0053] The first item determination subunit is configured to: determine the consistency degree between the question and the inspection key points as a first evaluation item;
[0054] A second item determination subunit is configured to: determine the similarity between the answer, the explanation and the background knowledge as a second evaluation item;
[0055] A third item determination subunit is configured to: evaluate the correctness of the answer and the explanation based on the background knowledge, and determine the correctness as a third scoring item;
[0056] a fourth item determination subunit, configured to: determine a question type according to the question requirement, evaluate a type similarity of the question according to the question type, and determine the type similarity as a fourth scoring item;
[0057] The fifth item determination subunit is configured to: determine the question sample and question difficulty according to the question requirements, evaluate the difficulty correctness of the question according to the question sample and the question difficulty, and determine the difficulty correctness as the fifth scoring item.
[0058] Optionally, the non-conformity problem handling module includes:
[0059] The first processing module is configured to: in response to the prompt optimization number being less than a preset number threshold, optimize the complete prompt according to the question data and the evaluation result based on a preset prompt optimization model to obtain an optimized prompt, regenerate question data based on the optimized prompt, increase the number of prompt optimization times by one to obtain a new prompt optimization time, and regenerate question data based on the new prompt optimization time when a new evaluation result indicates that the question is unqualified;
[0060] The second processing module is configured to: in response to the prompt optimization number being greater than or equal to a preset number threshold, determine a new complete prompt by modifying the precondition data, and regenerate question data according to the new complete prompt;
[0061] The prerequisite data includes the background knowledge, the examination points and the question requirements.
[0062] Based on the same inventive concept, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the above-mentioned method when executing the computer program.
[0063] Based on the same inventive concept, the present disclosure further provides a non-transitory computer-readable storage medium, which stores computer instructions for causing a computer to execute the method described above.
[0064] From the above, it can be seen that the method, device, equipment and medium for automatically generating a large language model evaluation set provided by the present application: obtain keywords, determine background knowledge based on keywords and retrieval enhancement generation technology, determine examination points based on initial prompts and background knowledge, and determine the question requirements corresponding to the examination points; based on a preset role template, assemble complete prompts according to background knowledge, examination points and question requirements, generate question data according to the complete prompts, and perform question quality evaluation on the question data according to the background knowledge, examination points and question requirements to obtain an evaluation result; when the evaluation result is that the question is qualified, the question data is stored in the question bank to obtain an evaluation set, and new keywords are obtained; wherein, the evaluation set is used to evaluate the application performance of the large language model in the data scheduling scenario; when the evaluation result is that the question is unqualified, determine the number of prompt optimizations, and regenerate the question data according to the number of prompt optimizations, and return to execute the step of performing question quality evaluation on the question data according to the background knowledge, examination points and question requirements to obtain the evaluation result. This system uses a search-enhanced generation approach to search for relevant professional documents for keywords. It automatically analyzes the retrieved content to derive key questions and presents answers. This eliminates the need for manual input of original questions as a starting condition, reducing labor costs and simplifying generation. It also leverages a large, unlabeled corpus from the dispatching domain to enhance the coverage of domain knowledge assessments. When questions are inappropriate, new question data can be generated in batches rather than rewriting the original, further increasing question diversity. Question data generation is based on knowledge points in the form of keywords, ensuring accurate and focused questions and avoiding the randomness inherent in large language model generation. After keyword input, evaluation sets can be automatically generated in batches, providing a benchmark for evaluating the performance of large dispatch model training and fine-tuning. This improves model development and testing efficiency, ensures the model's practicality in the field, and promotes the widespread application of large language model technology in the dispatching field. This will enhance dispatchers' information retrieval and knowledge integration capabilities, alleviate the pressure on dispatchers in emergency situations, and enhance the intelligence of business processes. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the technical solutions in this application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0066] Figure 1 This is a flowchart of a method for automatically generating a large language model evaluation set according to an embodiment of the present application;
[0067] Figure 2 This is a flowchart of another method for automatically generating a large language model evaluation set according to an embodiment of the present application;
[0068] Figure 3 A flowchart for determining background knowledge for embodiments of the present application;
[0069] Figure 4 The process of determining the key points to be examined for the embodiments of this application;
[0070] Figure 5 Flowchart for assembling a complete prompt for an embodiment of the present application;
[0071] Figure 6 A flowchart for question quality assessment for an embodiment of the present application;
[0072] Figure 7 This is a flowchart of regenerating question data according to the number of prompt optimizations in an embodiment of the present application;
[0073] Figure 8 This is a schematic diagram of the structure of the device for automatically generating a large language model evaluation set according to an embodiment of the present application;
[0074] Figure 9 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0075] In order to make the objectives, technical solutions and advantages of this application more clear, this application is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.
[0076] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the usual meanings understood by people with ordinary skills in the field to which this application belongs. The "first", "second" and similar words used in the embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0077] It should be understood herein that any number of elements in the drawings is for illustration only and not for limitation, and any naming is only for distinction and does not have any limiting meaning.
[0078] Based on the description of the above background technology, the following situations also exist in the related art:
[0079] Artificial intelligence (AI) big language model technology is rapidly developing. Currently, various industries are developing large models suitable for their specific fields to meet the needs of daily office and professional work, such as knowledge question answering, copywriting, logical reasoning, and engineering calculations, and to improve the efficiency of information acquisition and integration. With the accelerated construction of ultra-high voltage (UHV) interconnected power grids and the rapid integration of new energy sources, the amount of new knowledge and technologies in the field of power grid dispatching has increased exponentially. However, dispatching work requires extremely high accuracy and real-time results, making it increasingly difficult for professionals to grasp new knowledge and quickly apply it. Therefore, it is necessary to explore the application of AI big language models in dispatching scenarios to alleviate manual workload and improve work efficiency. To comprehensively, accurately, and objectively measure the performance of big language models in dispatching scenarios, it is necessary to construct evaluation test sets based on text corpora from the dispatching field and incorporating the characteristics of actual applications. Furthermore, methods should be explored to automatically generate large-scale question and answer sets, thereby accelerating the development process from basic models to big models for dispatching and professional scenarios.
[0080] The goal of large language model technology is to model a given text sequence in order to predict the probability distribution of a sentence or text. Conducting comprehensive evaluations of large language models provides a unified and objective benchmark for the model during pre-training, fine-tuning, and alignment. This helps accelerate model development and promotes the practical application of large models in a specific field. Unlike traditional software, large language models are data-driven and optimize themselves through learning and reasoning. They have higher autonomy and intelligence, and possess larger parameters, making them suitable for more complex and general tasks. Therefore, the functional evaluation of large language models needs to be richer and more comprehensive. The large-scale data-driven nature of large language models makes them more vulnerable to malicious attacks, so the security of large language models is also a key aspect of the evaluation.
[0081] Large language models for the dispatching domain, through continued pre-training, fine-tuning, and alignment of the base model, use a large amount of unlabeled data, such as work regulations, technical specifications, and dispatch logs, and a small amount of instruction data, to understand professional grid dispatching questions and generate relevant content. However, due to the complexity of dispatching scenarios, labeling questions and answers requires professional qualifications and consumes a considerable amount of time. This results in a lack of a unified benchmark for testing large dispatching models, making it difficult to accurately and objectively evaluate the models. It is necessary to explore methods for automatically generating evaluation sets for large dispatching models, combining the original dispatching data with publicly available models to generate large-scale question-and-answer datasets.
[0082] Question generation is a crucial task in natural language processing. Before the widespread adoption of deep learning technology, traditional question generation methods relied on manually designed rules and templates, requiring significant manpower and relying on deep grammatical knowledge, resulting in poor generalization and scalability. With the development of deep learning models, neural network models such as Seq2Seq are being used to generate questions based on answers, and the generation effect is optimized through methods such as optimizing encoder and decoder structures, optimizing semantic representations, and adding semantic annotations. However, these methods typically generate questions based on given answers or extract meaningful questions from contextual paragraphs. They are unable to truly analyze the question and simultaneously summarize the answer, and their intelligence remains insufficient.
[0083] Related technologies employ multiple agents to generate answers based on manually input prompts. Through inter-agent questioning, targeted revisions, and collective scoring, the optimal answer is ultimately selected, along with the optimal keyword set. However, these require manually crafted questions as prompts and cannot leverage professional documents as a source for answers, making them unsuitable for scheduling the automatic generation of large model evaluation sets.
[0084] Related technologies can also generate instructions based on the target output format and scenario requirements. These integrated instructions are used to guide target test data. The output performance of the large language model is determined by comparing the generated results after inputting them into the large model with the original format requirements. This test primarily targets the large model's ability to output formatted documents, focusing on comparing the original and generated documents. It cannot generate new questions and answers from the original documents and is not suitable for scheduling the automatic generation of large model evaluation sets.
[0085] Related technologies address data security or privacy issues that arise during large-scale model development by automatically rewriting the original test dataset to generate a new dataset. This ensures the validity and up-to-dateness of the test data, and corrects old test cases during large-scale model development. However, datasets primarily addressing data security and privacy issues cannot be applied to professional competence testing in the field of power grid dispatching. Furthermore, this method uses the original dataset as input and cannot spontaneously generate new datasets. Furthermore, this method primarily relies on content rewriting and cannot fully utilize large-scale original content documents.
[0086] The present application provides a method, apparatus, device and medium for automatically generating a large language model evaluation set: obtaining keywords, determining background knowledge based on keywords and retrieval enhancement generation technology, determining examination points based on initial prompts and background knowledge, and determining the question requirements corresponding to the examination points; assembling complete prompts based on preset role templates according to background knowledge, examination points and question requirements, generating question data based on the complete prompts, and performing question quality assessment on the question data based on the background knowledge, examination points and question requirements to obtain an evaluation result; when the assessment result is that the question is qualified, storing the question data in a question bank to obtain an evaluation set, and obtaining new keywords; wherein the evaluation set is used to evaluate the application performance of the large language model in a data scheduling scenario; when the assessment result is that the question is unqualified, determining the number of prompt optimizations, and regenerating the question data based on the number of prompt optimizations, and returning to execute the step of performing question quality assessment on the question data based on the background knowledge, examination points and question requirements to obtain an evaluation result. This system uses a search-enhanced generation approach to search for relevant professional documents for keywords. It automatically analyzes the retrieved content to derive key questions and presents answers. This eliminates the need for manual input of original questions as a starting condition, reducing labor costs and simplifying generation. It also leverages a large, unlabeled corpus from the dispatching domain to enhance the coverage of domain knowledge assessments. When questions are inappropriate, new question data can be generated in batches rather than rewriting the original, further increasing question diversity. Question data generation is based on knowledge points in the form of keywords, ensuring accurate and focused questions and avoiding the randomness inherent in large language model generation. After keyword input, evaluation sets can be automatically generated in batches, providing a benchmark for evaluating the performance of large dispatch model training and fine-tuning. This improves model development and testing efficiency, ensures the model's practicality in the field, and promotes the widespread application of large language model technology in the dispatching field. This will enhance dispatchers' information retrieval and knowledge integration capabilities, alleviate the pressure on dispatchers in emergency situations, and enhance the intelligence of business processes.
[0087] The following describes in detail the method for automatically generating a large language model evaluation set provided by the embodiments of the present application in conjunction with the accompanying drawings.
[0088] In some embodiments, as Figure 1 As shown, a method for automatically generating a large language model evaluation set includes:
[0089] Step 101: Obtain keywords, determine background knowledge based on keywords and search enhancement generation technology, determine examination points based on initial prompts and background knowledge, and determine the question requirements corresponding to the examination points.
[0090] In specific implementations, keywords serve as input data for the evaluation set generation process and serve as the foundational data for the process. A keyword in the evaluation set generation process can be a word, phrase, or sentence representing an entity or fact in the dispatching field, stored in a database or knowledge base. Examples include "interprovincial and interregional power grids" or "power grid dispatching organizations are divided into five levels." Keywords are user-entered data, and the evaluation set generation process is based on knowledge points in the form of keywords. This ensures that the evaluation set's topics are accurate and focused, avoiding the randomness inherent in the generation of large language model evaluation sets.
[0091] After acquiring keywords, background knowledge is first determined based on the keywords and Retrieval-Augmented Generation (RAG) technology. RAG is a technology that combines information retrieval with generative models. It dynamically incorporates external knowledge bases to enhance the generation capabilities of large language models (LLMs), improving the accuracy, timeliness, and domain adaptability of the generation process. RAG converts input keywords into vector embeddings, performs semantic searches using vector databases (such as ChromaDB and Faiss) or knowledge graphs, and returns the top-K documents or paragraphs.
[0092] For different usage scenarios, simply connect to the database corresponding to the scenario and determine the corresponding background knowledge based on the keywords in the corresponding usage scenario. For the power grid dispatch scenario, the vector database can be the power grid dispatch database. After obtaining the keywords, search for text corpus related to the keywords. Using knowledge enhancement methods, professional documents in the dispatch field are imported into the power grid dispatch database. The text corresponding to the keywords in the power grid dispatch database is searched based on vector similarity. Multiple results will be returned, each result being one or more paragraphs. These paragraphs serve as the background knowledge.
[0093] After determining the background knowledge, it's necessary to further determine the key points of the test based on the initial prompt and background knowledge, and then determine the corresponding question requirements for the key points. The key point generation model, after receiving a piece of background knowledge (e.g., a Top-K paragraph), automatically analyzes the prompt to identify key points for the test and generates multiple key points. Each key point can serve as a key point for the test, and each key point is typically in the form of a question.
[0094] The question requirements are a description of the large model generation task, including question type, question difficulty, and question examples. Question types include single-choice, multiple-choice, true / false, and short-answer questions, with difficulty levels ranging from easy, medium, and hard. Question examples are representative examples extracted from the historical scheduling exam question bank based on the selected question type and difficulty, providing reference for text length, written format, and language style during the large model generation process.
[0095] Among them, background knowledge, examination points and question requirements constitute the prerequisites for generating the evaluation set.
[0096] Step 102: Based on the preset role template, assemble a complete prompt according to background knowledge, examination points and question requirements, generate question data according to the complete prompt, and evaluate the question quality of the question data according to background knowledge, examination points and question requirements to obtain an evaluation result.
[0097] In specific implementation, the role template is defined as a domain expert in the usage scenario. For the grid dispatch usage scenario, the role template is "grid dispatch domain expert". Then, background knowledge, examination points and question requirements are brought into the role template, and the complete prompt required to generate the question is obtained. For example, the complete prompt can be: "As an expert in the field of grid dispatch, I now want to examine the staff's mastery of the knowledge in the background knowledge {background}. After analysis, it is known that the background knowledge {background} includes the following examination points {points}. Please prepare a question type of {question_type} and question difficulty of {question_difficulty}. The question format, question composition, question description and other aspects can refer to the following question examples {question_examples}, and generate question data including (1) question, (2) answer, and (3) explanation. Then, the complete prompt is input into the trained question generation model, so that the question generation model can output question data including (1) question, (2) answer, and (3) explanation according to the complete prompt.
[0098] Based on retrieval-enhanced generation technology, background knowledge is retrieved from domain documents based on keywords, and then questions, answers and explanations of the evaluation set are generated. The generation process is guided by keywords rather than being completely random and uncontrollable, so that the questions in the question data are more in line with the requirements of professional domain knowledge, increasing the targeted nature of the examination, while ensuring the factual correctness of the answers and explanations in the question data.
[0099] After obtaining the question data, it is necessary to evaluate the questions, answers, explanations, etc. in the question data based on the background knowledge, examination points and question requirements to determine the low-quality data and high-quality data in the batch-generated question data. For low-quality question data, the total evaluation score is low, and the corresponding evaluation result is that the question is unqualified. For high-quality question data, the total evaluation score is high, and the corresponding evaluation result is that the question is qualified. For example, data such as background knowledge, examination points and question requirements can be input into a pre-trained question quality assessment model, and the total score, sub-item score and reason data output by the question quality assessment model are determined as the evaluation result. Performing question quality assessment on question data can avoid the appearance of low-quality data in the evaluation set and ensure the accuracy of the evaluation set.
[0100] Step 103: In response to the evaluation result that the question is qualified, the question data is stored in the question bank, an evaluation set is obtained, and new keywords are acquired; wherein the evaluation set is used to evaluate the application performance of the large language model in the data scheduling scenario.
[0101] In specific implementation, for high-quality question data with qualified evaluation results, the question data is directly stored in the question bank. All the question data in the question bank constitute an evaluation set, and the obtained evaluation set can be used to evaluate the application performance of the large language model in the data scheduling scenario.
[0102] Step 104: In response to the evaluation result that the question is unqualified, determine the number of prompt optimization times, regenerate the question data based on the number of prompt optimization times, and return to execute the step of performing question quality evaluation on the question data based on background knowledge, examination points and question requirements to obtain the evaluation result.
[0103] During specific implementation, if the evaluation result is that the problem is unqualified, it means that the problem data is unqualified. At this time, it is necessary to determine the number of prompt optimization times. If the number of prompt optimization times is less than the preset threshold, it means that the problem with the data is relatively small. The complete prompt, problem data and evaluation results can be input into a pre-trained prompt optimization model so that the prompt optimization model can optimize the complete prompt according to the problem data and evaluation results to obtain the optimized prompt, and regenerate the problem data according to the optimized prompt, and continue to perform quality evaluation on the regenerated problem data.
[0104] If the prompt optimization count exceeds the preset threshold, the question data is unqualified and has significant quality issues. In this case, the precondition data must be modified to determine a new, complete prompt and regenerate the question data based on the new, complete prompt. Precondition data includes background knowledge, key examination points, and question requirements. Modifying the precondition data involves modifying one or more of these. If unqualified question data is encountered, the system returns to the step of evaluating the question data's quality based on the background knowledge, key examination points, and question requirements to obtain the evaluation results. Generating new questions rather than rewriting the original questions further increases the diversity of questions in the evaluation set.
[0105] In some embodiments, for power grid dispatch scenarios, such as Figure 2 As shown in Figure 2, the method for automatically generating a large language model evaluation set can also be described in the following way:
[0106] S1. Obtain the question keywords. Keywords are obtained by traversing the domain keyword knowledge base used to generate the question. A keyword can be a word, phrase, or sentence that represents an entity or fact in the dispatching field and is stored in a database or knowledge base. For example, keywords may be "interprovincial and interregional power grid" or "power grid dispatching organizations are divided into five levels."
[0107] S2, Acquire Background Knowledge. The process of acquiring background knowledge involves searching for keyword-related text corpora in a vector database. First, using a knowledge enhancement method, professional documents in the scheduling field are imported into the vector database. Based on vector similarity, text corresponding to the keywords is retrieved, and multiple results are returned, each containing one or more paragraphs, as background knowledge. For example, the Milvus vector database can be selected, and the vector model can use bge-small-zh, with a paragraph length set to 500 characters.
[0108] S3, obtain the key points of the examination. The process of obtaining the key points of the examination is the process of using the A-model (A-model represents any large model used for key point generation) to analyze the key points of the examination from the background knowledge. The A-model receives a piece of background knowledge and automatically analyzes it according to the prompt to obtain the content that can be examined. The large model may generate multiple key points of knowledge, each of which is usually in the form of a question. For example, the A-model can use the Qwen2.5-7B model, and the initial prompt is "As an expert in the field of power grid dispatching, I now want to examine the staff's mastery of the background knowledge {background}. Based on this purpose, please analyze the knowledge points that are worth examining, and each knowledge point is given in the form of a short question." The short questions output by the Qwen2.5-7B model are the key points of the examination.
[0109] S4. Obtain the question requirements. The question requirements describe the large model generation task and include question type, question difficulty, and question examples. Question types include single-choice, multiple-choice, true / false, and short-answer questions; question difficulty is divided into three levels: easy, medium, and hard. Question examples are representative examples extracted from the historical scheduling sample question bank based on the selected question type and difficulty. These question examples provide references for text length, written format, and language style during the generation of the assessment dataset's question data.
[0110] S5, assemble prompts. Use the predefined "power grid dispatching expert" role template, substitute background knowledge, examination points and question requirements, and obtain the complete prompts required to generate questions. For example, the prompts can be: "As a power grid dispatching expert, I now want to examine the staff's mastery of the background knowledge {background}. After analysis, it is known that the background knowledge includes examination points {points}. Please prepare a question type of {question_type}, question difficulty of {question_difficulty}, and the question format, question composition, question description and other aspects can refer to the following question examples {question_examples}, and generate question data including (1) question, (2) answer, and (3) explanation."
[0111] S6: Generate a question. Input the complete prompt into the large model B (large model B represents any model used for question generation). Large model B outputs the question, answer, and explanation based on the complete prompt. For example, large model B can use the Qwen2.5-72B model.
[0112] S7, evaluate the quality of the questions. Use the large model C (large model C represents the model used for question quality evaluation) to combine background knowledge, examination points and question requirements to score the newly generated questions, answers and explanations. The scoring process is based on the following points: (1) whether the question is consistent with the examination points; (2) whether the answer and explanation can be extracted from the background knowledge; (3) whether the answer and explanation of the question are correct based on the background knowledge; (4) by comparing with the example questions, determine whether the question type meets the requirements; (5) refer to the example questions to further determine whether the difficulty of the question meets the requirements. Scoring is based on a percentage system, with 80 points or above being considered a pass. The large model C needs to provide sub-item scores and reasons. For example, the C model can adopt the CRITIQUELLM model, and the prompt for inputting the C model is: "I have just compiled a question. Please rate the quality of the question. The score is based on a percentage system, and 80 points or above are qualified. The score considers five aspects, each with 20 points. Please give the total score, each score and reason. The five aspects include (1)(2)(3)(4)(5). The compiled questions, corresponding answers and explanations are as follows: Question: {question}; Answer: {answer} Explanation: {explanation}. The following contents were referred to in the process of compiling the questions: Background knowledge: {background}; Examination points: {points}; Question type: {question_type}; Question difficulty: {question_difficulty}; Example: {question_examples}."
[0113] S8: Conditional judgment. For qualified questions, load them into the question bank, obtain the evaluation set, return to S1, and proceed to the next keyword question generation process. For unqualified questions, proceed to S9 for prompt optimization. For questions that still fail after two prompt optimizations, proceed to S10 to modify the preconditions. For keywords that still fail to generate questions after two modifications, skip the process and return to S1 to proceed to the next keyword.
[0114] S9, optimize prompts. For questions that need to be modified, call the large model D (representing a type of model used for prompt optimization) to optimize the prompts, and input: the original prompt words used by S5 to generate questions; the questions, answers and explanations generated by S6; the scores and reasons of S7. The large model D modifies the prompt words based on these three parts and re-enters S6 to regenerate questions. For example, the large model D can use Qwen2.5-7B, and the prompt words are: "Please help me optimize the following prompt words to help the large model better generate questions. Prompt words: {prompt}. During the optimization process, you need to refer to the original question, answer, explanation and evaluation of the question. Original question: {question}; answer: {answer}; explanation: {explanation}; question evaluation: {grade_and_reason}."
[0115] S10, modify the precondition. Based on the currently selected keyword, reselect a cooperation as a precondition from the retrieved or generated background knowledge, examination points, and question requirements, and re-enter S6 to generate the question.
[0116] It can be seen that the core points of the embodiment of this application are: using text corpus and business keywords in the field of power grid dispatching as input, combining retrieval enhancement generation technology with large model technology, and constructing a process that can automatically generate evaluation questions to achieve efficient evaluation of the large dispatching model. Its features include generating questions based on keywords and retrieved related content to ensure the domain relevance of the questions and the factual correctness of the answers; using multiple different large models to respectively implement inspection point analysis, question generation, quality assessment and prompt optimization. The technical characteristics of multiple different models ensure that each key step achieves the best effect; performing quality assessment on the questions and optimizing the generation process based on the assessment, further optimizing the quality of the generated questions.
[0117] In the process of generating the evaluation set,
[0118] First, we generate assessment question data based on keyword and knowledge retrieval enhancement technology. This retrieval-enhanced generation technology retrieves background knowledge from domain documents based on keywords, and then generates assessment questions and answers. This advantage lies in the fact that the generation process is guided by information such as keywords, rather than being completely random and uncontrollable. This makes the questions more relevant to professional domain knowledge requirements, increases the relevance of the assessment, and ensures the factual accuracy of the answers.
[0119] Then, multiple large models are designed and used for different generation stages based on the characteristics of each task.
[0120] Different large models are used to respectively realize inspection point analysis, question generation, quality assessment and prompt optimization. The models have different parameter scales and areas of expertise. Compared with using only one large model, the advantage is that targeted models are used in each key step to obtain the best generation effect. At the same time, the hardware requirements of the model are more reasonable and the operating efficiency is higher.
[0121] Finally, the results of question quality assessment, prompt optimization, and precondition modification are optimized;
[0122] The model evaluates the quality of questions and optimizes the generation process based on the evaluation results, significantly improving the quality of questions and ensuring usability.
[0123] In summary, the present application provides a method for automatically generating a large language model evaluation set. The search for professional documents related to keywords is performed in a retrieval-enhanced generation manner, and key questions are automatically analyzed from the retrieved content and answers are given. There is no need to manually input the original question as a starting condition, which reduces labor costs and generation difficulty. At the same time, it can utilize large-scale unlabeled corpora in the scheduling field to improve the coverage of professional field knowledge examinations. When the question is not appropriate, new question data can be generated in batches instead of rewriting the original question, further improving the diversity of questions. The generation process of question data is based on knowledge points in the form of keywords to ensure that the topic of the question is accurate and focused, and avoid the randomness of the generation of large language models. After entering the keywords, evaluation sets can be generated in batches by automated means, providing a benchmark for capability evaluation in the training and fine-tuning process of large scheduling models, improving the efficiency of model development and testing, and providing guarantees for the practicality of the model in professional fields. It promotes the widespread application of large language model technology in the scheduling field, improves the information retrieval and knowledge integration capabilities of scheduling work, alleviates the pressure on dispatchers to deal with emergencies, and improves the intelligence level of business processes.
[0124] In some embodiments, as Figure 3 As shown, background knowledge is determined based on keywords and search enhancement generation technology, including:
[0125] Step 301: Determine the technical field corresponding to the keyword, and import professional documents in the technical field into a vector database.
[0126] During specific implementation, the technical field corresponding to the keyword is first determined, so as to import professional documents in the technical field into the database to provide data support for the subsequent determination of background knowledge.
[0127] Step 302: Retrieve at least one text data corresponding to the keyword in the vector database based on vector similarity, integrate the at least one text data, and obtain background knowledge.
[0128] In practice, we search for text corresponding to keywords based on vector similarity, returning multiple results, each consisting of one or more paragraphs, as background knowledge. For example, we can select the Milvus vector database, use the bge-small-zh vector model, and set the paragraph length to 500 characters.
[0129] In some embodiments, as Figure 4 As shown, based on the initial prompts and background knowledge, determine the key points of the examination, including:
[0130] Step 401: Analyze the background knowledge content according to a preset background analysis model to obtain available examination content.
[0131] Step 402: Filter available examination contents according to the initial prompt to obtain at least one examination key point.
[0132] In specific implementation, the process of obtaining the key points of the examination is the process of using the A-model (A-model represents any large model used for generating key points) to analyze the key points of the examination from the background knowledge. The A-model receives a piece of background knowledge and automatically analyzes and obtains the content that can be examined based on the prompt. The large model may generate multiple key points of knowledge, each of which is usually in the form of a question. For example, the A-model can use the Qwen2.5-7B model, and the initial prompt is "As an expert in the field of power grid dispatching, I now want to examine the staff's mastery of the background knowledge {background}. Based on this purpose, please analyze the knowledge points that are worth examining, and each knowledge point is given in the form of a short question." The short questions output by the Qwen2.5-7B model are the key points of the examination.
[0133] In some embodiments, as Figure 5 As shown, based on the preset role template, assemble complete prompts according to background knowledge, examination points and question requirements, including:
[0134] Step 501: Determine the question type and question difficulty according to the question requirements, and determine question samples in the assessment set based on the question type and question difficulty;
[0135] Step 502: Bring the background knowledge and examination points into the role template, and determine the complete prompt containing the examination points based on the sample questions.
[0136] In specific implementation, the questions require a description of the large model generation task, including question type, question difficulty, and question examples. Question types include single-choice, multiple-choice, true / false, and short-answer questions; question difficulty is divided into three levels: easy, medium, and hard. Question examples are representative examples extracted from a historical dispatch example question bank based on the selected question type and difficulty. These example questions provide references for text length, written format, and language style in the generation process of the assessment dataset's question data. Then, using the pre-defined "Power Grid Dispatching Expert" role template, background knowledge, key assessment points, and question requirements are substituted to obtain the complete prompts needed to generate the questions.
[0137] In some embodiments, the question data includes a question, an answer to the question, and an explanation of the answer; generating the question data based on the complete prompt includes:
[0138] The complete prompt is input into the trained question generation model, so that the question generation model can determine the question according to the question requirements and examination points in the complete prompt, and determine the answer and explanation corresponding to the question according to the question requirements and background knowledge in the complete prompt, and integrate and output the question, answer and explanation.
[0139] In specific implementation, the complete prompt is input into the trained question generation model, and the question generation model outputs the question, answer and explanation based on the complete prompt. For example, the B-large model can adopt the Qwen2.5-72B model.
[0140] In some embodiments, as Figure 6 As shown in the figure, the question data is evaluated for quality based on background knowledge, examination points, and question requirements, and the evaluation results are obtained, including:
[0141] Step 601: Determine the assessment items corresponding to the question data based on background knowledge, examination points, and question requirements;
[0142] In specific implementation, the large model C (large model C represents the model used for question quality assessment) is used to score the newly generated questions, answers and explanations in combination with background knowledge, examination points and question requirements. The scoring process is based on the following points: (1) whether the question is consistent with the examination points; (2) whether the answer and explanation can be extracted from the background knowledge; (3) whether the answer and explanation of the question are correct based on the background knowledge; (4) by comparing with the example questions, whether the question type meets the requirements; (5) referring to the example questions, further judging whether the difficulty of the question meets the requirements.
[0143] Then step 601 includes:
[0144] Step 6011: Determine the degree of consistency between the questions and the examination points as the first evaluation item.
[0145] During specific implementation, the first assessment item is used to assess whether the question is consistent with the inspection points. The higher the degree of consistency, the higher the score.
[0146] Step 6012: Determine the similarity between the answer, explanation, and background knowledge as the second evaluation item.
[0147] In specific implementation, the second assessment item is used to evaluate whether the answers and explanations can be extracted from background knowledge. The higher the similarity, the higher the score.
[0148] Step 6013: Evaluate the correctness of the answer and explanation based on background knowledge, and determine the correctness as the third scoring item.
[0149] In specific implementation, the third assessment item is used to evaluate whether the answers and explanations inferred based on background knowledge are correct. The higher the degree of correctness, the higher the score.
[0150] Step 6014: Determine the question type according to the question requirements, evaluate the type similarity of the question according to the question type, and determine the type similarity as the fourth scoring item.
[0151] In specific implementation, the fourth assessment item is used to evaluate whether the question type meets the requirements by comparing it with the example questions. The higher the type similarity, the higher the score.
[0152] Step 6015: Determine the sample question and the difficulty of the question according to the question requirements, and evaluate the difficulty accuracy of the question based on the sample question and the difficulty of the question, and determine the difficulty accuracy as the fifth scoring item.
[0153] In specific implementation, the fifth assessment item is used to evaluate whether the difficulty of the judgment question meets the requirements. The higher the degree of difficulty, the higher the score.
[0154] Step 602: Perform sub-item evaluation on the problem data package according to the evaluation items, obtain multiple sub-item scores and score reasons corresponding to the sub-item scores, and determine the sum of the sub-item scores as the total score. Integrate the sub-item scores, score reasons and total score to obtain the evaluation result.
[0155] In specific implementation, the sub-item evaluation process adopts a percentage system, with scores of 80 or above being considered qualified. The large model C needs to provide sub-item scores and reasons. For example, the C large model can adopt the CRITIQUELLM model, and the prompt for inputting the large model C is: "I have just compiled a question. Please rate the quality of the question. The scoring adopts a percentage system, with scores of 80 or above being considered qualified. The scoring considers five aspects, 20 points each. Please provide the total score, each score and reason. These five aspects include the first assessment item, the second assessment item, the third assessment item, the fourth assessment item, and the fifth assessment item. The compiled questions, corresponding answers and explanations are as follows: Question: {question}; Answer: {answer} Explanation: {explanation}. The following content was referred to in the process of compiling the questions: Background knowledge: {background}; Examination points: {points}; Question type: {question_type}; Question difficulty: {question_difficulty}; Example: {question_examples}."
[0156] After receiving the evaluation results, qualified questions are loaded into the question bank, generating an evaluation set, re-determining keywords, and then entering the next keyword question generation process. For unqualified questions, prompt optimization is performed. For questions that still fail after two prompt optimizations, the preconditions are modified. For keywords that still fail after two precondition modifications, the process of generating the next keyword question is continued.
[0157] In some embodiments, as Figure 7 As shown, the question data is regenerated based on the number of prompt optimizations, including:
[0158] Step 701: In response to the prompt word optimization times being less than a preset threshold, based on a preset prompt optimization model, the complete prompt word is optimized according to the question data and the evaluation result to obtain an optimized prompt word, and the question data is regenerated according to the optimized prompt word, and the number of prompt word optimization times is increased by one to obtain a new prompt word optimization times, and when the new evaluation result is that the question is unqualified, the question data is regenerated according to the new prompt word optimization times.
[0159] During specific implementation, if the number of prompt optimizations is less than the preset threshold, it means that the problem with the data is relatively small, and the complete prompt, question data, and evaluation results can be input into a pre-trained prompt optimization model so that the prompt optimization model can optimize the complete prompt based on the question data and evaluation results, and regenerate the question data based on the optimized prompt. Exemplarily, a large model D (representing a type of model used for prompt optimization) is called to optimize prompts, and the input is: the original prompt words used to generate questions; questions, answers, and explanations; scores and reasons. Based on these three parts, the large model D modifies the prompt words and regenerates the question. Exemplarily, the large model D can use Qwen2.5-7B, and the prompt words are: "Please help me optimize the following prompt words to help the large model better generate questions. Prompt words: {prompt}. During the optimization process, you need to refer to the original question, answer, explanation, and evaluation of the question. Original question: {question}; answer: {answer}; explanation: {explanation}; question evaluation: {grade_and_reason}."
[0160] After regenerating the question data according to the optimized prompt, set the prompt optimization times = prompt optimization times + 1. When the new evaluation result corresponding to the regenerated question data is that the question is unqualified, regenerate the question data according to the new prompt optimization times to achieve continuous optimization of the question data to meet user needs.
[0161] Step 702: In response to the prompt being optimized more than or equal to a preset threshold, a new complete prompt is determined by modifying the precondition data, and question data is regenerated based on the new complete prompt; wherein the precondition data includes background knowledge, examination points, and question requirements.
[0162] During specific implementation, if the number of prompt optimizations is greater than or equal to the preset threshold, it means that the question data is unqualified and there are major quality issues. At this time, it is necessary to re-select a collaboration as a prerequisite from the currently selected keywords, among the retrieved or generated background knowledge, examination points and question requirements, and regenerate the question.
[0163] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario and performed by multiple devices working together. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the method.
[0164] It should be noted that the above description is limited to some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0165] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a device for automatically generating a large language model evaluation set.
[0166] refer to Figure 8 The large language model evaluation set automatic generation device comprises:
[0167] The precondition determination module 10 is configured to: obtain keywords, determine background knowledge based on the keywords and the search enhancement generation technology, determine the key points of the examination based on the initial prompt and the background knowledge, and determine the question requirements corresponding to the key points of the examination;
[0168] The question generation and evaluation module 20 is configured to: assemble a complete prompt based on a preset role template according to background knowledge, examination key points, and question requirements; generate question data based on the complete prompt; and perform a quality evaluation on the question data based on the background knowledge, examination key points, and question requirements to obtain an evaluation result;
[0169] The qualified question processing module 30 is configured to: in response to the evaluation result that the question is qualified, store the question data in an evaluation set and obtain new keywords; wherein the evaluation set is used to evaluate the application performance of the large language model in the data scheduling scenario;
[0170] The unqualified question processing module 40 is configured to: in response to the evaluation result that the question is unqualified, determine the number of prompt optimization times, regenerate the question data according to the number of prompt optimization times, and return to execute the question quality evaluation of the question data based on background knowledge, examination points and question requirements to obtain the evaluation result.
[0171] Optionally, the precondition determination module 10 includes:
[0172] The document importing unit is configured to: determine the technical field corresponding to the keyword and import the professional documents in the technical field into the vector database;
[0173] The data retrieval unit is configured as follows: the document importing unit is configured as follows: searching at least one text data corresponding to the keyword based on vector similarity in the vector database, integrating the at least one text data, and obtaining background knowledge.
[0174] Optionally, the precondition determination module 10 further includes:
[0175] The inspection content determination unit is configured to: perform content analysis on the background knowledge according to a preset background analysis model to obtain available inspection content;
[0176] The inspection key point determination unit is configured to filter available inspection content according to the initial prompt to obtain at least one inspection key point.
[0177] Optionally, the question generation and evaluation module 20 includes:
[0178] The question requirement parsing unit is configured to: determine a question type and a question difficulty according to the question requirement, and determine a question sample in the assessment set according to the question type and the question difficulty;
[0179] The prompt assembly unit is configured to: bring background knowledge and examination points into the role template, and determine a complete prompt containing the examination points based on the question sample.
[0180] Optionally, the question data includes a question, an answer to the question, and an explanation of the answer; the question generation and evaluation module 20 further includes:
[0181] The question data generation unit is configured to: input the complete prompt into the trained question generation model, so that the question generation model can determine the question according to the question requirements and examination points in the complete prompt, and determine the answer and explanation corresponding to the question according to the question requirements and background knowledge in the complete prompt, and integrate and output the question, answer and explanation.
[0182] Optionally, the question generation and evaluation module 20 further includes:
[0183] The assessment item determination unit is configured to: determine the assessment items corresponding to the question data according to the background knowledge, the examination points and the question requirements;
[0184] The evaluation score determination unit is configured to: perform sub-item evaluation on the problem data package according to the evaluation items, obtain multiple sub-item scores and score reasons corresponding to the sub-item scores, and determine the sum of the sub-item scores as the total score, integrate the sub-item scores, score reasons and total score to obtain the evaluation result.
[0185] Optionally, the question data includes questions, answers to the questions, and explanations of the answers; and the assessment item determination unit further includes:
[0186] The first item determination subunit is configured to: determine the consistency degree between the question and the examination points as the first evaluation item;
[0187] A second item determination subunit is configured to: determine the similarity between the answer, the explanation, and the background knowledge as a second assessment item;
[0188] A third item determination subunit is configured to: evaluate the correctness of the answer and explanation based on background knowledge, and determine the correctness as a third scoring item;
[0189] The fourth item determination subunit is configured to: determine the question type according to the question requirements, evaluate the type similarity of the question according to the question type, and determine the type similarity as the fourth scoring item;
[0190] The fifth item determination subunit is configured to: determine the question sample and question difficulty according to the question requirements, evaluate the difficulty correctness of the question based on the question sample and question difficulty, and determine the difficulty correctness as the fifth scoring item.
[0191] Optionally, the non-conformity problem processing module 40 includes:
[0192] The first processing module is configured to: in response to the prompt optimization number being less than a preset number threshold, optimize the complete prompt according to the question data and the evaluation result based on a preset prompt optimization model to obtain an optimized prompt, regenerate question data based on the optimized prompt, increase the number of prompt optimization times by one to obtain a new prompt optimization time, and regenerate question data based on the new prompt optimization time when the new evaluation result indicates that the question is unqualified;
[0193] The second processing module is configured to: in response to the prompt optimization number being greater than or equal to a preset number threshold, determine a new complete prompt by modifying the precondition data, and regenerate question data according to the new complete prompt;
[0194] Among them, the prerequisite data includes background knowledge, examination points and question requirements.
[0195] For the convenience of description, the above devices are described as being divided into various modules according to their functions. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0196] The apparatus of the above embodiment is used to implement the corresponding large language model evaluation set automatic generation method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0197] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements the method for automatically generating a large language model evaluation set as described in any of the above-mentioned embodiments.
[0198] Figure 9 10 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.
[0199] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0200] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0201] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0202] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0203] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0204] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0205] The electronic device of the above embodiment is used to implement the corresponding large language model evaluation set automatic generation method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0206] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method for automatically generating a large language model evaluation set as described in any of the above embodiments.
[0207] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0208] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the method for automatically generating a large language model evaluation set as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0209] It is understandable that before using the technical solutions of each embodiment of the present disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.
[0210] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operation of the disclosed technical solution based on the prompt message.
[0211] As an optional but non-limiting implementation, in response to a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0212] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0213] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application is limited to these examples. In line with the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.
[0214] In addition, for simplicity of description and discussion, and in order not to make the embodiment of the application difficult to understand, the known power supply / ground connection with integrated circuit (IC) chip and other components may or may not be shown in the accompanying drawings provided. In addition, the device can be shown in the form of a block diagram to avoid making the embodiment of the application difficult to understand, and this also takes into account the following fact, that is, the details of the embodiment of these block diagram devices are highly dependent on the platform to be implemented in the embodiment of the application (that is, these details should be fully within the scope of understanding of those skilled in the art). When specific details (for example, circuit) are set forth to describe exemplary embodiments of the application, it will be apparent to those skilled in the art that the embodiment of the application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.
[0215] Although the present invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may utilize the embodiments discussed.
[0216] The embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the present application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of the present application.
Claims
1. A method for automatically generating a large language model evaluation set, characterized in that: include: Obtain keywords, determine background knowledge based on the keywords and search enhancement generation technology, determine examination points based on the initial prompt and the background knowledge, and determine the question requirements corresponding to the examination points; Based on a preset role template, assembling a complete prompt according to the background knowledge, the examination key points, and the question requirements, generating question data according to the complete prompt, and performing a question quality assessment on the question data according to the background knowledge, the examination key points, and the question requirements to obtain an assessment result; In response to the evaluation result that the question is qualified, the question data is stored in a question bank to obtain an evaluation set, and new keywords are obtained; wherein the evaluation set is used to evaluate the application performance of the large language model in the data scheduling scenario; In response to the evaluation result that the question is unqualified, the number of prompt optimization times is determined, and the question data is regenerated according to the number of prompt optimization times, and the step of returning to execute the question quality evaluation of the question data based on the background knowledge, the examination points and the question requirements to obtain the evaluation result is returned.
2. The method for automatically generating a large language model evaluation set according to claim 1, characterized in that: The determining of background knowledge based on the keywords and the search enhancement generation technology includes: Determine the technical field corresponding to the keyword, and import professional documents in the technical field into a vector database; At least one text data corresponding to the keyword is retrieved from the vector database based on vector similarity, and the at least one text data is integrated to obtain the background knowledge.
3. The method for automatically generating a large language model evaluation set according to claim 1, characterized in that: Determining the examination points based on the initial prompt and the background knowledge includes: Performing content analysis on the background knowledge according to a preset background analysis model to obtain usable examination content; The available examination contents are screened according to the initial prompt to obtain at least one examination key point.
4. The method for automatically generating a large language model evaluation set according to claim 1, characterized in that: The complete prompt is assembled based on the preset role template according to the background knowledge, the examination points and the question requirements, including: Determining a question type and a question difficulty according to the question requirements, and determining a question sample in the assessment set according to the question type and the question difficulty; The background knowledge and the examination points are introduced into the role template, and the complete prompt containing the examination points is determined according to the question sample.
5. The method for automatically generating a large language model evaluation set according to claim 1, characterized in that: The question data includes a question, an answer to the question, and an explanation of the answer; Generating question data according to the complete prompt includes: The complete prompt is input into a trained question generation model, so that the question generation model determines the question according to the question requirements and examination points in the complete prompt, and determines the answer and the explanation corresponding to the question according to the question requirements and the background knowledge in the complete prompt, and integrates and outputs the question, the answer, and the explanation.
6. The method for automatically generating a large language model evaluation set according to claim 1, characterized in that: The question quality assessment of the question data is performed based on the background knowledge, the examination key points and the question requirements to obtain an assessment result, including: Determine the assessment items corresponding to the question data according to the background knowledge, the examination points and the question requirements; The problem data packet is evaluated item by item according to the evaluation items to obtain multiple item scores and score reasons corresponding to the item scores, and the sum of the item scores is determined as the total score. The item scores, the score reasons and the total score are integrated to obtain the evaluation result.
7. The method for automatically generating a large language model evaluation set according to claim 6, characterized in that: The question data includes a question, an answer to the question, and an explanation of the answer; The step of determining the assessment items corresponding to the question data based on the background knowledge, the examination points, and the question requirements includes: Determining the consistency between the question and the inspection points as the first evaluation item; determining the similarity between the answer, the explanation, and the background knowledge as a second evaluation item; evaluating the correctness of the answer and the explanation based on the background knowledge, and determining the correctness as a third scoring item; Determining a question type according to the question requirement, evaluating the type similarity of the question according to the question type, and determining the type similarity as a fourth scoring item; Determine the sample questions and the difficulty of the questions according to the question requirements, evaluate the correctness of the difficulty of the questions according to the sample questions and the difficulty of the questions, and determine the correctness of the difficulty as the fifth scoring item.
8. The method for automatically generating a large language model evaluation set according to claim 1, wherein: The regenerating question data according to the number of optimization times of the prompt includes: In response to the prompt optimization number being less than a preset number threshold, optimizing the complete prompt according to the question data and the evaluation result based on a preset prompt optimization model to obtain an optimized prompt, regenerating question data according to the optimized prompt, adding one to the number of prompt optimization times to obtain a new prompt optimization number, and regenerating question data according to the new prompt optimization number when a new evaluation result indicates that the question is unqualified; In response to the prompt being optimized a number of times greater than or equal to a preset number threshold, determining a new complete prompt by modifying the precondition data, and regenerating question data according to the new complete prompt; The prerequisite data includes the background knowledge, the examination points and the question requirements.
9. A device for automatically generating a large language model evaluation set, characterized in that: include: The precondition determination module is configured to: obtain keywords, determine background knowledge based on the keywords and the search enhancement generation technology, determine examination points based on the initial prompt and the background knowledge, and determine the question requirements corresponding to the examination points; The question generation and evaluation module is configured to: assemble a complete prompt based on a preset role template according to the background knowledge, the examination key points, and the question requirements, generate question data based on the complete prompt, and perform a quality evaluation on the question data based on the background knowledge, the examination key points, and the question requirements to obtain an evaluation result; a qualified question processing module configured to: in response to the evaluation result that the question is qualified, store the question data in an evaluation set and obtain new keywords; wherein the evaluation set is used to evaluate the application performance of the large language model in the data scheduling scenario; The unqualified question processing module is configured to: in response to the evaluation result that the question is unqualified, determine the number of prompt optimization times, regenerate question data according to the number of prompt optimization times, and return to execute the question quality evaluation of the question data based on the background knowledge, the examination points and the question requirements to obtain the evaluation result.
10. The automatic generation device for a large language model evaluation set according to claim 9, characterized in that: Precondition determination module, including: A document importing unit is configured to: determine a technical field corresponding to the keyword, and import professional documents in the technical field into a vector database; The data retrieval unit is configured as: a document importing unit is configured as: retrieving at least one text data corresponding to the keyword based on vector similarity in the vector database, integrating the at least one text data, and obtaining the background knowledge.
11. The apparatus for automatically generating a large language model evaluation set according to claim 9, wherein: The precondition determination module also includes: The inspection content determination unit is configured to: perform content analysis on the background knowledge according to a preset background analysis model to obtain available inspection content; The examination key point determination unit is configured to: filter the available examination content according to the initial prompt to obtain at least one examination key point.
12. The automatic generation device for a large language model evaluation set according to claim 9, characterized in that: Question generation and evaluation module, including: a question requirement parsing unit, configured to: determine a question type and a question difficulty according to the question requirement, and determine a question sample in the assessment set according to the question type and the question difficulty; The prompt assembly unit is configured to: bring the background knowledge and the examination points into the role template, and determine the complete prompt containing the examination points based on the question sample.
13. The automatic generation device of a large language model evaluation set according to claim 9, characterized in that: The question data includes a question, an answer to the question, and an explanation of the answer; The question generation and evaluation module also includes: The question data generation unit is configured to: input the complete prompt into a trained question generation model, so that the question generation model determines the question according to the question requirements and examination points in the complete prompt, and determines the answer and the explanation corresponding to the question according to the question requirements and the background knowledge in the complete prompt, and integrates and outputs the answer and the explanation of the question.
14. The automatic generation device of a large language model evaluation set according to claim 9, characterized in that: The question generation and evaluation module also includes: An assessment item determination unit is configured to: determine an assessment item corresponding to the question data according to the background knowledge, the examination points and the question requirements; The evaluation score determination unit is configured to: perform item-by-item evaluation on the problem data packet according to the evaluation items, obtain multiple item scores and score reasons corresponding to the item scores, and determine the sum of the item scores as the total score, integrate the item scores, the score reasons and the total score to obtain the evaluation result.
15. The automatic generation device for a large language model evaluation set according to claim 14, characterized in that: The question data includes a question, an answer to the question, and an explanation of the answer; The assessment project determines the unit, which also includes: The first item determination subunit is configured to: determine the consistency degree between the question and the inspection key points as a first evaluation item; A second item determination subunit is configured to: determine the similarity between the answer, the explanation and the background knowledge as a second evaluation item; A third item determination subunit is configured to: evaluate the correctness of the answer and the explanation based on the background knowledge, and determine the correctness as a third scoring item; a fourth item determination subunit, configured to: determine a question type according to the question requirement, evaluate a type similarity of the question according to the question type, and determine the type similarity as a fourth scoring item; The fifth item determination subunit is configured to: determine the question sample and question difficulty according to the question requirements, evaluate the difficulty correctness of the question according to the question sample and the question difficulty, and determine the difficulty correctness as the fifth scoring item.
16. The apparatus for automatically generating a large language model evaluation set according to claim 9, wherein: Non-conformity problem handling module, including: The first processing module is configured to: in response to the prompt optimization number being less than a preset number threshold, optimize the complete prompt according to the question data and the evaluation result based on a preset prompt optimization model to obtain an optimized prompt, regenerate question data based on the optimized prompt, increase the number of prompt optimization times by one to obtain a new prompt optimization time, and regenerate question data based on the new prompt optimization time when a new evaluation result indicates that the question is unqualified; The second processing module is configured to: in response to the prompt optimization number being greater than or equal to a preset number threshold, determine a new complete prompt by modifying the precondition data, and regenerate question data according to the new complete prompt; The prerequisite data includes the background knowledge, the examination points and the question requirements.
17. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
18. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 8.
Citation Information
Cited By
Problem set generation method and device, storage medium and program product
CN121615623A
Question set generation method, device, storage medium, and program product
CN121615623B