Low-carbon energy fine tuning data set generation method and system based on Pareto multi-model collaborative evaluation mechanism

By adopting the data set generation method of Pareto multi-model collaborative evaluation mechanism in the field of low-carbon energy, the problem of lack of dedicated data sets in this field is solved, high-quality data generation and optimization are achieved, and the application capabilities of large language models are significantly improved.

CN119918656APending Publication Date: 2025-05-02CHINA DATANG GRP TECH INNOVATION CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411840421.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The lack of dedicated large-model fine-tuning data sets in the field of low-carbon energy has led to insufficient application of models in this field, and the existing data generation methods lack an evaluation system, feedback optimization mechanism and data interpretability.

Method used

A low-carbon energy fine-tuned data set generation method based on the Pareto multi-model collaborative evaluation mechanism is adopted to collect and clean multiple data sources in the low-carbon energy field, design prompt words to generate Q&A data, and ensure data quality through multi-dimensional evaluation and feedback optimization mechanisms.

Benefits of technology

The first large-scale model fine-tuning data set dedicated to the low-carbon energy field was generated, which significantly improved the knowledge mastery and intelligent performance of large-language models in this field, and provided reliable data support for low-carbon energy intelligent systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119918656A_ABST
    Figure CN119918656A_ABST
Patent Text Reader

Abstract

The invention discloses a low-carbon energy fine tuning data set generation method based on a Pareto multi-model collaborative evaluation mechanism. The method comprises the following steps: step 1, collecting low-carbon energy field data; step 2, carrying out preliminary cleaning, preprocessing and blocking on the collected low-carbon energy data to filter useless information in the low-carbon energy data; 3, designing cue words by using the data processed in the step 2 to generate question and answer pair data which can be used for fine tuning; 4, evaluating the data generated in the step 3 to ensure the quality of the generated data; and step 5, obtaining evaluation result feedback of the step 4, and performing dynamic optimization and iterative improvement on the data generation process according to the evaluation result feedback. According to the method, the first large model fine tuning data set special for the low-carbon energy field is established, the knowledge mastering ability and the intelligent performance of the large language model in the field are remarkably improved, and reliable data support is provided for development of a low-carbon energy intelligent system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of energy large models, and in particular to a method and system for generating a low-carbon energy fine-tuning data set based on a Pareto multi-model collaborative evaluation mechanism. Background Art

[0002] In the field of large language models, general large models (such as GPT-4, Claude, etc.) have demonstrated excellent language understanding and generation capabilities and are applicable to a variety of scenarios. However, despite the excellent performance of these models in general tasks, their application in the field of low-carbon energy still faces challenges. The lack of knowledge in this field makes it difficult for general models to accurately understand the terminology and complex situations of specific industries, which may cause the model to produce "hallucinations" and thus affect user decisions.

[0003] One of the main reasons for the lack of application of large models in the field of low-carbon energy is the lack of dedicated fine-tuning datasets. Traditional manual labeling methods are not only costly, but also may introduce personal biases during the labeling process, which will affect the quality and applicability of the data. In recent years, with the development of large language models, a data-centric solution has emerged that can effectively alleviate the scarcity of real-world data by synthesizing high-quality data.

[0004] Currently, there is no dedicated large-scale model fine-tuning dataset in the field of low-carbon energy. The existing data generation methods have the following shortcomings: First, there is a lack of an evaluation system for data generated in this field, and it is impossible to evaluate the quality and effectiveness of the generated data. Second, there is a lack of feedback optimization mechanism, and the generated data cannot be optimized, which makes it possible that the generated data may not meet the specific needs of the industry, thus affecting the performance of the model in practical applications. Third, the synthetic data is not interpretable enough, and it is difficult for users to understand the source and generation process of the generated data. This lack of transparency may cause users to have doubts during use, affecting their acceptance of the model and its output. Summary of the invention

[0005] The purpose of the present invention is to provide a method and system for generating a low-carbon energy fine-tuning dataset based on a Pareto multi-model collaborative evaluation mechanism. The method and system utilize the contextual understanding ability of a large language model and combine a variety of data sources in the field of low-carbon energy, including degree theses, industry reports, and professional books. The system converts documents in the field of low-carbon energy into a question-answer pair format that can be used for fine-tuning through prompt word engineering. The system ensures the quality of generated data through data collection, preliminary cleaning and preprocessing, data generation and evaluation. The system realizes dynamic optimization and iterative improvement through a feedback mechanism. The system improves the accuracy, relevance, and innovation of generated data through a hierarchical data generation structure, a few-sample prompt strategy, and a systematic data quality evaluation system. The system constructs an efficient and scalable framework for generating fine-tuning datasets in the field of low-carbon energy and generates the first fine-tuning dataset in the field of low-carbon energy.

[0006] The present invention provides a method for generating a low-carbon energy fine-tuning data set based on a Pareto multi-model collaborative evaluation mechanism, comprising the following steps:

[0007] Step 1: Collect data in the field of low-carbon energy;

[0008] Step 2: Preliminary cleaning, preprocessing and segmentation of the collected low-carbon energy data to filter out useless information;

[0009] Step 3: Using the data processed in step 2, design prompt words to generate question-answer pair data that can be used for fine-tuning;

[0010] Step 4, evaluating the data generated in step 3 to ensure the quality of the generated data;

[0011] Step 5: Obtain the evaluation result feedback of step 4, and dynamically optimize and iteratively improve the data generation process based on the evaluation result feedback.

[0012] Furthermore, the low-carbon energy field data includes academic papers, industry reports, professional books, policy and regulatory documents, infrastructure data, user feedback data, and market trend analysis reports in the field of low-carbon energy.

[0013] Furthermore, the step 2 comprises:

[0014] 1) Filter out useless information in low-carbon energy data to ensure the quality of subsequent data generation;

[0015] 2) For the massive amount of text data collected, a set of delimiters is used to split the text into smaller blocks based on a recursive block strategy, while meeting the input format of data generation and maintaining semantic integrity;

[0016] 3) Store the segmented text blocks in an efficient data structure for access and calling in subsequent data generation processes.

[0017] Furthermore, the step 3 comprises:

[0018] Based on the preset prompt words and segmented text as context information, a large language model is used to generate fine-tuning data in the field of low-carbon energy.

[0019] Furthermore, in step 3, the prompt word design adopts a few-sample prompt strategy to construct diverse prompt word templates to guide the large language model to generate data formats.

[0020] Furthermore, the step 4 comprises:

[0021] A multi-model collaborative scoring mechanism is adopted, and the Pareto optimization method is introduced to comprehensively evaluate the generated data based on the three evaluation indicators of relevance, innovation, and correctness;

[0022] Design corresponding evaluation agents for different evaluation indicators. The agent uses a large language model as the core and scores the generated question-answer pairs through prompt words.

[0023] Use the large language model to build an automated evaluation tool, call the large language model to quickly evaluate the generated data, score the generated data according to the set evaluation criteria, and provide quantitative and qualitative evaluation results.

[0024] Furthermore, in step 4, the Pareto optimization method is introduced to achieve the optimal balance among the three indicators of relevance, innovation, and correctness, and the optimization goal is set as:

[0025] Maximize:F=[S rel ,S cre ,S cor ]

[0026] Among them, S rel Score the relevance, S cre Score innovation, S cor Score for correctness.

[0027] The scoring values ​​come from the collaborative scoring of multiple models, and the scores are formed into a solution set X = {x1, x2, ..., x N}; Each solution x i Corresponding to a rating vector:

[0028] F(x i )=[S rel (x i ),S cre (x i ),S cor (x i )]

[0029] In the solution on the Pareto front, the comprehensive score S is further calculated overall , introduce the following method:

[0030] Calculate the weights based on the distribution of each rating:

[0031]

[0032] in

[0033] The overall rating is:

[0034]

[0035] According to the comprehensive score S overall , select the optimal solution x from the Pareto front solution set * :

[0036]

[0037] The Pareto optimal scoring results are fed back to guide data screening to ensure that only question-answer pairs with higher scores enter the dataset and are then used for further fine-tuning of the model.

[0038] Furthermore, among the evaluation indicators: the relevance is used to measure whether the generated answer is highly relevant to the question content and whether it provides a meaningful response; the innovation is used to evaluate whether the answer is unique and whether it can answer the question in an innovative way; the correctness is used to detect the accuracy of the answer to ensure that the generated content is correct and consistent with domain knowledge.

[0039] Furthermore, the step 5 comprises:

[0040] (1) Output data analysis:

[0041] Analyze the generated data and evaluation results, judge the quality of the generated data, set different thresholds for the three different evaluation indicators, and if all three meet the threshold requirements, it proves that the generated data is of high quality and is saved in the data set; for question-answer pairs that do not meet the requirements, enter the iterative optimization process;

[0042] (2) Iterative optimization process:

[0043] An iterative optimization strategy for data generation is formulated based on the evaluation results. For evaluation indicators that do not meet the threshold requirements, the prompt words and generation conditions are adjusted to rewrite the questions and regenerate the answers based on the original question and answer pairs to generate new question and answer pairs. The new question and answer pairs then enter the evaluation process again until all three evaluation indicators meet the threshold requirements and the question and answer pairs are saved in the dataset.

[0044] The present invention also provides a low-carbon energy fine-tuning data set generation system based on the Pareto multi-model collaborative evaluation mechanism, comprising:

[0045] Data collection module, used to collect data in the field of low-carbon energy;

[0046] The data preprocessing module is used to perform preliminary cleaning, preprocessing and segmentation of the collected low-carbon energy data to filter out useless information;

[0047] A data generation module is used to utilize the processed data to design prompt words and generate question-answer pair data that can be used for fine-tuning;

[0048] A data evaluation module, used to evaluate the generated data to ensure the quality of the generated data;

[0049] The feedback optimization module is used to obtain feedback on evaluation results and dynamically optimize and iteratively improve the data generation process based on the feedback on evaluation results.

[0050] Through the above scheme, through the low-carbon energy fine-tuning dataset generation method and system based on the Pareto multi-model collaborative evaluation mechanism, the first large model fine-tuning dataset dedicated to the low-carbon energy field was established, which significantly improved the knowledge mastery and intelligent performance of the large language model in this field, and provided reliable data support for the development of low-carbon energy intelligent systems.

[0051] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a flowchart of a method for generating a low-carbon energy field fine-tuning dataset based on a large language model according to an embodiment of the present invention;

[0053] Figure 2 It is the Json format of fine-tuning data according to an embodiment of the present invention;

[0054] Figure 3 is an application paradigm diagram for generating data based on a large language model according to an embodiment of the present invention;

[0055] Figure 4 is a schematic diagram of a data preprocessing module according to an embodiment of the present invention;

[0056] Figure 5 is a schematic diagram of a data generation module according to an embodiment of the present invention;

[0057] Figure 6 is a schematic diagram of a data evaluation module according to an embodiment of the present invention;

[0058] Figure 7 is a flowchart of data set feedback optimization according to an embodiment of the present invention;

[0059] Figure 8 4 is a diagram of a data generation framework based on a large language model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0060] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0061] Ginseng Figure 1 As shown, this embodiment provides a method for generating a fine-tuning dataset in the field of low-carbon energy based on a large language model, and the method includes the following steps:

[0062] Step S1, collecting data in the field of low-carbon energy;

[0063] Step S2, preliminarily cleaning, preprocessing and segmenting the collected low-carbon energy data to filter out useless information;

[0064] Step S3, using the data processed in step S2, designing prompt words to generate question-answer pair data for fine-tuning;

[0065] Step S4, evaluating the data generated in step S3 to ensure the quality of the generated data;

[0066] Step S5, obtaining the evaluation result feedback of step S4, and dynamically optimizing and iteratively improving the data generation process according to the evaluation result feedback.

[0067] This method can generate question-and-answer data covering a wide range of problem scenarios by collecting and cleaning multi-source data; it can ensure the high quality and applicability of data by screening the generated data through a strict multi-dimensional evaluation mechanism; it builds a closed-loop system for data generation and evaluation based on the feedback optimization mechanism to continuously optimize the data generation effect. This invention has established the first large model fine-tuning dataset dedicated to the field of low-carbon energy, which significantly improves the knowledge mastery and intelligent performance of large language models in this field, and provides reliable data support for the development of low-carbon energy intelligent systems.

[0068] In one embodiment of the present invention, the present invention needs to first collect data in the field of low-carbon energy, and according to the application scenarios and objectives of the large model, clearly identify the data sources (academic papers, etc.) that need to be collected to ensure that the data obtained can effectively support the training and optimization of the model.

[0069] Specifically, in the process of data collection, the specific sources of various types of data should be determined first. Academic papers provide peer-reviewed research results, which can provide important references for the theoretical basis and practical experience in the field of low-carbon energy. Industry reports are usually issued by authoritative organizations and contain information such as market trends, technological developments and policy analysis. These data can help the model better understand market dynamics and policy environment. Secondly, government statistical data is an important basis for policy formulation and implementation, containing multi-dimensional information such as energy production, consumption, and emissions. The accuracy and authority of such data can enhance the reliability of the model.

[0070] In one embodiment of the present invention, low-carbon energy field data is cleaned and divided into fixed-size text blocks through a block strategy to meet the format of the data generation module, including:

[0071] Content filtering and selection: The collected data contains a lot of useless information, such as citations in academic papers, style tags in web pages, etc. These useless information will affect the quality of subsequent data generation and need to be filtered out before data segmentation.

[0072] Data segmentation,For the massive text data collected, a set of delimiters (different between Chinese and English texts) is used to segment the text into smaller blocks based on a recursive segmentation strategy, which not only meets the input format of the data generation module, but also maintains the semantic integrity.

[0073] Block storage and management, storing the segmented text blocks in an efficient data structure to facilitate subsequent rapid access and call, ensuring that the data generation module can obtain the required information in a timely manner.

[0074] In one embodiment of the present invention, the present invention can convert various low-carbon energy field data collected into a data format that can be used for fine-tuning. The specific format is as follows: Figure 2 As shown. Among them, "question" refers to various questions that may be raised in the field of low-carbon energy, and "answer" is the corresponding answer provided by the model based on domain knowledge. Through this result-based data format, the present invention can support the fine-tuning of large models in the field of low-carbon energy.

[0075] Specifically, during the fine-tuning process, the large model will perform supervised learning based on this data format, enabling the model to better understand and generate answers that meet the needs of the low-carbon energy field. Specific application paradigms include Figure 3As shown in the figure. To achieve this goal, the "question" part covers the core issues in low-carbon energy applications, while the "answer" provides accurate answers based on scientific research, policy information and industry data, thereby providing the model with rich contextual knowledge. This formatted data structure enables the large model to capture the context, key concepts and industry terms in the field of low-carbon energy during the fine-tuning process, improving the model's knowledge mastery and answer accuracy in specific fields. At the same time, through this design, the fine-tuned model will be able to provide more targeted answers to complex issues involving low-carbon energy, thereby improving its intelligence level and practicality in low-carbon energy application scenarios.

[0076] In one embodiment of the present invention, in order to improve the validity and utilization of data, the collected data needs to be preprocessed before being converted into a fine-tuned data format. Figure 4 shown.

[0077] Specifically, the data cleaning process mainly includes removing invalid information such as duplicate text and useless text. Duplicate text includes content that appears repeatedly in multiple data sources, such as similar paragraphs in different reports or repeated statistical data obtained from different documents. Removing duplicate text can reduce data redundancy and prevent model deviation caused by multiple appearances of the same information during training. Useless text refers to content that has no practical value for large model training. Common useless texts include copyright statements, reference lists, annotation information in table contents, etc. In addition, low-relevance content, such as narratives or background introductions that are not related to low-carbon energy, will also be considered useless text and removed. The removal of these useless texts helps reduce data noise and improve the purity of model training data.

[0078] The available data obtained after cleaning will be processed into text blocks. The block strategy is different for Chinese and English texts. For Chinese text, due to the high information density of Chinese characters, the size of the text block is relatively small; for English text, its character information density is low, and the size of the text block is relatively large. At the same time, the recursive block method is used to ensure the semantic integrity of the text block.

[0079] Through this block segmentation strategy, the final generated block text can cover the core knowledge in the field of low-carbon energy, and at the same time facilitate the processing and learning of the large model during the fine-tuning process, effectively improving the model's knowledge mastery and response capabilities in the field of low-carbon energy.

[0080] In one embodiment of the present invention, after obtaining the divided text, it is necessary to design reasonable prompt words (the format of guiding the model to generate data) to call the existing general large model to convert the text block into a fine-tuned question-answer pair, that is, a data generation module, such as Figure 5 shown.

[0081] Specifically, the present invention will first design different levels of prompt word templates based on the core knowledge points, key terms and typical problems in the field of low-carbon energy to ensure that the model can focus on information-intensive content when understanding text blocks. At the same time, in order to improve the diversity of question-answer pairs, the system introduces a sample pool mechanism, which helps the model obtain more references when generating new question-answer pairs by adding representative question-answer examples (few-shot learning), so that the generated question-answer pairs are richer in content, more diverse in types, and cover a wider range of problem scenarios.

[0082] The sample pool includes not only some common low-carbon energy field questions, but also a small number of samples for different topics (such as policies and regulations, market trends, and technological innovations) so that the model can refer to typical answer patterns in these fields when generating diversified questions. This few-shot learning mechanism enables the generated question-answer pairs to be more in line with actual application scenarios and meet professional needs.

[0083] The final generated question-answer pairs are saved in a standardized format, providing an accurate and rich source of training data for subsequent fine-tuning of the large model, enabling the model to have stronger knowledge mastery and intelligent response capabilities in the field of low-carbon energy.

[0084] In one embodiment of the present invention, the quality of the data generated by the data generation module varies, so it is necessary to screen and score its quality through the data evaluation module to ensure that the final question-answer pair used for fine-tuning has high reliability and validity. Figure 6 As shown in Figure 2, the core of the data evaluation module is to conduct a comprehensive analysis of the generated question-answer pairs by setting reasonable evaluation indicators.

[0085] Specifically, the evaluation module first extracts the "question-answer" pairs from the generated question-answer pairs and inputs them into the evaluation process. The evaluation process includes three main indicators: Relevance, Creativity, and Correctness. Each indicator is used to evaluate the performance of the question-answer pair in different dimensions:

[0086] 1. Relevance: whether the generated data comes from the provided text: mainly used to measure whether the generated answer is highly relevant to the question content and whether it provides a meaningful response.

[0087] 2. Creativity, the novelty and originality of the generated data: evaluate whether the answer is unique and whether it can answer the question in an innovative way, thereby enriching the diversity of the data set.

[0088] 3. Correctness: whether the generated data is consistent with the facts: used to detect the accuracy of the answer and ensure that the generated content is correct and consistent with domain knowledge.

[0089] The evaluation process is carried out by calling agents based on general large models (such as OpenAI models or Zhipu large models). These agents are configured with dedicated prompt words to generate comprehensive scores for each question and answer pair through different scoring criteria, including Relevance Score, Creativity Score and Correctness Score. Taking relevance as an example, the core of the prompt word design is: Your task is to provide an "overall score" to evaluate how well you can answer the given question clearly and unambiguously in a given context. Please give your score on a scale of 1 to 5, where 1 means that the question cannot be answered at all in the given context, and 5 means that the question is clear and unambiguous in the context.

[0090] In order to avoid the problem of reduced innovation due to increased relevance and reduced accuracy due to increased innovation, the optimal balance between the three indicators is achieved, the concept of Pareto optimization is introduced, and the optimization goal is set as follows:

[0091] Maximize:F=[S rel ,S cre ,S cor ]

[0092] Among them, S rel Score the relevance, S cre Score innovation, S cor Rate accuracy.

[0093] The above-mentioned scoring values ​​come from the collaborative scoring of multiple models (GPT, Llama, mass spectrometry, etc.), and the scores are formed into a solution set X = {x1, x2, ..., x N}. Each solution x i Corresponding to a rating vector:

[0094] F(x i )=[S rel (x i ),S cre (x i ),S cor (x i )]

[0095] In the solution on the Pareto front, the comprehensive score S is further calculated overall , the following methods can be introduced:

[0096] Calculate the weights based on the distribution of each rating:

[0097]

[0098] in

[0099] The overall rating is:

[0100]

[0101] According to the comprehensive score S overall , select the optimal solution x from the Pareto front solution set * :

[0102]

[0103] These Pareto optimal scoring results will eventually be fed back to the system to guide data screening, ensuring that only question-answer pairs with higher scores will enter the data set and be used for further fine-tuning of the model. In this way, the system achieves strict control of data quality and effectively improves the professionalism and practicality of the fine-tuning data set.

[0104] In one embodiment of the present invention, the feedback optimization module is used to improve the overall quality of the question-answer pairs generated by the data generation module, such as Figure 7 shown.

[0105] Specifically, the data generation module first generates initial question-answer pairs based on text blocks and pre-trained large language models (LLM). However, since the automatically generated question-answer pairs may differ in relevance, creativity, and correctness, the data evaluation module conducts multi-dimensional quality assessments on these question-answer pairs, including content accuracy, answer innovation, and context relevance.

[0106] After the quality assessment, the data evaluation module will feed back the generated evaluation results to the feedback optimization module. If the evaluation results show that the question-answer pair does not meet the set quality standards, the feedback optimization module will analyze it and take a series of optimization measures. For example, it can use the few-shot learning method to select high-quality question-answer pairs from the sample pool as a reference to enhance the diversity of answers in the generation model. In addition, the feedback optimization module can also improve the performance of question-answer pairs in specific situations by modifying the prompt words of the generation model, adding contextual information, adjusting model parameters, etc.

[0107] The optimized question-answer pairs will be quality-checked again through the data evaluation module to ensure their professionalism, practicality and accuracy in the field of low-carbon energy. Once the data meets all quality requirements, it will be included in the fine-tuning data set in the field of low-carbon energy to provide high-quality data support for further fine-tuning of the model.

[0108] The introduction of this feedback optimization module not only builds a closed-loop optimization mechanism for data generation and evaluation, but also significantly improves the output quality of the data generation module, ensuring the accuracy, richness and applicability of the fine-tuning data set, thereby laying a solid data foundation for the research and application of low-carbon energy intelligent systems.

[0109] In summary, the beneficial effects of the present invention are:

[0110] 1. Data acquisition and cleaning: Through multi-channel data collection and preprocessing processes, high-quality data sources in the field of low-carbon energy are effectively screened. Through data cleaning and block processing, useless information and redundant data are removed to ensure that the model can obtain pure and highly relevant domain knowledge during the training process, greatly improving the utilization rate of data and the training effect of the model.

[0111] 2. Prompt word design and diversity generation: We designed different levels of prompt word templates for key issues and terms in the field of low-carbon energy, and combined them with the few-shot learning mechanism of the sample pool to generate diverse question-answer pairs. Through this process, the generated data not only covers the core issues of low-carbon energy, but is also innovative and extensive, providing rich data support for model fine-tuning, so that the model can more comprehensively grasp the knowledge in this field.

[0112] 3. Data quality assessment and screening: The data assessment module scores and screens the generated question-answer pairs through multi-dimensional indicators such as relevance, innovation, and correctness to ensure that only high-quality data is included in the fine-tuning dataset. The introduction of this module allows data quality to be strictly controlled, further improving the professionalism and practicality of the dataset and ensuring the performance of the model in specific fields.

[0113] 4. Feedback optimization mechanism: Through the feedback optimization module, dynamic optimization and iterative improvement of generated data are achieved. Through few-shot learning, prompt word adjustment and parameter optimization, the diversity and accuracy of generated data are further improved. This module builds a closed-loop mechanism for data generation and evaluation, so that the system can improve the overall quality of data in each iteration and provide continuously optimized data support for fine-tuning.

[0114] In order to implement the above embodiment, Figure 8 As shown, this embodiment also provides a fine-tuning dataset generation framework based on a large language model, including:

[0115] A data collection module 100 is used to obtain high-quality data in the field of low-carbon energy from multiple channels such as academic papers, industry reports and government statistics;

[0116] The data preprocessing module 200 is used to clean and segment the collected data to improve the data purity and enhance the learning effect of the model through a reasonable segmentation strategy;

[0117] The data generation module 300 is used to design prompt words according to the core issues in the field of low-carbon energy, and generate diversified question-answer pairs through a sample pool mechanism;

[0118] The data evaluation module 400 is used to evaluate the generated question-answer pairs in terms of relevance, innovation, and correctness, and select high-quality question-answer pairs to enter the fine-tuning dataset;

[0119] The feedback optimization module 500 is used to dynamically optimize and iteratively improve the generated question-answer pairs according to the data evaluation results.

[0120] The framework calls the data generation module, uses the prompt word engineering to synthesize data, and then calls the data evaluation module to evaluate the quality of the generated data, and uses the evaluation feedback mechanism to optimize the data that does not meet the quality requirements. The data generation module adopts a hierarchical structure, which includes four main parts: data input, prompt word generation, model calling, and output generation, to ensure the efficiency and scalability of the data generation process.

[0121] The feedback optimization mechanism calls the output of the generation module and the evaluation module to achieve iterative optimization of the generated data, including:

[0122] Output data analysis: Analyze the output results of the generation module and the evaluation module, judge the quality of the generated data, set different thresholds for the three different evaluation indicators, and if all three meet the threshold requirements, it proves that the generated data is of high quality and can be saved in the data set. For question-answer pairs that do not meet the requirements, enter the iterative optimization process.

[0123] Based on the evaluation results, an iterative optimization strategy for data generation is formulated. For evaluation indicators that do not meet the threshold requirements, a series of strategies such as rewriting questions and regenerating answers are performed on the basis of the original question and answer pairs by adjusting the prompt words and generation conditions to generate new question and answer pairs. After that, the new question and answer pairs will enter the evaluation process again until all three evaluation indicators meet the threshold requirements, and then the question and answer pairs can be saved in the data set.

[0124] According to the fine-tuning dataset generation framework based on the large language model of the embodiment of the present invention, the large language model technology is used to extract fine-tuned question and answer pairs from the low-carbon energy field data. At the same time, a fine-tuning dataset evaluation system is established, and the generated question and answer pairs are dynamically optimized and iteratively improved in combination with the feedback optimization mechanism to generate the first fine-tuning dataset in the low-carbon energy field.

[0125] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the technical principles of the present invention, and these improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A method for generating a low-carbon energy fine-tuning dataset based on a Pareto multi-model collaborative evaluation mechanism, characterized in that: The steps include: Step 1: Collect data in the field of low-carbon energy; Step 2: Preliminary cleaning, preprocessing and segmentation of the collected low-carbon energy data to filter out useless information; Step 3: Using the data processed in step 2, design prompt words to generate question-answer pair data that can be used for fine-tuning; Step 4, evaluating the data generated in step 3 to ensure the quality of the generated data; Step 5: Obtain the evaluation result feedback of step 4, and dynamically optimize and iteratively improve the data generation process based on the evaluation result feedback.

2. The method for generating a low-carbon energy fine-tuning dataset based on a Pareto multi-model collaborative evaluation mechanism according to claim 1 is characterized in that: The low-carbon energy data include academic papers, industry reports, professional books, policy and regulatory documents, infrastructure data, user feedback data, and market trend analysis reports in the field of low-carbon energy.

3. The method for generating a low-carbon energy fine-tuning dataset based on a Pareto multi-model collaborative evaluation mechanism according to claim 2 is characterized in that: The step 2 comprises: 1) Filter out useless information in low-carbon energy data to ensure the quality of subsequent data generation; 2) For the massive amount of text data collected, a set of delimiters is used to split the text into smaller blocks based on a recursive block strategy, while meeting the input format of data generation and maintaining semantic integrity; 3) Store the segmented text blocks in an efficient data structure for access and calling in subsequent data generation processes.

4. The method for generating a low-carbon energy fine-tuning dataset based on a Pareto multi-model collaborative evaluation mechanism according to claim 3 is characterized in that: The step 3 comprises: Based on the preset prompt words and segmented text as context information, a large language model is used to generate fine-tuning data in the field of low-carbon energy.

5. The method for generating a low-carbon energy fine-tuning dataset based on a Pareto multi-model collaborative evaluation mechanism according to claim 4 is characterized in that: In step 3, the prompt word design adopts a few-sample prompt strategy to build a variety of prompt word templates to guide the large language model to generate data format.

6. The method for generating a low-carbon energy fine-tuning dataset based on a Pareto multi-model collaborative evaluation mechanism according to claim 5 is characterized in that: The step 4 comprises: A multi-model collaborative scoring mechanism is adopted, and the Pareto optimization method is introduced to comprehensively evaluate the generated data based on the three evaluation indicators of relevance, innovation, and correctness; Design corresponding evaluation agents for different evaluation indicators. The agent uses a large language model as the core and scores the generated question-answer pairs through prompt words. Use the large language model to build an automated evaluation tool, call the large language model to quickly evaluate the generated data, score the generated data according to the set evaluation criteria, and provide quantitative and qualitative evaluation results.

7. The method for generating a low-carbon energy fine-tuning dataset based on a Pareto multi-model collaborative evaluation mechanism according to claim 6 is characterized in that: In step 4, the Pareto optimization method is introduced to achieve the optimal balance between the three indicators of relevance, innovation, and correctness, and the optimization goal is set as: Maximize:F=[S rel ,S cre ,S cor ] Among them, S rel Score the relevance, S cre Score innovation, S cor Score for correctness. The scoring values ​​come from the collaborative scoring of multiple models, and the scores are formed into a solution set X = {x1, x2, ..., x N }; Each solution x i Corresponding to a rating vector: F(x i )=[S rel (x i ),S cre (x i ),S cor (x i )] In the solution on the Pareto front, the comprehensive score S is further calculated overall , introduce the following method: Calculate the weights based on the distribution of each rating: in The overall rating is: According to the comprehensive score S overall , select the optimal solution x from the Pareto front solution set * : The Pareto optimal scoring results are fed back to guide data screening to ensure that only question-answer pairs with higher scores enter the dataset and are then used for further fine-tuning of the model.

8. The method for generating a low-carbon energy fine-tuning dataset based on a Pareto multi-model collaborative evaluation mechanism according to claim 7 is characterized in that: In the evaluation indicators: the relevance is used to measure whether the generated answer is highly relevant to the question content and whether it provides a meaningful response; the innovation is used to evaluate whether the answer is unique and whether it can answer the question in an innovative way; the correctness is used to detect the accuracy of the answer to ensure that the generated content is correct and consistent with domain knowledge.

9. The method for generating a low-carbon energy fine-tuning dataset based on a Pareto multi-model collaborative evaluation mechanism according to claim 8 is characterized in that: The step 5 comprises: (1) Output data analysis: Analyze the generated data and evaluation results, judge the quality of the generated data, set different thresholds for the three different evaluation indicators, and if all three meet the threshold requirements, it proves that the generated data is of high quality and is saved in the data set; for question-answer pairs that do not meet the requirements, enter the iterative optimization process; (2) Iterative optimization process: An iterative optimization strategy for data generation is formulated based on the evaluation results. For evaluation indicators that do not meet the threshold requirements, the prompt words and generation conditions are adjusted to rewrite the questions and regenerate the answers based on the original question and answer pairs to generate new question and answer pairs. The new question and answer pairs then enter the evaluation process again until all three evaluation indicators meet the threshold requirements and the question and answer pairs are saved in the dataset.

10. A low-carbon energy fine-tuning dataset generation system based on Pareto multi-model collaborative evaluation mechanism, characterized in that: include: Data collection module, used to collect data in the field of low-carbon energy; The data preprocessing module is used to perform preliminary cleaning, preprocessing and segmentation of the collected low-carbon energy data to filter out useless information; A data generation module is used to utilize the processed data to design prompt words and generate question-answer pair data that can be used for fine-tuning; A data evaluation module, used to evaluate the generated data to ensure the quality of the generated data; The feedback optimization module is used to obtain feedback on evaluation results and dynamically optimize and iteratively improve the data generation process based on the feedback on evaluation results.

Citation Information

Cited By

  • Water resource scheduling instruction fine tuning data set construction method and device, equipment and medium

    CN121745319A