Energy decision generation method and system based on large language model
Through the energy decision generation method based on large language model, energy data sets and reinforcement learning algorithms are used to solve the problems of low efficiency and generation of fictional regulations in the existing technology, and efficient and accurate energy decision generation is achieved.
Patent Information
- Application Number
- CN202510078515.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has time-consuming and labor-intensive and error-prone problems in energy decision management, and general large models and decision tree methods are prone to generate non-existent fictitious regulations when processing energy decision files, which affects the generation efficiency of energy decisions.
A method of energy decision generation based on large language models is proposed. By obtaining energy data sets, the large language model is fine-tuned, user input tasks are received, prompt words are generated through reinforcement learning algorithms, and the model is trained using reward model to generate efficient energy decisions.
It improves the generation efficiency of energy decisions, enhances the logical judgment ability, anti-interference ability and generalization ability of the model, reduces the "illusion" phenomenon, and ensures the accuracy and reliability of the extraction results.
Smart Images

Figure CN120013465A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method and system for generating energy decisions based on a large language model. Background Art
[0002] With the growing demand for clean energy, the rapid development of renewable energy, especially wind and solar energy, has become an important way to achieve a low-carbon economy and address climate change, by making and managing a large number of energy decisions to ensure the environmental and social sustainability of energy projects. In related technologies, the method of manually maintaining databases for energy decision management is not only time-consuming and labor-intensive, but also prone to errors. And general large models and decision trees are prone to generate non-existent fictitious regulations when processing energy decision documents, affecting the efficiency of energy decision generation.
[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the invention
[0004] The main purpose of the embodiments of the present application is to propose an energy decision generation method and system based on a large language model, which can improve the efficiency of energy decision generation.
[0005] To achieve the above-mentioned purpose, an embodiment of the present application proposes an energy decision generation method based on a large language model, the method comprising:
[0006] Obtain energy data sets to fine-tune the large language model and obtain the initial energy decision-making large model;
[0007] Receiving a task input by a user, and performing prompt word generation processing on the task input by the user according to a reinforcement learning algorithm to obtain a task prompt word;
[0008] Performing task generation processing on the initial energy decision-making model according to the task prompt words to obtain generated data;
[0009] Training the initial energy decision-making model according to the generated data through a reward model to obtain a target energy decision-making model;
[0010] The user input data is input into the target energy decision model to perform energy decision generation processing to obtain an energy decision.
[0011] In some embodiments, the step of obtaining an energy data set to fine-tune the large language model to obtain an initial energy decision large model includes the following steps:
[0012] A task data set is constructed based on the energy data set; the task data set includes multiple tasks, each task includes a question text and an answer text;
[0013] Inputting the question text into the large language model for question-answering processing to obtain a question-answering result;
[0014] Perform cross entropy loss function calculation on the question and answer result according to the answer text to obtain a loss value;
[0015] The parameters of the large language model are updated according to the loss value to obtain the initial energy decision large model.
[0016] In some embodiments, the step of constructing a task dataset based on the energy dataset comprises the following steps:
[0017] Performing keyword recognition and context screening processing on the energy data set to obtain a corpus data set;
[0018] According to the data generation prompt words, the corpus data set is input into the large language model for data generation processing to obtain the task data set; the task data set includes yes / no question and answer data, extractive question and answer data set and natural language inference data set.
[0019] In some embodiments, the step of receiving a task input by a user and performing prompt word generation processing on the task input by the user according to a reinforcement learning algorithm to obtain a task prompt word includes the following steps:
[0020] Performing prompt word matching processing on the user input task through an intelligent agent to obtain a matching result;
[0021] When the matching result is empty, performing prompt word generation processing on the user input task through the prompt word large model to obtain a first prompt word;
[0022] When the matching result is the current prompt word, adjusting the current prompt word according to the depth search algorithm or the breadth search algorithm to obtain a second prompt word;
[0023] Updating the first prompt word or the second prompt word to the current prompt word;
[0024] Inputting the user input task and the current prompt word into the initial energy decision-making model for generation processing to obtain prediction data;
[0025] A predicted reward score is calculated based on the predicted data, and a strategy update process is performed on the agent based on the predicted reward score, and the process returns to the step of matching the prompt word of the user input task through the agent until a preset condition is met, and the current prompt word is determined as the task prompt word.
[0026] In some embodiments, the step of performing prompt word generation processing on the user input task by using the prompt word big model to obtain the first prompt word includes the following steps:
[0027] Performing content recognition processing on the user input task through the prompt word large model to obtain recognized content;
[0028] Performing rule extraction and sorting processing on the identified content to obtain sorted content;
[0029] The first prompt word is obtained according to the sorted content definition, and the first prompt word includes role, background, skill, goal, constraint and output format.
[0030] In some embodiments, the training process of the initial energy decision model according to the generated data by the reward model to obtain the target energy decision model includes the following steps:
[0031] Constructing a standard data set to perform initialization training processing on the initial energy decision-making large model to obtain an initialization training model;
[0032] According to the reward model, the generated data is input into the initialization training model for reinforcement learning processing to obtain the target energy decision model.
[0033] In some embodiments, the step of constructing a standard data set to perform initialization training processing on the initial energy decision-making large model to obtain an initialization training model includes the following steps:
[0034] The standard data set is constructed according to the positive sample data and the negative sample data;
[0035] The initial energy decision-making large model is directly subjected to preference optimization processing according to the standard data set to obtain the initialization training model.
[0036] To achieve the above objectives, another aspect of the embodiment of the present application proposes an energy decision-making system based on a large language model, the system comprising:
[0037] The first module is used to obtain energy data sets to fine-tune the large language model and obtain the initial energy decision-making large model;
[0038] The second module is used to receive a task input by a user, and perform prompt word generation processing on the task input by the user according to a reinforcement learning algorithm to obtain a task prompt word;
[0039] The third module is used to perform task generation processing on the initial energy decision-making model according to the task prompt words to obtain generated data;
[0040] The fourth module is used to train the initial energy decision-making model according to the generated data through a reward model to obtain a target energy decision-making model;
[0041] The fifth module is used to input the user input data into the target energy decision model to perform energy decision generation processing to obtain an energy decision.
[0042] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above-mentioned method when executing the computer program.
[0043] To achieve the above objective, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0044] The embodiments of the present application include at least the following beneficial effects: The present application provides a method and system for generating energy decisions based on a large language model. The scheme fine-tunes the large language model by acquiring an energy data set to obtain an initial energy decision large model, and can enhance the large language model's ability to understand the logical relationship of energy industry content through the energy data set. In addition, the scheme receives a user input task, generates prompt words for the user input task according to a reinforcement learning algorithm, and obtains a task prompt word. The task prompt word can be used to extract feature information in the data set, so that the large language model can dynamically adjust the logical rules according to new data and information, and can adapt to the ever-changing environment. In addition, the scheme trains the initial energy decision large model according to the generated data through a reward model to obtain a target energy decision large model, which can enable the model to have stronger logical judgment capabilities, while enhancing anti-interference capabilities and generalization capabilities, and improving the efficiency of energy decision generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a flow chart of an energy decision-making method based on a large language model provided in an embodiment of the present application;
[0046] Figure 2 is a flowchart of constructing a task data set provided by an embodiment of the present application;
[0047] Figure 3 This is a flow chart of generating task prompt words provided by an embodiment of the present application;
[0048] Figure 4 It is a structural schematic diagram of an energy decision-making system based on a large language model provided in an embodiment of the present application;
[0049] Figure 5 It is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of systems and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.
[0051] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".
[0052] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0054] Before describing the embodiments of the present application in detail, some nouns and terms involved in the embodiments of the present application are first described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0055] 1) Large Language Model (LLM): refers to deep learning models that are trained with a large amount of data and can generate and understand natural language. These models usually have a large number of parameters and can perform well in a variety of natural language processing tasks.
[0056] 2) Logical Programming Prompt: It is a technical means to guide large language models to perform complex logical tasks by designing specific input prompts. These prompts usually contain clear instructions and structured information to help the model understand the task requirements and accurately extract and process specific feature information from the text.
[0057] 3) Direct Preference Optimization (DPO): This is a training method that optimizes the model's generation capabilities by comparing high-quality and low-quality outputs. When constructing a DPO dataset, it usually contains three parts: prompt (input question or instruction), chosen (high-quality output), and rejected (low-quality output). By learning these data pairs, the model improves its ability to generate high-quality outputs and reduces errors and hallucinations that may occur during the generation process, thereby improving the overall performance of the model.
[0058] In the related technologies, general database methods, large language models, and decision trees have many deficiencies in processing energy decision-making documents. These methods have obvious disadvantages in information processing capabilities, illusion control, complexity, speed, dependency, update cost, scalability, and automation.
[0059] For example, the methods have obvious limitations in dealing with ambiguous or conflicting information. For example, when using databases and open source large language models, these methods are not very effective in extracting accurate structured information from large texts. The regulations on wind and solar facilities in legal documents often have large heterogeneity and complex fallback value table formats, which makes data extraction difficult. In addition, when faced with specifications that stipulate both small and large wind energy systems or require comparison logic between multiple values, the complexity is beyond the scope of what can be handled.
[0060] In the relevant database approach, updating and maintaining the energy policy database requires a lot of manual effort, which not only increases costs but also limits the flexibility of the system. Manual updates are prone to fatigue-induced errors and are difficult to keep up with the rapidly changing policy environment. In contrast, by automating the extraction and update process with a large language model, maintenance costs can be significantly reduced, and the responsiveness and flexibility of the system can be improved. Although the relevant large language model performs well in processing text and extracting information, frequent calls to the large language model will result in slower processing. Especially when adopting a decision tree framework, each node may need to interact with the large language model, which increases the total processing time. This delay is particularly evident when processing large document sets, affecting the overall efficiency of the system.
[0061] Related methods usually split documents into smaller chunks to facilitate large language models to process. However, this approach may cause important information to be lost during the cutting process, because the large language model can only process one chunk at a time. This information loss may affect the accuracy of the extraction results, especially when dealing with complex terms that require contextual understanding, the extraction accuracy is low.
[0062] In order to implement the logic of automatically extracting energy policy information from legal documents, related research relies on a specific Python framework to build structures such as decision trees. Although this approach can effectively improve the extraction accuracy, it also means that the implementation of the system is highly dependent on programming technology, which may limit its application scope among users with non-technical backgrounds. In addition, the reliance on a specific programming framework may also increase the difficulty of maintaining the system.
[0063] During the feature extraction process, the large language model sometimes exhibits "hallucinations", that is, generates fictitious regulations that do not exist. This is more common in initial testing, especially when the large language model is required to output a large amount of text to faithfully reproduce the input prompt. Although this phenomenon can be largely reduced by explicitly instructing the large language model not to add extra text, extract only relevant small paragraphs, and implement heuristic N-Gram similarity checks, it still occurs 1% to 3% of the time in the final experiment. These "hallucinations" mainly manifest themselves in the large language model misapplying the real regulations to the wrong objects, such as assuming that a certain setback distance applies to buildings and property lines, when in fact it only applies to buildings.
[0064] In view of this, an energy decision generation method and system based on a large language model is provided in an embodiment of the present application. The scheme fine-tunes the large language model by acquiring an energy data set to obtain an initial energy decision large model, and can enhance the large language model's ability to understand the logical relationship of energy industry content through the energy data set. In addition, the scheme receives user input tasks, generates prompt words for the user input tasks according to a reinforcement learning algorithm, and obtains task prompt words. The feature information in the data set can be extracted through the task prompt words, so that the large language model can dynamically adjust the logical rules according to new data and information, and can adapt to the ever-changing environment. In addition, the scheme trains the initial energy decision large model according to the generated data through a reward model to obtain a target energy decision large model, which can enable the model to have stronger logical judgment capabilities, while enhancing anti-interference capabilities and generalization capabilities, and improving the efficiency of energy decision generation.
[0065] An energy decision-making method based on a large language model provided in an embodiment of the present application relates to the field of artificial intelligence technology. An energy decision-making method based on a large language model provided in an embodiment of the present application can be applied to a terminal, can be applied to a server, or can be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and can also be configured to provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and cloud servers for basic cloud computing services such as big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements an energy decision-making method based on a large language model, etc., but is not limited to the above forms.
[0066] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0067] Figure 1 is an optional flow chart of an energy decision-making method based on a large language model provided in an embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S105.
[0068] Step S101, obtaining an energy data set to fine-tune the large language model to obtain an initial energy decision-making large model;
[0069] Step S102, receiving a task input by a user, and performing prompt word generation processing on the task input by the user according to a reinforcement learning algorithm to obtain a task prompt word;
[0070] Step S103, performing task generation processing on the initial energy decision-making model according to the task prompt words to obtain generated data;
[0071] Step S104, training the initial energy decision-making model according to the generated data through a reward model to obtain a target energy decision-making model;
[0072] Step S105, inputting the user input data into the target energy decision big model to perform energy decision generation processing to obtain an energy decision.
[0073] Steps S101 to S105 shown in the embodiment of the present application are to fine-tune the large language model by acquiring an energy data set. By acquiring an energy data set, the energy data set may include data in multiple dimensions such as energy policy, planning, regulations, production, environment, and weather. Then, the large language model is trained and fine-tuned using the data set, and the model performance is optimized by continuous iteration, so that the large language model can enhance the logical relationship of understanding the industry content and obtain the initial energy decision-making large model. Then, the user input task is received, and the prompt word generation process is performed on the user input task according to the reinforcement learning algorithm to obtain the task prompt word. The generated task prompt word can identify and integrate the relationship between multiple factors in the input data, such as economy, environment, technology, etc., and support the model to identify more complex decision-making tasks. By inputting the task prompt word into the initial energy decision-making large model for task generation processing to obtain generated data, and the initial energy decision-making large model is trained and processed according to the data by the reward model, and the target energy decision-making large model is obtained. The target energy decision-making large model is a general decision-making large language model for the energy industry with automation capabilities, reliability, flexibility, and strong interpretability.
[0074] In step S101 of some embodiments, the energy data set is obtained to fine-tune the large language model to obtain an initial energy decision large model, including the following steps:
[0075] A task data set is constructed based on the energy data set; the task data set includes multiple tasks, each task includes a question text and an answer text;
[0076] Inputting the question text into the large language model for question-answering processing to obtain a question-answering result;
[0077] Perform cross entropy loss function calculation on the question and answer result according to the answer text to obtain a loss value;
[0078] The parameters of the large language model are updated according to the loss value to obtain the initial energy decision large model.
[0079] In an embodiment of the present application, the energy data set includes data of multiple dimensions such as energy policy, planning, regulations, production, environment, weather, etc. The task data set is a data set for task processing of the energy data set. The task data set is obtained by filtering out the required context or paragraph construction from the energy data set, which may include articles, paragraphs or other forms of text content, suitable for generating tasks that require understanding of text content, and the task data set is used to train the large language model. Among them, the task data set may include multiple tasks, each of which includes a question text and an answer text. Then the question text is input into the large language model for question and answer processing, and the large language model is a pre-trained large language model, such as qwen2.5, Mistral, llama and other models as the basic model. The answer text is output by the large language model, and then the cross entropy loss function is calculated and processed according to the answer text to obtain the loss value, so as to update the parameters of the large language model according to the loss value, and the updated parameters may include learning rate, batch size, number of training rounds, etc., so as to obtain the initial energy decision large model after training. It can be imagined that the embodiment of the present application can also divide the task data set into a training set, a validation set and a test set to evaluate the performance of the model. Fine-tune the model on the training set and optimize the model parameters using the cross entropy loss function. Regularly evaluate the performance of the model on the validation set to prevent overfitting.
[0080] One of the above technical solutions has the following advantages or beneficial effects: the embodiment of the present application can further improve the performance of the large language model on specific tasks by constructing a task data set to perform model fine-tuning on the large language model.
[0081] In some embodiments, the step of constructing a task dataset based on the energy dataset comprises the following steps:
[0082] Performing keyword recognition and context screening processing on the energy data set to obtain a corpus data set;
[0083] According to the data generation prompt words, the corpus data set is input into the large language model for data generation processing to obtain the task data set; the task data set includes yes / no question and answer data, extractive question and answer data set and natural language inference data set.
[0084] In an embodiment of the present application, a task data set is obtained by constructing an energy data set, and specifically, a corpus data set is obtained by filtering out the content that requires context or paragraphs to be completed from the energy data set. The corpus data set may include general energy industry corpus and energy industry regulatory corpus, wherein the general energy industry corpus can be collected through designated energy websites and industry websites. The energy industry regulatory corpus can collect documents related to energy construction regulations, including wind energy, solar energy, hydrogen energy and other data, by specifying a regulatory database or judging whether it is a regulatory data corpus based on keywords. The general corpus and regulatory corpus are then analyzed to obtain keywords related to the wind energy siting regulations of concern, such as energy capacity, plot or plot size, turbine height, spacing, setback distance, noise level and other data in documents related to solar energy siting. Please refer to Figure 2 , the embodiment of the present application takes the corpus dataset as input and provides data generation prompt words to the large language model for data generation processing. The large language model will give three different types of results, including a yes or no question answering (YNQA) dataset, an extractive question answering (ExQA) dataset, and a natural language inference (NLI) dataset. The generated yes or no question answering (YNQA) dataset, extractive question answering (ExQA) dataset, and natural language inference (NLI) dataset are determined as task datasets. The embodiment of the present application can define prompt words for different tasks separately, or use one prompt word for data generation for all tasks. In a feasible embodiment, the embodiment of the present application provides a data generation prompt word, and the content of the data generation prompt word is as follows:
[0085] Please generate three types of question-answering tasks based on the following text:
[0086] Please generate questions and answers in the following format:
[0087] 1. Yes-No Question Answering (YNQA) Task:
[0088] Question: [Ask a question that can be answered with "yes" or "no" based on the text];
[0089] Answer: [yes / no];
[0090] 2. Extractive Question Answering (ExQA) Task:
[0091] Question: [asks a question whose answer can be directly extracted from the text, usually involving specific data, numbers, or explicit statements];
[0092] Answer: [Answer extracted directly from the text];
[0093] 3. Natural Language Inference (NLI) Task:
[0094] Prerequisite: [Choose a key statement from the text];
[0095] Hypothesis: [make a reasonable assumption based on the premise];
[0096] Reasoning: [Judge whether the hypothesis is valid, the answer is "yes" or "no"];
[0097] Please ensure that all questions and answers are strictly based on the text content and avoid subjective speculation.
[0098] One of the above technical solutions has the following advantages or beneficial effects: the embodiment of the present application can generate high-quality tasks from various texts, which can meet the needs of different application scenarios. These tasks are strictly based on the text content, avoiding subjective speculation and ensuring the accuracy and reliability of the generated tasks.
[0099] In some embodiments, the step of receiving a task input by a user and performing prompt word generation processing on the task input by the user according to a reinforcement learning algorithm to obtain a task prompt word includes the following steps:
[0100] Performing prompt word matching processing on the user input task through an intelligent agent to obtain a matching result;
[0101] When the matching result is empty, performing prompt word generation processing on the user input task through the prompt word large model to obtain a first prompt word;
[0102] When the matching result is the current prompt word, adjusting the current prompt word according to the depth search algorithm or the breadth search algorithm to obtain a second prompt word;
[0103] Updating the first prompt word or the second prompt word to the current prompt word;
[0104] Inputting the user input task and the current prompt word into the initial energy decision-making model for generation processing to obtain prediction data;
[0105] A predicted reward score is calculated based on the predicted data, and a strategy update process is performed on the agent based on the predicted reward score, and the process returns to the step of matching the prompt word of the user input task through the agent until a preset condition is met, and the current prompt word is determined as the task prompt word.
[0106] In the embodiment of the present application, prompt words are generated for user input tasks according to the reinforcement learning algorithm. Figure 3, first receive the user input task, which can be text input or voice input. In a feasible embodiment, the user input task is: "In a certain county, a wind power facility is to be built. The support tower is 377 feet high, the blade is 279 feet long, the rotor diameter is 558 feet, and the total height of the system is 656 feet. What is the safe distance?" Then the intelligent agent performs prompt word matching processing on the user input task to obtain a matching result. Among them, the intelligent agent will build a state (state) according to the user input task, including the task description and the current existing prompt words, for example: S = {"Analyze the characteristics of wind energy site selection based on the energy regulations of the county: find rule-related features"}, S represents the state, and the intelligent agent will select an action (action) according to the current state. The actions include the following: generate a new prompt word (prompt), adjust the existing prompt word (prompt) from the perspective of deep search, and adjust the existing prompt word (prompt) from the perspective of breadth search. Among them, the generation of new prompt words can be generated by the prompt word large model, which is a model that obtains prompt words for guiding the large model to perform corresponding tasks according to the task type. When the matching result is empty, the prompt word generation process is performed on the user input task through the prompt word big model to obtain the first prompt word. When the matching result is a match to obtain the current prompt word, the current prompt word is adjusted and processed according to the depth search algorithm or the breadth search algorithm to obtain the second prompt word. Among them, the depth search algorithm can add steps, such as finding irrelevant information that needs to be excluded, finding numerical information, determining the type of numerical value, etc. The breadth search algorithm can increase the types of relevant features, such as structures, property lines, roads, railways, etc., or increase the types of irrelevant information that are excluded, such as ignoring all information about small, mini or private wind energy systems. Then the first prompt word or the second prompt word is updated to the current prompt word, and the user input task and the current prompt word are input into the initial energy decision big model for generation processing to obtain the predicted data. The prediction reward score can be calculated based on the accuracy of the predicted data and other performance indicators, for example: accuracy reward: if the prediction accuracy reaches 95%, the reward is +10; if the accuracy is less than 80%, the penalty is -5. The agent's strategy is updated according to the predicted reward score to optimize future decisions, and the process returns to the step of matching prompt words for the user input task through the agent until a preset condition is met, which may be the number of iterations, etc., and the current prompt word finally obtained is determined as the task prompt word.
[0107] One of the above technical solutions has the following advantages or beneficial effects: the embodiment of the present application generates task prompt words according to the reinforcement learning algorithm, so that the generated task prompt words can extract the required feature information from the data set, adapt to different task processing, and improve the efficiency of data generation.
[0108] In some embodiments, the step of performing prompt word generation processing on the user input task by using the prompt word big model to obtain the first prompt word includes the following steps:
[0109] Performing content recognition processing on the user input task through the prompt word large model to obtain recognized content;
[0110] Performing rule extraction and sorting processing on the identified content to obtain sorted content;
[0111] The first prompt word is obtained according to the sorted content definition, and the first prompt word includes role, background, skill, goal, constraint and output format.
[0112] In an embodiment of the present application, the prompt word big model performs content recognition processing on the user input task to obtain the recognized content, and the prompt word big model is used to read and understand the document content to identify the chapters or paragraphs related to energy site selection. The recognized content is then subjected to rule extraction and organization processing to obtain the organized content, and the prompt word big model is used to extract specific regulations and rules to ensure that each rule has a clear legal basis and document references. A clear list is formed by organizing the extracted rules, and an audit is performed to ensure accuracy. Finally, the prompt word big model will obtain the first prompt word according to the organized content definition. The first prompt word includes role, background, skills, goals, constraints and output format. In a feasible embodiment, the prompt word big model receives the user input task, and the format of the first prompt word outputted is as follows:
[0113] The role: Regulatory Analyst and Document Review Specialist
[0114] Background: Users need to extract regulatory rules on energy site selection from a large amount of documents, which requires a deep understanding of the regulatory content and accurate extraction capabilities.
[0115] Skills: You have the legal knowledge, document analysis skills, and information extraction techniques to quickly scan large amounts of text, identify key regulatory provisions, and accurately extract rules related to energy siting.
[0116] Objective: Accurately extract all regulatory rules related to energy siting from documents and provide users with clear and accurate regulatory information.
[0117] Constraints: The extracted rules must be accurate and must not contain any irrelevant or misleading information, while ensuring the integrity and legal validity of the information. Only commercial large-scale facility requirements are extracted, excluding private and public facility requirements.
[0118] Output format: The output json result provides a clear list of regulatory rules. Each rule should contain a reference to the source document to ensure the traceability of the information.
[0119] Among them, this JSON object can contain the keys in Table 1, which is a JSON object content table, as shown in the following table:
[0120]
[0121] Table 1
[0122] One of the above technical solutions has the following advantages or beneficial effects: the embodiment of the present application can systematically extract the required characteristic information from the document in the field of energy regulations by generating prompt words. This method not only improves the efficiency of information extraction, but also enhances the accuracy and availability of information, providing a solid foundation for subsequent regulatory implementation and supervision.
[0123] In some embodiments, the training process of the initial energy decision model according to the generated data by the reward model to obtain the target energy decision model includes the following steps:
[0124] Constructing a standard data set to perform initialization training processing on the initial energy decision-making large model to obtain an initialization training model;
[0125] According to the reward model, the generated data is input into the initialization training model for reinforcement learning processing to obtain the target energy decision model.
[0126] In an embodiment of the present application, the goal of the reward model is to automatically extract feature data from the energy data set, such as the maximum system height, minimum backoff distance, noise limit and other regulations for wind power projects. The energy decision-making model is then evaluated based on accuracy, completeness and consistency. For example, the extracted wind power project regulations must be completely consistent with the description in the document; ensure that all relevant wind power project regulations are correctly extracted; and correctly identify and classify different statements of the same regulation in different documents. In an embodiment of the present application, a standard data set is first constructed to perform initialization training processing on the initial energy decision-making model to obtain an initialization training model. Then, according to the reward model, the generated data is input into the initialization training model for reinforcement learning processing. The goal of reinforcement learning is to use the reinforcement learning model to allow the model to explore more complex and more diverse business scenarios on its own. The specific mathematical expression of the reward model is as follows:
[0127] R(s,a)=f(θ,x);
[0128] Among them, R(s,a) represents the reward obtained by taking action a in state s, f(θ,x) is a function defined by parameter θ, and input x includes but is not limited to text generated by the model, context information, etc. Among them, the state refers to the answer or explanation generated by the large language model based on a given question or context. The state can include a direct answer to the question, an explanation of related concepts, a reference to specific data, etc. The reinforcement learning model will select actions based on the data type of the state. For example, for positive sample data, the action includes extracting the original text or conducting in-depth association analysis, covering the identification of key terms, clarifying the target object, understanding the context, etc.; for negative sample data, the action is mainly to identify and analyze various error types, such as degree errors, conditional errors, etc. The selection action of negative sample data is to explore different error types, such as degree errors, conditional errors, subject object errors, numerical errors, lack of integrity, etc.
[0129] In a feasible embodiment, the selection actions of positive sample data can be: extracting original text, correlation analysis, concept adaptation scenario distinction, multi-criteria decision analysis (MCDA), scenario planning, case comparison, risk assessment, etc.
[0130] For association analysis, such as:
[0131] (1) Identify key terms: Identify the professional words or phrases in the document that are directly related to the problem and understand their basic meaning.
[0132] (2) Identify the target audience: Identify the specific entity (e.g., wind turbine, temporary meteorological tower) or concept (e.g., noise limit, energy capacity, etc.) that the problem is concerned with.
[0133] (3) Understand the context: Consider the context in which the regulation or standard being discussed was developed, including geographic location, industry characteristics, etc.
[0134] (4) Determine the scope of application: Find out what types of projects or activities a particular regulation is intended for.
[0135] (5) Understand the relevant rule system: If there are multiple interrelated regulations, you need to understand how the entire rule system is constructed.
[0136] (6) Compare the differences in different scenarios: Analyze the similarities and differences in the requirements of the same concept in different application scenarios (for example, different types of wind energy systems).
[0137] (7) Evaluate influencing factors: Explore the impact of various external conditions (such as weather changes, land use patterns, etc.) on the implementation of a certain regulation.
[0138] (8) Explore exceptions: Identify whether there are certain special situations in which the original rules may not be fully applicable or need to be adjusted.
[0139] The embodiment of the present application performs reinforcement learning on the initialization training model through the reinforcement learning model, outputs the answer through the initialization training model, and performs reward evaluation on the answer. The model is updated according to the evaluation results until the preset conditions are met, thereby obtaining the target energy decision-making model that is finally trained.
[0140] In order to define the reward function more specifically, the embodiment of the present application can set specific scoring criteria for each quality attribute (accuracy, completeness, relevance). The following are the detailed scoring rules and corresponding scores:
[0141] (1) Accuracy
[0142] Completely accurate: The generated text information is accurate, factual, and logical. Score: +1.0;
[0143] Partially accurate: The generated text is mostly accurate, but contains some minor errors or is not completely accurate. Score: +0.5;
[0144] Inaccurate: Generated text contains obvious errors or factual inaccuracies. Score: -1.0;
[0145] (2) Integrity
[0146] Complete Coverage: The generated answer covers all necessary details without missing any critical information. Score: +1.0;
[0147] Partial Coverage: The generated answer covers most of the necessary details but leaves out some critical information.
[0148] Score: +0.5;
[0149] Not Covered: The generated answer leaves out a lot of critical information. Score: -1.0;
[0150] (3) Relevance
[0151] Highly relevant: The generated content is closely related to the questioner's needs and directly responds to the question. Score: +1.0;
[0152] Partially relevant: The generated content is somewhat relevant to the questioner’s needs, but not directly or comprehensively.
[0153] Score: +0.5;
[0154] Irrelevant: The generated content is irrelevant or off-topic to the questioner’s needs. Score: -1.0;
[0155] Reward function formula:
[0156] The reward function R(s,a) can be expressed as:
[0157] R(s,a)=w(accuracy)+w(completeness)+w(relevance);
[0158] Among them, w(accuracy) represents the accuracy score, w(completeness) represents the completeness score, and w(relevance) represents the relevance score.
[0159] One of the above technical solutions has the following advantages or beneficial effects: By introducing a reward model method, the embodiment of the present application can effectively reduce the probability of "hallucination" in the large language model during feature extraction. The reliability and accuracy of the data are further improved, ensuring the authenticity and credibility of the extraction results.
[0160] In some embodiments, the step of constructing a standard data set to perform initialization training processing on the initial energy decision-making large model to obtain an initialization training model includes the following steps:
[0161] The standard data set is constructed according to the positive sample data and the negative sample data;
[0162] The initial energy decision-making large model is directly subjected to preference optimization processing according to the standard data set to obtain the initialization training model.
[0163] In an embodiment of the present application, a standard data set is constructed based on positive sample data and negative sample data. The positive sample data is the content that is consistent with the original paragraph extracted based on keywords. These contents will be used as positive samples for training model recognition and generating high-quality outputs. Negative sample data are answers that contain hallucinations, wrong information, or logical incoherence. These answers will be used as negative samples to help the model learn to distinguish and avoid generating low-quality outputs. Finally, the initial energy decision-making large model is directly optimized based on the standard data set to obtain an initialized training model. The direct preference optimization goal is to increase the logarithmic probability of the positive sample data response and reduce the logarithmic probability of the negative sample data response.
[0164] One of the above technical solutions has the following advantages or beneficial effects: The embodiment of the present application is based on direct preference optimization of the data set, and the extraction results of the large language model are evaluated and fine-tuned through high-quality extraction results and low-quality extraction results. This helps to reduce the hallucination phenomenon generated by the model and improve the accuracy and reliability of the extraction results.
[0165] The following is a detailed description of the embodiments of the present invention with reference to specific application examples:
[0166] The embodiment of the present application is applied to the field of artificial intelligence technology and can be applied to the scene of energy decision generation. With the increase of renewable energy projects, the number of site selection regulations issued by local governments has also increased dramatically, and local regulations are characterized by diversity and complexity. Therefore, it is necessary to build a general energy decision system with high accuracy and flexibility. The energy decision generation system based on the large language model provided by the embodiment of the present application can receive the energy decision corresponding to the task input by the user, thereby improving the generation efficiency of energy decisions. Specifically, the large language model is fine-tuned by the energy data set, so that the large language model can generate high-quality tasks from various texts to meet the needs of different application scenarios, and obtain the initial energy decision large model. Then, the prompt words are adjusted by the reinforcement learning algorithm so that the generated task prompt words can enable the large model to extract specific rule information from the data set. The generated task prompt words can enable the model to identify and integrate the relationship between multiple factors and support more complex decision tasks. The transparent decision-making process can clearly show the basis and reasoning process of the decision, enhance the transparency and credibility of the decision-making system, and improve the interpretability of the decision. The initial energy decision large model can dynamically adjust the logical rules according to new data and information, so that the decision system can adapt to the changing environment. Finally, by introducing the reward model, the large language model has stronger logical judgment ability, while enhancing the anti-interference ability and generalization ability, and obtaining the final target energy decision-making large model.
[0167] See also Figure 4 The embodiment of the present application also provides an energy decision-making system based on a large language model, which can implement the above-mentioned energy decision-making method based on a large language model. The system includes:
[0168] The first module 401 is used to obtain an energy data set to fine-tune the large language model to obtain an initial energy decision large model;
[0169] The second module 402 is used to receive a task input by a user, and perform prompt word generation processing on the task input by the user according to a reinforcement learning algorithm to obtain a task prompt word;
[0170] The third module 403 is used to perform task generation processing on the initial energy decision-making model according to the task prompt words to obtain generated data;
[0171] The fourth module 404 is used to train the initial energy decision-making model according to the generated data through a reward model to obtain a target energy decision-making model;
[0172] The fifth module 405 is used to input the user input data into the target energy decision big model to perform energy decision generation processing to obtain an energy decision.
[0173] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0174] The embodiment of the present application also provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned energy decision-making method based on a large language model when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.
[0175] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0176] See also Figure 5 , Figure 5 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:
[0177] The processor 501 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0178] The memory 502 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 502 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 502, and the processor 501 calls and executes the energy decision generation method based on a large language model in the embodiment of this application;
[0179] Input / output interface 503, used to implement information input and output;
[0180] Communication interface 504, used to realize communication interaction between the device and other devices, which can be realized by wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);
[0181] A bus 505 that transmits information between the various components of the device (e.g., the processor 501, the memory 502, the input / output interface 503, and the communication interface 504);
[0182] The processor 501 , the memory 502 , the input / output interface 503 and the communication interface 504 are connected to each other in communication within the device via the bus 505 .
[0183] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned energy decision-making method based on the large language model.
[0184] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0185] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0186] The embodiment of the present application provides a method and system for generating energy decisions based on a large language model. The scheme fine-tunes the large language model by acquiring an energy data set to obtain an initial energy decision large model, and can enhance the large language model's ability to understand the logical relationship of energy industry content through the energy data set. In addition, the scheme receives user input tasks, generates prompt words for the user input tasks according to a reinforcement learning algorithm, and obtains task prompt words. The task prompt words can be used to extract feature information in the data set, so that the large language model can dynamically adjust the logical rules according to new data and information, and can adapt to the ever-changing environment. In addition, the scheme uses a reward model to train the initial energy decision large model according to the generated data to obtain a target energy decision large model, which can enable the model to have stronger logical judgment capabilities, while enhancing anti-interference capabilities and generalization capabilities, and improving the efficiency of energy decision generation.
[0187] The embodiment of this application solves the adaptability problem of large language models in professional fields by fine-tuning the large language model with energy data sets. This technology can automatically convert unlabeled text into training data sets for specific tasks, enabling the model to better understand and process the content in the professional field of energy regulations, significantly improving the performance of the model in specific fields.
[0188] The embodiment of the present application also reduces the number of calls to the large language model by optimizing the task generation and processing logic. This not only speeds up the processing speed, but also improves the overall efficiency of the system, especially when processing large-scale document sets.
[0189] At the same time, by optimizing text segmentation and processing logic, the embodiment of the present application can fully utilize the ability of a large language model to process long texts. This avoids information loss caused by text segmentation and ensures the accuracy and completeness of the extraction results. The embodiment of the present application enhances the model's ability to understand context by introducing specially designed prompt words. This enables the model to more accurately capture and process key information in complex texts, improving the accuracy of information extraction.
[0190] In addition, the embodiment of the present application realizes the automation of the logic process through the carefully designed task prompt words, which not only improves the accuracy of information extraction, but also makes the implementation of the system more flexible and can adapt to different application scenarios and needs.
[0191] The embodiment of the present application makes the system implementation more flexible, easier to maintain and expand by not completely relying on a specific programming framework for the writing of prompt words, which provides convenience for future technical upgrades and functional expansions.
[0192] Furthermore, the embodiment of the present application can effectively reduce the probability of "hallucination" in the large language model during feature extraction by introducing the reward model method, which further improves the reliability and accuracy of the data and ensures the authenticity and credibility of the extraction results.
[0193] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0194] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0195] The system embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.
[0196] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0197] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0198] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0199] In the several embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or units, which can be electrical, mechanical or other forms.
[0200] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0201] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0202] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.
[0203] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.
Claims
1. A method for generating energy decisions based on a large language model, characterized in that: The method comprises the following steps: Obtain energy data sets to fine-tune the large language model and obtain the initial energy decision-making large model; Receiving a task input by a user, and performing prompt word generation processing on the task input by the user according to a reinforcement learning algorithm to obtain a task prompt word; Performing task generation processing on the initial energy decision-making model according to the task prompt words to obtain generated data; Training the initial energy decision-making model according to the generated data through a reward model to obtain a target energy decision-making model; The user input data is input into the target energy decision model to perform energy decision generation processing to obtain an energy decision.
2. The method according to claim 1, characterized in that The step of obtaining the energy data set to fine-tune the large language model to obtain an initial energy decision large model includes the following steps: A task data set is constructed based on the energy data set; the task data set includes multiple tasks, each task includes a question text and an answer text; Inputting the question text into the large language model for question-answering processing to obtain a question-answering result; Perform cross entropy loss function calculation on the question and answer result according to the answer text to obtain a loss value; The parameters of the large language model are updated according to the loss value to obtain the initial energy decision large model.
3. The method according to claim 2, characterized in that The step of constructing a task dataset based on the energy dataset comprises the following steps: Performing keyword recognition and context screening processing on the energy data set to obtain a corpus data set; According to the data generation prompt words, the corpus data set is input into the large language model for data generation processing to obtain the task data set; the task data set includes yes / no question and answer data, extractive question and answer data set and natural language inference data set.
4. The method according to claim 1, characterized in that: The receiving of a task input by a user and performing prompt word generation processing on the task input by the user according to a reinforcement learning algorithm to obtain a task prompt word includes the following steps: Performing prompt word matching processing on the user input task through an intelligent agent to obtain a matching result; When the matching result is empty, performing prompt word generation processing on the user input task through the prompt word large model to obtain a first prompt word; When the matching result is the current prompt word, adjusting the current prompt word according to the depth search algorithm or the breadth search algorithm to obtain a second prompt word; Updating the first prompt word or the second prompt word to the current prompt word; Inputting the user input task and the current prompt word into the initial energy decision-making model for generation processing to obtain prediction data; A predicted reward score is calculated based on the predicted data, and a strategy update process is performed on the agent based on the predicted reward score, and the process returns to the step of matching the prompt word of the user input task through the agent until a preset condition is met, and the current prompt word is determined as the task prompt word.
5. The method according to claim 4, characterized in that The step of performing prompt word generation processing on the user input task by using the prompt word big model to obtain the first prompt word comprises the following steps: Performing content recognition processing on the user input task through the prompt word large model to obtain recognized content; Performing rule extraction and sorting processing on the identified content to obtain sorted content; The first prompt word is obtained according to the sorted content definition, and the first prompt word includes role, background, skill, goal, constraint and output format.
6. The method according to claim 1, characterized in that The step of training the initial energy decision-making model according to the generated data through the reward model to obtain the target energy decision-making model includes the following steps: Constructing a standard data set to perform initialization training processing on the initial energy decision-making large model to obtain an initialization training model; According to the reward model, the generated data is input into the initialization training model for reinforcement learning processing to obtain the target energy decision model.
7. The method according to claim 6, characterized in that The step of constructing a standard data set to perform initialization training processing on the initial energy decision-making large model to obtain an initialization training model includes the following steps: The standard data set is constructed according to the positive sample data and the negative sample data; The initial energy decision-making large model is directly subjected to preference optimization processing according to the standard data set to obtain the initialization training model.
8. An energy decision-making system based on a large language model, characterized in that: The system comprises: The first module is used to obtain energy data sets to fine-tune the large language model and obtain the initial energy decision-making large model; The second module is used to receive a task input by a user, and perform prompt word generation processing on the task input by the user according to a reinforcement learning algorithm to obtain a task prompt word; The third module is used to perform task generation processing on the initial energy decision-making model according to the task prompt words to obtain generated data; The fourth module is used to train the initial energy decision-making model according to the generated data through a reward model to obtain a target energy decision-making model; The fifth module is used to input the user input data into the target energy decision model to perform energy decision generation processing to obtain an energy decision.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Satellite energy system intelligent control method and storage medium
CN120335379A
Model cooperative processing method and device, electronic equipment and computer storage medium
CN120671795A
Modelica model data set construction method based on execution feedback
CN120781074A
Method and system for extracting ecological environment access list rule of intelligent agent
CN121615756A