A large model adaptation method, device, equipment and storage medium

By fine-tuning the parameters of the large language model and using reinforcement learning, and by optimizing the model using historical power grid operation command data, the problems of understanding and execution errors of the large model in the main power grid operation tasks were solved, thereby improving the accuracy and efficiency of the tasks.

CN119558377BActive Publication Date: 2025-12-16GUANGDONG POWER GRID CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411610779.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-12-16
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

Large language models are prone to misunderstanding or execution errors in power grid main network operation tasks, resulting in low accuracy of scheduling task execution.

Method used

By fine-tuning the parameters of the pre-trained large model based on historical mainnet operation command data, a fine-tuned mainnet large model is obtained. Then, by using reinforcement learning to optimize the model with feedback data, a large model for mainnet operation tasks is formed.

Benefits of technology

This improved the large model's ability to understand power grid commands, making its output more in line with actual business needs and enhancing the accuracy and efficiency of power grid main grid scheduling tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119558377B_ABST
    Figure CN119558377B_ABST
Patent Text Reader

Abstract

The application discloses a large model adaptation method and device, equipment and a storage medium. The large model adaptation method comprises the following steps: based on historical main network operation instruction data, fine-tuning parameters of a pre-training large model to obtain a fine-tuned main network large model; using the fine-tuned main network large model to execute a main network operation task to obtain an output result of the fine-tuned main network large model and feedback data based on the output result; based on the output result and the feedback data, reinforcement learning is performed on the fine-tuned main network large model to obtain a main network operation task large model. The technical scheme of the embodiment of the application can improve the understanding ability of the large model to the power grid instruction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of machine learning, and in particular to a large model adaptation method and device, equipment and a storage medium. BACKGROUND

[0002] Large language model (LLM for short, hereinafter referred to as "large model") refers to a deep learning model with a large number of parameters and complex structure used in the field of natural language processing and machine learning. Large models are widely used in various fields.

[0003] Large models are good at text processing, but they are prone to bias when directly solving specific tasks. For example, large models may have understanding or execution errors during the execution of power grid main network operation tasks. Therefore, how to improve the understanding and execution of large language models for main network operation task instructions is very important for improving the accuracy and efficiency of power grid main network dispatching task execution. SUMMARY

[0004] The present application provides a large model adaptation method, device, equipment and storage medium to solve the problem of low accuracy of large model in power grid main network dispatching task execution.

[0005] According to one aspect of the present application, a large model adaptation method is provided, comprising:

[0006] Based on historical main network operation instruction data, the pre-trained large model is fine-tuned to obtain a fine-tuned main network large model;

[0007] The fine-tuned main network large model is used to execute the main network operation task to obtain the output result of the fine-tuned main network large model and feedback data based on the output result;

[0008] Based on the output result and the feedback data, the fine-tuned main network large model is subjected to reinforcement learning to obtain a main network operation task large model.

[0009] According to another aspect of the present application, a large model adaptation device is provided, comprising:

[0010] The parameter fine-tuning module is configured to fine-tune the pre-trained large model based on historical main network operation instruction data to obtain a fine-tuned main network large model;

[0011] The feedback acquisition module is configured to use the fine-tuned main network large model to execute the main network operation task to obtain the output result of the fine-tuned main network large model and feedback data based on the output result;

[0012] The reinforcement learning module is configured to perform reinforcement learning on the fine-tuned main network large model based on the output result and the feedback data, to obtain a main network operation task large model.

[0013] According to another aspect of the present application, an electronic device is provided, comprising:

[0014] at least one processor; and

[0015] a memory in communication with the at least one processor; wherein

[0016] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the large model adaptation method according to any one of the embodiments of the present application.

[0017] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to perform the large model adaptation method according to any one of the embodiments of the present application.

[0018] According to another aspect of the present application, a computer program product is provided, comprising a computer program for enabling a processor to perform the large model adaptation method according to any one of the embodiments of the present application.

[0019] The technical solution of the embodiments of the present application performs parameter fine-tuning on the pre-trained large model based on historical main network operation instruction data to obtain a fine-tuned main network large model, and then uses the fine-tuned main network large model to perform a main network operation task to obtain an output result of the fine-tuned main network large model and feedback data based on the output result. Finally, reinforcement learning is performed on the fine-tuned main network large model based on the output result and the feedback data to obtain a main network operation task large model. By performing parameter fine-tuning on the pre-trained large model based on historical main network operation instruction data, the understanding ability of the large model for power grid instructions can be improved, and further reinforcement learning based on the feedback data can make the output of the large model more in line with actual business requirements.

[0020] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description only show some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0022] Figure 1 is a flow chart of a large model adaptation method according to an embodiment of the present application;

[0023] Figure 2 is a flow chart of a large model adaptation method according to an embodiment of the present application;

[0024] Figure 3 is a structural schematic diagram of a large model adaptation device according to an embodiment of the present application;

[0025] Figure 4 is a structural schematic diagram of an electronic device for implementing a large model adaptation method according to an embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to make the person skilled in the art better understand the present application, the following will combine the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily mean a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0028] Embodiment one

[0029] Figure 1A flowchart of a large model adaptation method is provided for the first embodiment of the application. The embodiment can be applied to the parameter optimization of a large model using grid instruction data and user feedback data. The method can be executed by a large model adaptation device, which can be implemented in hardware and / or software. The large model adaptation device can be configured in various general-purpose computing devices. As shown in FIG. 1, the method comprises the following steps. Figure 1

[0030] In S110, the pre-trained large model is fine-tuned based on historical main grid operation instruction data to obtain a fine-tuned main grid large model.

[0031] The historical main grid operation instruction data is used to fine-tune the pre-trained large model. The historical main grid operation instruction data is instruction-type data obtained from the grid main grid dispatching system. To improve the accuracy of the large model in understanding the grid main grid operation instruction, the historical main grid operation instruction data needs to cover common operation scenarios in the grid main grid, such as line switching, device start-stop or load adjustment, etc. For example, the historical main grid operation instruction data is text instruction, operation command or device control instruction, etc. The text instruction can be "close line 123" or "start generator set A". The operation command is a specific system instruction, such as "close_line(123)". The device control instruction is, for example, "open_circuit_breaker(CB12)".

[0032] The pre-trained large model is a large model that has been preliminarily trained through a large-scale data set. The pre-trained large model learns general features and knowledge in the pre-training process, and does not learn specific knowledge based on the operation instruction of the grid main grid. This makes the pre-trained large model prone to understanding errors and execution deviations when processing grid main grid operation tasks.

[0033] In the embodiment of the application, to enable the large model to more accurately understand and execute the grid main grid operation task, historical main grid operation instruction data can be collected, and the pre-trained large model can be further fine-tuned using the historical main grid operation instruction data to obtain a fine-tuned main grid large model. Specifically, the historical main grid operation instruction data can be input into the pre-trained large model to obtain the output result of the pre-trained large model. Then, based on the deviation between the output result and the actual operation corresponding to the historical main grid operation instruction, a loss function is formed. Finally, the loss function is optimized through back propagation to fine-tune the parameters of the pre-trained large model, so that the performance of the pre-trained large model in understanding and generating grid main grid operation task instructions is improved.

[0034] The loss function can include the deviation between the model output result and the actual task, and can also include the deviation between the model output result and the actual task, as well as a grid safety degree penalty term of the model output result.​

[0035] In addition, during the parameter fine-tuning process of the pre-trained large model, only a certain number of model parameters close to the input layer of the model can be fine-tuned. For example, only the layers in the model used for feature extraction are fine-tuned, and the model parameters of the remaining layers are frozen, which can improve the understanding ability of the model for the power grid main network operation task, reduce the calculation amount in the model fine-tuning process, and improve the parameter fine-tuning efficiency.

[0036] S120, using the fine-tuned main network large model to perform the main network operation task to obtain an output result of the fine-tuned main network large model and feedback data based on the output result.

[0037] In the embodiment of the application, after the pre-trained large model is fine-tuned to obtain the fine-tuned main network large model, the fine-tuned main network large model is used to perform the main network operation task, and the output result of the fine-tuned main network large model is obtained. In addition, in the power grid main network scheduling operation based on the output result, task execution feedback data is collected, for example, the feedback data includes task completion rate, abnormal practice record, and operator feedback, etc.

[0038] In order to improve the model training efficiency, the feedback data can be further cleaned to remove noise and invalid information, and the remaining valid feedback data and the corresponding generated task are labeled to form a training data set for reinforcement learning.

[0039] S130, based on the output result and the feedback data, reinforcement learning is performed on the fine-tuned main network large model to obtain a main network operation task large model.

[0040] In the embodiment of the application, after the feedback data is obtained, based on the output result of the fine-tuned main network large model and the feedback data for the output result, the fine-tuned main network large model is further subjected to reinforcement learning to obtain a main network operation task large model. Specifically, based on the feedback data, a reinforcement learning key is constructed, the output result of the fine-tuned main network large model is taken as an action, and the feedback data is taken as a reward to construct a reward function. Based on the reward function, a deep Q network (Deep Q-Network, DQN for short) is used for iterative training to adjust the task generation strategy of the model, so that the model output is more in line with the actual business needs of the power grid main network scheduling, and finally a main network operation task large model is obtained.

[0041] The reward function can include a task completion reward and a task execution safety degree reward. Based on reinforcement learning, the model parameters are gradually adjusted so that the model output is more in line with the actual operation demand, and the generation strategy of the model is optimized to avoid high-risk operation task generation.

[0042] In addition, after the master grid operation task model is obtained, the effect of parameter fine-tuning can be evaluated through indicators such as success rate of task execution and safety score, so that it is ensured that the model can generate efficient and safe operation tasks in various grid master dispatching scenarios.

[0043] The technical scheme of the embodiment of the application is based on historical master grid operation instruction data, and parameters of a pre-trained large model are fine-tuned to obtain a fine-tuned master grid large model. Then, the fine-tuned master grid large model is used to execute a master grid operation task, and output results of the fine-tuned master grid large model and feedback data based on the output results are obtained. Finally, the fine-tuned master grid large model is subjected to reinforcement learning based on the output results and the feedback data, and a master grid operation task large model is obtained. The pre-trained large model is fine-tuned through historical master grid operation instruction data, which can improve the understanding ability of the large model for grid instructions. Furthermore, the large model can output results that are more in line with actual business requirements through further reinforcement learning based on feedback data.

[0044] Embodiment two

[0045] Figure 2 A flowchart of a large model adaptation method provided by the second embodiment of the application is further refined on the basis of the above-mentioned embodiment, and specific steps of fine-tuning a pre-trained large model based on historical master grid operation instruction data to obtain a fine-tuned master grid large model, and specific steps of reinforcing learning of the fine-tuned master grid large model based on output results and feedback data to obtain a master grid operation task large model are provided. As shown in Figure 2 The method comprises the following steps:

[0046] In S210, historical master grid operation instruction data is preprocessed to obtain an instruction text sequence.

[0047] In the embodiment of the application, in order to improve the model fine-tuning efficiency and enable the model to correctly understand and execute grid instructions, the historical master grid operation instruction data is first preprocessed to obtain an instruction text sequence. Specifically, historical master grid operation instruction data is obtained from a grid master dispatching system, for example, historical master grid operation instructions can be text instructions such as “close line 123” and “start generator set A”, or operation commands such as “close_line(123)”, or device control instructions such as “open_circuit_breaker(CB12)”.

[0048] Further, the historical master grid operation instruction data in different formats is converted into a standard text sequence format, for example, [CLS] close line 123 [SEP], so as to facilitate processing by the large model.

[0049] After being converted into the text sequence format, each instruction in the text sequence can be further data cleaned to remove invalid content and normalized in quality into a unified format, with all instructions using the same operator and parameter representation, such as "[CLS] operation content operation object type operation object identification [SEP]". Finally, the standardized instruction text sequence is obtained, such as "[CLS] start the generator set A [SEP]".

[0050] S220, input the instruction text sequence into the pre-trained large model to obtain an output result of the pre-trained large model.

[0051] S230, form a loss function based on a deviation of the output result and the instruction text sequence in association with an actual task and a power grid operation safety degree of the output result.

[0052] In the embodiment of the application, after the instruction text sequence is obtained, the instruction text sequence is input into the pre-trained large model to obtain an output result of the pre-trained large model. A loss function is formed based on a deviation of the output result and the quality text sequence in association with an actual task and a power grid operation safety degree of the output result.

[0053] Optionally, the loss function is specifically as follows:

[0054] L(θ)=α×L task +β×L safety

[0055] Wherein, L task is the deviation of the output result of the pre-trained large model and the instruction text sequence in association with the actual task, L safety is the power grid operation safety degree penalty term of the output result of the pre-trained large model, and α and β are weight coefficients.

[0056] In the optional embodiment, the specific formula of the loss function is provided as follows:

[0057] L(θ)=α×L task +β×L safety

[0058] Wherein, L task is the deviation of the output result of the pre-trained large model and the instruction text sequence in association with the actual task, L safety is the power grid operation safety degree penalty term of the output result of the pre-trained large model, and α and β are weight coefficients.

[0059] By simultaneously considering the output deviation and the power grid operation safety degree in the loss function, the accuracy of the model in understanding the power grid instruction can be improved, and the safety of the model output can be improved.

[0060] S240, based on the loss function, the pre-training large model is parameter fine-tuned to obtain a fine-tuned main grid large model.

[0061] In the embodiment of the application, after the loss function is constructed, the pre-training large model is parameter fine-tuned through the instruction text sequence, the loss function is optimized through back propagation, the model parameters are fine-tuned, and the performance of the model in understanding and generating power grid main grid operation task instructions is improved.

[0062] Specifically, to improve the parameter adjustment efficiency, part of the levels in the pre-training large model can be fine-tuned according to actual needs, for example, only the parameters of the convolutional layer are adjusted, so that the model can better learn the characteristics of the power grid main grid operation task and improve the adaptability of the model in the power grid field.

[0063] Optionally, based on the loss function, the pre-training large model is parameter fine-tuned to obtain a fine-tuned main grid large model, comprising:

[0064] The model parameters of a set number of layers from the output layer in the pre-training large model are frozen;

[0065] Based on the loss function, the other model parameters in the pre-training large model except the frozen model parameters are fine-tuned to obtain a fine-tuned main grid large model.

[0066] In the optional embodiment, a specific way of fine-tuning the pre-training large model based on the loss function to obtain a fine-tuned main grid large model is provided: the model parameters of a set number of layers from the output layer in the pre-training large model are frozen, for example, the model parameters of a set number of layers from the output layer to the input layer are frozen. Then, based on the loss function, the other model parameters in the pre-training large model except the frozen model parameters are fine-tuned to obtain a fine-tuned main grid large model. Through the above method, only the model parameters close to the input layer in the model can be fine-tuned, the efficiency of model parameter fine-tuning is improved, and since the model parameters close to the output layer are used for feature extraction, fine-tuning only this part of the model parameters can improve the adaptability of the model to the power grid main grid operation task.

[0067] Optionally, after the pre-training large model is parameter fine-tuned to obtain a fine-tuned main grid large model, it further comprises:

[0068] The verification main grid operation instruction data is input into the fine-tuned main grid large model to obtain an output result of the fine-tuned main grid large model, and the parameter index of the fine-tuned main grid large model is determined according to the output result;

[0069] In the case that the parameter index does not meet the training requirements, the operation of fine-tuning the pre-training large model based on the historical main grid operation instruction data to obtain a fine-tuned main grid large model is returned until the parameter index meets the training requirements.

[0070] In this optional embodiment, specific steps after parameter fine-tuning of the pre-trained large model to obtain the fine-tuned main network large model are provided: after collecting the historical main network operation instruction data, the data can be divided into two parts, one part is used as the training set for model parameter fine-tuning, and the other part is used as the validation main network operation instruction data for model parameter index verification.

[0071] The validation main network operation instruction data is input into the fine-tuned main network large model to obtain the output result of the fine-tuned main network large model, and the parameter index of the fine-tuned main network large model is determined according to the output result. For example, the parameter index is the model performance that can be quantified, such as precision and recall rate. In the case that the parameter index does not meet the training requirement, the operation of fine-tuning the pre-trained large model based on the historical main network operation instruction data to obtain the fine-tuned main network large model is returned until the parameter index meets the training requirement. In the iteration process, all parameters in the current fine-tuned main network large model can also be fine-tuned, or only part of the model parameters that need to be optimized can be fine-tuned, for example, only the convolution layer parameters are fine-tuned, so that the model can better learn the specific features of the input data.

[0072] S250, using the fine-tuned main network large model to perform the main network operation task to obtain the output result of the fine-tuned main network large model and the feedback data based on the output result.

[0073] S260, taking the output result as an action, and determining the task reward based on the feedback data and the reward function; the reward function includes a task completion reward and a task execution safety degree reward.

[0074] In the embodiment of the application, the output result is taken as an action, and the task reward is determined based on the feedback data and the reward function. The reward function includes a task completion reward and a task execution safety degree reward. Considering the task completion and the task execution safety degree, the model can complete the task on the basis of ensuring the safe operation of the power grid.

[0075] Optionally, the reward function is as follows:

[0076] R=γ×S task +(1-γ)×S safety

[0077] Wherein, S task is the task completion reward of the fine-tuned main network large model, S safety is the task execution safety degree reward, and γ is a weight factor.

[0078] In this optional embodiment, the specific formula of the reward function is provided:

[0079] R=γ×S task +(1-γ)×Ssafety

[0080] wherein S task is the task completion reward of fine-tuning the main grid large model, S safety is the task execution safety reward, and γ is a weight factor.

[0081] By considering the task completion reward and the task execution safety reward in the reward function, the model can complete the task on the basis of ensuring the safe operation of the power grid.

[0082] S270, based on the action and task reward, the fine-tuning main grid large model is reinforced learning, and the main grid operation task large model is obtained.

[0083] In the embodiment of the application, the output result is taken as the action, based on the action and task reward, the DQN method is used to reinforce the learning of the fine-tuning main grid large model, and the main grid operation task large model is obtained.

[0084] Optionally, on the basis of reinforcement learning, the task generation strategy of the model is adjusted through multiple iterations of training, so that the output is more in line with the actual business needs of the main grid dispatching of the power grid.

[0085] Optionally, the running state of the main grid of the power grid can be monitored in real time, and the changed scene can be identified, including the device state, the line topology change and the load change. According to the identified scene change, the fine-tuning strategy is dynamically adjusted. For example, when a device fault or an emergency power outage is detected, the system will preferentially call the strategy optimized by reinforcement learning to ensure the safety of task generation and the success rate of execution.

[0086] The closed-loop analysis of the adapted task generation result and the actual execution feedback is carried out, and the model parameters are further adjusted and optimized to form a continuous improvement closed-loop feedback mechanism. Through multiple cycles, the model adaptation performance is continuously enhanced.

[0087] Dynamic adaptation effect evaluation: by monitoring the task completion rate, operation risk and other indicators in real time, the effect of dynamic adaptation is evaluated to ensure that the model can maintain efficient task generation capability in various complex scenarios.

[0088] The technical scheme of the embodiment of the present application pre-processes historical main network operation instruction data to obtain an instruction text sequence, inputs the instruction text sequence into a pre-trained large model to obtain an output result of the pre-trained large model, forms a loss function based on the output result and the instruction text sequence, and the deviation of the actual task associated with the output result and the power grid operation safety degree of the output result, performs parameter fine-tuning on the pre-trained large model based on the loss function to obtain a fine-tuned main network large model, further, executes a main network operation task by using the fine-tuned main network large model to obtain an output result of the fine-tuned main network large model and feedback data based on the output result, takes the output result as an action, and determines a task reward based on the feedback data and a reward function, and performs reinforcement learning on the fine-tuned main network large model based on the action and the task reward to obtain a main network operation task large model, which adjusts the model by simultaneously considering the accuracy and safety of the model in executing the task, and improves the adaptability of the large model to the main network operation task of the power grid.

[0089] Embodiment three

[0090] Figure 3 A structural schematic diagram of a large model adaptation device provided for the third embodiment of the present application is shown in FIG. 3. Figure 3 As shown in the figure, the device comprises:

[0091] The parameter fine-tuning module 310 is configured to perform parameter fine-tuning on the pre-trained large model based on the historical main network operation instruction data to obtain a fine-tuned main network large model.

[0092] The feedback acquisition module 320 is configured to execute a main network operation task by using the fine-tuned main network large model to obtain an output result of the fine-tuned main network large model and feedback data based on the output result.

[0093] The reinforcement learning module 330 is configured to perform reinforcement learning on the fine-tuned main network large model based on the output result and the feedback data to obtain a main network operation task large model.

[0094] The technical scheme of the embodiment of the present application performs parameter fine-tuning on the pre-trained large model based on the historical main network operation instruction data to obtain a fine-tuned main network large model, further executes a main network operation task by using the fine-tuned main network large model to obtain an output result of the fine-tuned main network large model and feedback data based on the output result, and finally performs reinforcement learning on the fine-tuned main network large model based on the output result and the feedback data to obtain a main network operation task large model, which can improve the understanding ability of the large model to the power grid instruction by performing parameter fine-tuning on the pre-trained large model based on the historical main network operation instruction data, and can make the output of the large model more in line with the actual business requirements by further reinforcement learning based on the feedback data.

[0095] Optionally, the parameter fine-tuning module 310 comprises:

[0096] A data preprocessing unit is configured to preprocess historical main grid operation instruction data to obtain an instruction text sequence.

[0097] A model output result acquisition unit is configured to input the instruction text sequence into the pre-trained large model to obtain an output result of the pre-trained large model.

[0098] A loss function construction unit is configured to form a loss function based on a deviation of the output result and the instruction text sequence associated with an actual task and a power grid operation safety degree of the output result.

[0099] A parameter fine-tuning unit is configured to fine-tune parameters of the pre-trained large model based on the loss function to obtain a fine-tuned main grid large model.

[0100] Optionally, the loss function is specifically as follows:

[0101] L(θ)=α×L task +β×L safety

[0102] wherein L task is a deviation of the output result of the pre-trained large model and the instruction text sequence associated with the actual task, L safety is a power grid operation safety degree penalty term of the output result of the pre-trained large model, and α and β are weight coefficients.

[0103] Optionally, the reinforcement learning module 330 is specifically configured to:

[0104] take the output result as an action, and determine a task reward based on the feedback data and a reward function; the reward function includes a task completion reward and a task execution safety degree reward;

[0105] perform reinforcement learning on the fine-tuned main grid large model based on the action and the task reward to obtain a main grid operation task large model.

[0106] Optionally, the reward function is specifically as follows:

[0107] R=γ×S task +(1-γ)×S safety

[0108] wherein S task is the task completion reward of the fine-tuned main grid large model, S safety is the task execution safety degree reward, and γ is a weight factor.

[0109] Optionally, the parameter fine-tuning unit is specifically configured to:

[0110] freeze model parameters of a set number of layers starting from an output layer in the pre-trained large model;

[0111] Based on the loss function, the other model parameters in the pre-training large model except the frozen model parameters are fine-tuned to obtain a fine-tuned main network large model.

[0112] Optionally, the large model adaptation apparatus further comprises:

[0113] The parameter index acquisition module is configured to, after fine-tuning the pre-training large model to obtain the fine-tuned main network large model, input the verification main network operation instruction data into the fine-tuned main network large model to obtain an output result of the fine-tuned main network large model, and determine a parameter index of the fine-tuned main network large model according to the output result.

[0114] The fine-tuned main network large model iteration module is configured to, in a case where the parameter index does not meet the training requirement, return to perform the operation of fine-tuning the pre-training large model based on the historical main network operation instruction data to obtain the fine-tuned main network large model until the parameter index meets the training requirement.

[0115] The large model adaptation apparatus provided in the embodiments of the present application can perform the large model adaptation method provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of performing the method.

[0116] In the technical solution of the present application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solution comply with relevant laws and regulations and do not violate public order and good customs.

[0117] Embodiment four

[0118] According to the embodiments of the present application, the present application further provides an electronic device, a readable storage medium and a computer program product.

[0119] Figure 4 A structural schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, appliances, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.

[0120] As Figure 4As shown, the electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., communicatively connected to the at least one processor 11, where the memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 12 or loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0121] Various components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc., an output unit 17, such as various types of displays, a speaker, etc., a storage unit 18, such as a magnetic disk, an optical disk, etc., and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0122] The processor 11 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the large model adaptation method.

[0123] In some embodiments, the large model adaptation method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the large model adaptation method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the large model adaptation method by any other appropriate means, such as by means of firmware.

[0124] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0125] Computer programs used to implement the processes of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program

[0126] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store computer programs for use by or in connection with an instruction execution system, apparatus, or device. Computer-readable storage media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0127] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0128] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data application server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0129] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing application system, and solves the defects of large management difficulty and weak business scalability in traditional physical host and VPS application.

[0130] It should be understood that various forms of flow shown above can be used, with steps reordered, added, or removed. For example, the steps recited in the present invention can be performed in parallel, in series, or in a different order, without limitation herein, as long as the desired results of the technical solutions of the present invention can be achieved.

[0131] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and principles of the present invention. Accordingly, the disclosure is intended to embrace all such alternatives, modifications and variances that fall within the scope of the present invention, including the patent claims as allowed.

Claims

1. A method for adapting large models, characterized in that, include: Based on historical mainnet operation command data, the parameters of the pre-trained large model are fine-tuned to obtain the fine-tuned mainnet large model. The mainnet operation task is executed using the fine-tuned mainnet large model to obtain the output results of the fine-tuned mainnet large model and the feedback data based on the output results; Based on the output results and the feedback data, reinforcement learning is performed on the fine-tuned mainnet large model to obtain the mainnet operation task large model. Based on historical mainnet operation command data, the parameters of the pre-trained large model are fine-tuned to obtain the fine-tuned mainnet large model, including: The historical mainnet operation command data is preprocessed to obtain a command text sequence; The instruction text sequence is input into the pre-trained large model to obtain the output result of the pre-trained large model; Based on the deviation between the output results and the actual task associated with the instruction text sequence, and the power grid operation safety of the output results, a loss function is formed; Based on the loss function, the parameters of the pre-trained large model are fine-tuned to obtain the fine-tuned mainnet large model; The loss function is as follows: ; in, It refers to the discrepancy between the output of the pre-trained large model and the actual task in relation to the instruction text sequence. It is the power grid operation safety penalty term in the output of the pre-trained large model. and These are weighting coefficients; Based on the output results and the feedback data, reinforcement learning is performed on the fine-tuned mainnet large model to obtain a large model for mainnet operation tasks, including: The output is used as an action, and the task reward is determined based on the feedback data and the reward function; the reward function includes a task completion reward and a task execution safety reward. Based on the action and task rewards, reinforcement learning is performed on the fine-tuned mainnet model to obtain the mainnet operation task model. The reward function is as follows: ; in, It's a reward for completing the task of fine-tuning the mainnet's large model. It is a reward for the safety of task execution. It is a weighting factor.

2. The method according to claim 1, characterized in that, Based on the loss function, the parameters of the pre-trained large model are fine-tuned to obtain the fine-tuned main network large model, including: The model parameters of the pre-trained large model, starting from the output layer, are frozen for a set number of layers. Based on the loss function, the model parameters other than the frozen model parameters in the pre-trained large model are fine-tuned to obtain the fine-tuned mainnet large model.

3. The method according to claim 1, characterized in that, After fine-tuning the parameters of the pre-trained large model to obtain the fine-tuned mainnet large model, the following steps are also included: The mainnet operation command data is input into the fine-tuned mainnet large model to obtain the output result of the fine-tuned mainnet large model, and the parameter index of the fine-tuned mainnet large model is determined based on the output result; If the parameter indicators do not meet the training requirements, return to the execution based on historical mainnet operation instruction data to fine-tune the parameters of the pre-trained large model, and obtain the operation of fine-tuning the mainnet large model until the parameter indicators meet the training requirements.

4. A large model adaptation device, characterized in that, include: The parameter fine-tuning module is used to fine-tune the parameters of the pre-trained large model based on historical mainnet operation command data, so as to obtain the fine-tuned mainnet large model. The feedback acquisition module is used to perform mainnet operation tasks using the fine-tuned mainnet large model, and obtain the output results of the fine-tuned mainnet large model and feedback data based on the output results; The reinforcement learning module is used to perform reinforcement learning on the fine-tuned mainnet large model based on the output results and the feedback data to obtain the mainnet operation task large model. The parameter fine-tuning module includes: The data preprocessing unit is used to preprocess historical mainnet operation command data to obtain command text sequences; The model output result acquisition unit is used to input the instruction text sequence into the pre-trained large model and obtain the output result of the pre-trained large model. The loss function construction unit is used to form a loss function based on the deviation between the output result and the instruction text sequence in relation to the actual task, and the power grid operation safety of the output result; The parameter fine-tuning unit is used to fine-tune the parameters of the pre-trained large model based on the loss function to obtain the fine-tuned main network large model; The loss function is as follows: ; in, It refers to the discrepancy between the output of the pre-trained large model and the actual task in relation to the instruction text sequence. It is the power grid operation safety penalty term in the output of the pre-trained large model. and These are weighting coefficients; The reinforcement learning module is specifically used for: The output is used as an action, and the task reward is determined based on the feedback data and the reward function; the reward function includes a task completion reward and a task execution safety reward. Based on the action and task rewards, reinforcement learning is performed on the fine-tuned mainnet model to obtain the mainnet operation task model. The reward function is as follows: ; in, It's a reward for completing the task of fine-tuning the mainnet's large model. It is a reward for the safety of task execution. It is a weighting factor.

5. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the large model adaptation method according to any one of claims 1-3.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the large model adaptation method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Traditional Chinese medicine large model preference alignment method, device and medium

    CN118155860A

  • Traditional Chinese medicine large model based on reinforcement learning and preference alignment method

    CN118230908A