Starting scheme optimization method and system

Through the combination of meta-learning model and reinforcement learning, a power grid startup solution suitable for multiple scenarios is generated, which solves the problem of instarting efficiency caused by complex system training in the existing technology, and achieves rapid adaptation and efficient resource utilization.

CN120454167APending Publication Date: 2025-08-08ZHONGSHAN POWER SUPPLY BUREAU OF GUANGDONG POWER GRID

Patent Information

Application Number
CN202510612632.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing power grid startup solution generation method has high requirements for system training and complex system implementation, resulting in low startup efficiency.

Method used

The meta-learning model is used for data preprocessing and global parameter training, combined with reinforcement learning, generate the optimal startup scheme, and reinforcement learning of scene data is carried out through target task data, and global initialization parameters suitable for multiple scenarios are generated.

Benefits of technology

It significantly shortens the optimization time, improves the system startup efficiency, reduces resource consumption and operating costs, and provides more comprehensive startup solution optimization support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120454167A_ABST
    Figure CN120454167A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power grid starting scheme optimization, discloses a starting scheme optimization method and system, and ensures the accuracy and integrity of data by performing data preprocessing on initial historical starting data. And performing model training on the global parameters of the initial meta-learning model based on the target historical starting data and the target task data to obtain a target meta-learning model, and performing reinforcement learning based on the scene data corresponding to the target task data through the target meta-learning model. Multiple starting scenes are trained by using a meta-learning method, global initialization parameters suitable for different scenes are generated, and rapid adaptive capacity is provided for a subsequent model, so that the system starting efficiency is improved. The technical problem that a starting scheme generation method adopted in the prior art is high in system training requirement and complex in system implementation, and consequently the system starting efficiency is low is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power grid startup scheme optimization, and in particular to a startup scheme optimization method and system. Background Art

[0002] The startup strategy used by power grids is crucial for ensuring safe and efficient operation. Its core lies in rationally formulating startup sequences and load distribution strategies to minimize startup shock and equipment losses. During startup, the insulation state and load response of high-voltage equipment are key factors affecting system stability. Experimental scenarios are needed to study the startup characteristics and transient fluctuations of different equipment, optimize startup strategies, and improve system stability and reliability.

[0003] Patent number CN117609769A currently discloses a method, system, device, and medium for generating startup plans for new power grid equipment. This method simulates the startup process of new power grid equipment by constructing a generative model based on an expert rule knowledge fusion network. By acquiring startup range data and power grid topology information, features are extracted to generate a startup plan. This method utilizes the expert rule knowledge fusion network to automatically analyze power grid topology relationships and equipment startup requirements. By integrating expert rules with a deep learning-based text generation model, the method optimizes the startup steps and ensures that the plan meets safety and compliance standards. Furthermore, dynamic text generation technology based on offline training and cross-entropy loss function optimization verifies the generated plan in real time, confirming that key steps in the startup process meet expectations. The final plan is output as text, which power grid control personnel can use directly or further adjust. However, this method places high demands on feature extraction, text generation accuracy, and system training, making the system implementation complex and resulting in low system startup efficiency.

[0004] Patent number CN117540741A discloses a system and method for intelligent startup plan generation based on deep learning. This method simulates the grid startup process to construct an intelligent startup plan generation system. Deep learning technology is used to establish a knowledge base of startup steps and conditions, and combined with the grid topology model, it analyzes the dynamic impact of new equipment commissioning on the grid structure. By combining deep learning with a graph attention neural network, this method automatically analyzes grid topology relationships and equipment startup requirements, optimizes the startup sequence, and ensures that the plan meets safety and compliance standards. Furthermore, dynamic verification technology is used to verify the generated plan in real time, confirming that key startup steps such as charging, phase verification, and load testing meet expectations. Finally, the plan is presented in an intuitive and visual format for control personnel to review and implement. This method places high demands on deep learning model training, topology feature extraction accuracy, and real-time dynamic response capabilities. Feature extraction of the system model and training of the deep learning model rely on high-quality and diverse historical data. Furthermore, high hardware computing power and algorithm optimization requirements are required, increasing the deployment cost and complexity of the system implementation, resulting in low system startup efficiency. Summary of the Invention

[0005] The present invention provides a startup scheme optimization method and system, which solves the technical problems that the startup scheme generation method adopted in the prior art has high requirements for system training, complex system implementation, and leads to low system startup efficiency.

[0006] A first aspect of the present invention provides a method for optimizing a startup scheme, comprising:

[0007] Acquire initial historical startup data and target task data, perform data preprocessing on the initial historical startup data, and generate target historical startup data;

[0008] Performing model training on global parameters of an initial meta-learning model based on the target historical startup data and the target task data to generate a target meta-learning model;

[0009] The target meta-learning model performs reinforcement learning based on the scenario data corresponding to the target task data to generate an optimal startup plan.

[0010] Optionally, the step of preprocessing the initial historical startup data to generate target historical startup data includes:

[0011] Deleting duplicate data in the initial historical startup data to generate first processed data;

[0012] Using a mean filling method to fill in missing values in the first processed data to generate second processed data;

[0013] removing noise from the second processed data to generate third processed data;

[0014] Perform feature extraction on the third processed data according to preset extraction indicators to generate a startup plan indicator;

[0015] Target historical startup data is constructed using the third processed data and the startup plan indicator.

[0016] Optionally, the step of performing model training on global parameters of the initial meta-learning model based on the target historical startup data and the target task data to generate a target meta-learning model includes:

[0017] Performing gradient updates on global parameters of the initial meta-learning model using a support set of the device in the target task data to generate an intermediate meta-learning model;

[0018] Constructing a device query set set by using a query set of each device in the target task data and a query set of each device in the target historical startup data;

[0019] Using each device query set in the device query set set to respectively calculate the loss value of the global parameter of the intermediate meta-model to generate multiple loss values;

[0020] Calculate the sum of all the loss values to generate a query set loss value;

[0021] The query set loss value and a preset learning rate are used to perform gradient update on the global parameters of the intermediate meta-learning model to generate a target meta-learning model.

[0022] Optionally, the step of using each device query set in the device query set to calculate the loss value of the global parameter of the intermediate meta-model to generate multiple loss values includes:

[0023] Substituting the startup time data and expected startup time corresponding to the device query sets in the device query set into a preset startup time loss formula respectively, and calculating an initial startup time loss value corresponding to the device query set;

[0024] Calculating a multiplication value between the initial startup time loss value and a first preset weight coefficient to generate a target startup time loss value corresponding to the device query set;

[0025] Substituting the load data and preset load distribution data in the device query set into a preset balance loss formula to calculate an initial balance loss value corresponding to the device query set;

[0026] Calculating a multiplication value between the initial instantaneous balance value and a second preset weight coefficient to generate a target balance loss value corresponding to the device query set;

[0027] Substituting the device state data corresponding to the device query set into a preset device state loss formula to calculate the initial state loss value corresponding to the device query set;

[0028] Calculating a multiplication value between the initial state loss value and a third preset weight coefficient to generate a target state loss value corresponding to the device query set;

[0029] The sum of the target startup time loss value, the target balance loss value, and the target state loss value is calculated to generate a loss value corresponding to the device query set.

[0030] Optionally, the step of performing reinforcement learning based on the scenario data corresponding to the target task data by the target meta-learning model to generate an optimal startup plan includes:

[0031] Constructing a comprehensive reward function using startup minimization data, load balancing optimization data, and equipment safety data corresponding to the target task data;

[0032] Reinforcing and updating the target meta-learning model using the scenario data corresponding to the target task data and the comprehensive reward function to generate a reinforcement model;

[0033] The startup plan is iteratively optimized using the policy gradient method through the reinforcement model to generate the optimal startup plan.

[0034] Optionally, the step of iteratively optimizing the startup plan using the policy gradient method through the reinforcement model to generate the optimal startup plan includes:

[0035] Perform action selection using a policy network through the reinforcement model to generate a first action;

[0036] Using the state data corresponding to the first action to perform reward calculation to generate a reward value;

[0037] Calculating the cumulative reward corresponding to the state data using the value network and the reward value to generate a cumulative reward value;

[0038] Using the cumulative reward value and policy gradient method to update the corresponding global parameters of the reinforcement model to generate policy network parameters;

[0039] When the policy network parameters do not meet the preset targets, the execution is skipped to the step of selecting an action using the policy network through the reinforcement model to generate a first action;

[0040] When the strategic network parameters meet the preset targets, the startup plan corresponding to the strategic network parameters is used as the optimal startup plan.

[0041] Optionally, it also includes:

[0042] Using the startup data corresponding to the optimal startup solution to update the target meta-learning model to generate an optimal meta-learning model;

[0043] The optimal meta-learning model is used as a new initial meta-model, and the step of obtaining the initial historical startup data and target task data is skipped.

[0044] A second aspect of the present invention provides a startup plan optimization system, comprising:

[0045] A data processing module is used to obtain initial historical startup data and target task data, perform data preprocessing on the initial historical startup data, and generate target historical startup data;

[0046] A model training module, configured to perform model training on global parameters of an initial meta-learning model based on the target historical startup data and the target task data to generate a target meta-learning model;

[0047] A solution generation module is used to perform reinforcement learning based on the scenario data corresponding to the target task data through the target meta-learning model to generate an optimal startup solution.

[0048] A third aspect of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the startup scheme optimization method as described in any one of the above items.

[0049] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed, implements the startup scheme optimization method as described in any one of the above items.

[0050] It can be seen from the above technical solutions that the present invention has the following advantages:

[0051] The present invention trains the global parameters of the meta-learning model and uses scenario data corresponding to the target task data for reinforcement learning, fully considering multiple requirements and generating the optimal startup plan. The meta-learning optimization method generates global initialization parameters applicable to a variety of scenarios by learning from historical substation startup scenarios, enabling deep learning or reinforcement learning models to quickly adapt to new scenarios, significantly shortening optimization time. The meta-learning model is trained on multiple task distributions and can extract common features during the substation startup process, ensuring that the generated initialization parameters have good generalization capabilities and adapt to the complex and changing substation startup requirements.

[0052] The present invention combines the dynamic optimization capabilities of reinforcement learning, and by feeding the startup strategy generated by reinforcement learning back to the meta-learning model, it continuously updates the initialization parameters to make it more in line with the dynamic characteristics of the actual scenario, further improving the optimization effect. The meta-learning optimization method can also simultaneously consider multiple objective requirements such as minimizing startup time, optimizing load balancing, and improving equipment safety, providing more comprehensive startup solution optimization support for substations. The optimal solution can be generated quickly, which improves the system startup efficiency and reduces resource consumption and operating costs during the substation startup process. It solves the technical problem that the startup solution generation method adopted in the existing technology has high requirements for system training, complex system implementation, and low system startup efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0054] Figure 1 A flowchart of the steps of a startup scheme optimization method provided in Example 1 of the present invention;

[0055] Figure 2 A flowchart of a method for optimizing a startup solution according to a second embodiment of the present invention;

[0056] Figure 3 This is the overall flow chart of meta-learning provided in the second embodiment of the present invention;

[0057] Figure 4 A flowchart of a startup solution optimization provided in the second embodiment of the present invention;

[0058] Figure 5 This is a structural block diagram of a startup solution optimization system provided in Example 3 of the present invention;

[0059] Figure 6 This is a structural block diagram of a computer device provided in Example 4 of the present invention. DETAILED DESCRIPTION

[0060] The embodiments of the present invention provide a startup scheme optimization method and system for solving the technical problems that the startup scheme generation method adopted in the prior art has high requirements for system training, complex system implementation, and leads to low system startup efficiency.

[0061] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0062] Against the backdrop of the intelligent development of power grids and the growth of complex load demands, existing power grid startup schemes are plagued by problems such as irrational startup sequences, high equipment insulation risks, and a lack of dynamic optimization capabilities. Therefore, it is necessary to develop intelligent startup control methods, optimize startup schemes, and establish a performance evaluation platform to standardize startup efficiency and safety. This will facilitate the intelligent upgrade of power grid operations and improve the system's startup efficiency, stability, and reliability.

[0063] See also Figure 1 , Figure 1 This is a flowchart of the steps of a startup scheme optimization method provided in Example 1 of the present invention.

[0064] The present invention provides a startup scheme optimization method, comprising:

[0065] Step 101: Acquire initial historical startup data and target task data, perform data preprocessing on the initial historical startup data, and generate target historical startup data.

[0066] In the grid startup of a large substation, the meta-learning model requires multi-dimensional data input to fully reflect the key characteristics of the substation startup process. The meta-learning model requires multi-dimensional data input to fully reflect the key characteristics of the substation startup process.

[0067] Initial historical startup data includes various data records from past substation startup processes. This data is used to train the meta-learning model, allowing it to learn common characteristics across different startup scenarios and generate global initialization parameters applicable to multiple scenarios. This data includes historical startup plan data, grid topology data, startup process monitoring data, and equipment status and performance data.

[0068] Historical startup plan data: This covers the history of different startup sequences and strategies, including parameters such as startup time, sequence, load distribution, and device status. Furthermore, performance indicators such as startup success rate, failure conditions, and abnormal load changes need to be recorded. This data can be obtained from grid operation logs and dispatch center records.

[0069] Grid topology data: This data details the connections and relationships between substation equipment (such as transformers, circuit breakers, and switches). This data takes into account dynamic changes such as equipment failures, load variations, and line disconnections, ensuring that startup plans can be adjusted in a timely manner. This data typically comes from a grid management system or simulation platform.

[0070] Startup process monitoring data: This includes transient data during startup (such as voltage, current, and frequency fluctuations) and abnormal phenomena (such as equipment tripping or overload signals). This data is obtained through a real-time monitoring system, which can be a phasor measurement unit (PMU) or a dynamic voltage restorer (DVR).

[0071] Equipment status and performance data: This includes pre-startup equipment health (e.g., insulation strength, temperature rise, and partial discharge characteristics) and dynamic response during startup (e.g., current and voltage changes). Equipment aging and historical maintenance records are also required. This data is collected through online monitoring equipment and equipment management systems.

[0072] Target task data refers to the specific substation startup task that needs to be solved at the moment. It is used to further adjust and optimize the model based on the global initialization parameters generated by the meta-learning model to generate a startup plan that meets the current task requirements.

[0073] The target historical startup data is obtained by cleaning and extracting features from the initial historical startup data to obtain multidimensional data.

[0074] Furthermore, the meta-learning model requires diverse and structured data input to support its rapid adaptability across different tasks. Each of these tasks represents a specific substation startup scenario, such as varying load distributions, equipment state combinations, or startup sequences. Each task, corresponding to the initial historical startup data and target task data, consists of a support set and a query set.

[0075] Assume that the task set is , where each task Contains support set (Support Set) and query set (Query Set), that is, tasks for:

[0076] ;

[0077] in, is the i-th task; The support set is used to quickly tune model parameters within the task; is the query set used to evaluate the model after in-task fine-tuning.

[0078] Each task Data Consists of input and output:

[0079] ;

[0080] in, is the input feature of the jth sample in the i-th task (topology in the substation, etc.); For the corresponding target output (start-up time of a device or transformer, load balance, etc.); is the number of samples in the task.

[0081] In an embodiment of the present invention, a data input module corresponding to a meta-learning model is used to obtain multidimensional initial historical startup data and target task data. This multidimensional data is then preprocessed, specifically, the initial historical startup data. Specifically, the data preprocessing process involves removing duplicate data from the initial historical startup data to generate first processed data. Then, a mean-filling method is used to fill missing values in the first processed data to generate second processed data. Noise is then removed from the second processed data to generate third processed data. Feature extraction is then performed on the third processed data according to preset extraction indicators to generate startup scenario indicators. Finally, the third processed data and the startup scenario indicators are used to construct the target historical startup data.

[0082] Step 102: Perform model training on the global parameters of the initial meta-learning model based on the target historical startup data and the target task data to generate a target meta-learning model.

[0083] The target meta-learning model refers to the model obtained by meta-training the initial meta-learning model using the target historical startup data and target task data.

[0084] In an embodiment of the present invention, the global parameters of the initial meta-learning model are first gradient-updated using the support set of the devices in the target task data to generate an intermediate meta-learning model. Next, a device query set is constructed using the query set of each device in the target task data and the query set of each device in the target historical startup data. Loss values are then calculated for each device query set in the device query set to generate multiple loss values. The sum of all loss values is then calculated to generate a query set loss value. Finally, the query set loss values and a preset learning rate are used to gradient-update the global parameters of the intermediate meta-learning model to generate a target meta-learning model.

[0085] Step 103: Perform reinforcement learning based on the scenario data corresponding to the target task data through the target meta-learning model to generate an optimal startup plan.

[0086] In this embodiment of the present invention, a comprehensive reward function is constructed using startup minimization data, load balancing optimization data, and device safety data corresponding to the target task data. The target meta-learning model is then enhanced and updated using the scenario data corresponding to the target task data and the comprehensive reward function to generate a reinforced model. Finally, the reinforced model is used to iteratively optimize the startup plan using the policy gradient method to generate the optimal startup plan.

[0087] In an embodiment of the present invention, initial historical startup data and target task data are obtained, and the initial historical startup data is preprocessed to generate target historical startup data. Based on the target historical startup data and target task data, the global parameters of the initial meta-learning model are trained to generate a target meta-learning model. The target meta-learning model performs reinforcement learning based on the scenario data corresponding to the target task data to generate an optimal startup plan. By training the global parameters of the meta-learning model and performing reinforcement learning on the scenario data corresponding to the target task data, the meta-learning optimization method is adopted to fully consider multiple requirements and generate an optimal startup plan. It has the following technical advantages:

[0088] 1. Rapid adaptability: The meta-learning optimization method generates global initialization parameters applicable to multiple scenarios by learning from historical substation startup scenarios. This enables deep learning or reinforcement learning models to quickly adapt to new scenarios, significantly shortening optimization time.

[0089] 2. Efficient generalization capability: The meta-learning model is trained on multiple task distributions and can extract common features during substation startup. This ensures that the generated initialization parameters have good generalization capabilities and can adapt to the complex and changing substation startup requirements.

[0090] 3. Dynamic Feedback Optimization: This method combines the dynamic optimization capabilities of reinforcement learning. By feeding the startup strategy generated by reinforcement learning back to the meta-learning model, it continuously updates the initialization parameters to make them more consistent with the dynamic characteristics of the actual scenario, further improving the optimization effect.

[0091] 4. Multi-objective optimization support: The meta-learning optimization method can also simultaneously consider multiple objectives such as minimizing startup time, optimizing load balancing, and improving equipment safety, providing more comprehensive startup plan optimization support for substations.

[0092] 5. Efficient resource utilization: By quickly generating optimization solutions, this method reduces unnecessary trial-and-error processes, improves optimization efficiency, and reduces resource consumption and operating costs during substation startup.

[0093] In summary, the meta-learning approach to optimizing substation startup strategies, through its rapid adaptability, efficient generalization, dynamic feedback optimization, and multi-objective support, provides an efficient and reliable optimization solution for substation startup, demonstrating significant practicality and cost-effectiveness. This approach addresses the technical issues faced by existing startup strategy generation methods, which often require high system training requirements and are complex to implement, resulting in low startup efficiency.

[0094] See also Figure 2 , Figure 2 This is a flowchart of the steps of a startup scheme optimization method provided in Example 2 of the present invention.

[0095] The present invention provides a startup scheme optimization method, comprising:

[0096] Step 201: Acquire initial historical startup data and target task data, perform data preprocessing on the initial historical startup data, and generate target historical startup data.

[0097] Furthermore, step 201 may include the following sub-steps:

[0098] S11: Delete duplicate data in the initial historical startup data to generate first processed data.

[0099] S12. Use the mean filling method to supplement the missing values in the first processed data to generate second processed data.

[0100] S13: Remove noise from the second processed data to generate third processed data.

[0101] S14: Extract features from the third processed data according to preset extraction indicators to generate a startup solution indicator.

[0102] S15. Use the third processed data and the startup plan indicator to construct target historical startup data.

[0103] The preset extraction indicators may be one or more of load peak, startup time delay, startup energy consumption, etc.

[0104] In the embodiment of the present invention, since the initial historical startup data is multi-dimensional data, it is necessary to pre-process the multi-dimensional data, including data cleaning and feature extraction.

[0105] Data cleaning: The collected initial historical startup data is subjected to noise removal, deduplication, and missing value processing. Preferably, the initial historical startup data is first deduplicated to obtain first processed data. Missing values in the first processed data are then corrected to obtain second processed data. Missing values can be filled using mean-filling to prevent data defects from interfering with model training. Finally, the second processed data is de-noised to obtain third processed data.

[0106] Feature extraction: Extract key features from the startup data, i.e., the third processed data obtained after data cleaning, such as load peak, startup time delay, startup energy consumption, etc., as important indicators for startup plan optimization, thereby obtaining startup plan indicators.

[0107] Step 202: Perform model training on the global parameters of the initial meta-learning model based on the target historical startup data and the target task data to generate a target meta-learning model.

[0108] Furthermore, step 202 may include the following sub-steps:

[0109] S21. Use the support set of the device in the target task data to gradient update the global parameters of the initial meta-learning model to generate an intermediate meta-learning model.

[0110] S22: Construct a device query set set by using the query set of each device in the target task data and the query set of each device in the target historical startup data.

[0111] S23. Using each device query set in the device query set set to perform loss value calculation on the global parameters of the intermediate meta-model, a plurality of loss values are generated.

[0112] S24. Calculate the sum of all loss values to generate the query set loss value.

[0113] S25. Use the query set loss value and the preset learning rate to gradient update the global parameters of the intermediate meta-learning model to generate the target meta-learning model.

[0114] In the embodiment of the present invention, Figure 3 As shown, firstly, task collection is performed to obtain the target history startup data and each task data in the target task data. The goal of the meta-learning model is to Learning optimal global parameters , making the task From the task distribution Extraction, where the task distribution In the meta-learning scenario, the probability distribution of all possible tasks T. In the specific application of substation startup plan optimization, each device startup task Can be regarded as a distribution from this task The samples extracted from .

[0115] Deep learning model with optimal global parameters When it is the initial parameter, only a small amount of training is required to Higher performance can be achieved and a better substation startup plan can be output.

[0116] ;

[0117] in, is the optimal global parameter, which enables the model to achieve better performance on related tasks; To pass the task Support set The task-specific parameters obtained by the above update are the global parameters corresponding to the intermediate meta-learning model; is a global parameter, which is a parameter to be learned and optimized in the meta-learning model. The value will affect the performance of the model on the task; For task distribution The i-th specific task extracted from ; For the task The loss function of To find the parameter that minimizes the value of the following expression , that is, by adjusting the parameters To minimize the subsequent objective function; For the mission The query set contains the query used to evaluate the model on the task Data samples of the above performance; The task distribution describes the probability distribution of all possible tasks in the target historical startup data and target task data. The model samples specific tasks from this distribution for learning.

[0118] In-task update: Use the support set of the device in the target task data to perform gradient update on the global parameters of the initial meta-learning model to obtain the intermediate meta-learning model. Above, the global parameters corresponding to the initial meta-learning model Perform one or more gradient updates to generate task-specific parameters , thus obtaining the intermediate meta-learning model.

[0119] ;

[0120] in, is the intra-task learning rate, which controls the step size of each parameter update. If the value is too large, the parameter update may skip the optimal solution; if the value is too small, the parameter update speed will be very slow and more training steps will be required to converge; is the gradient on the support set; To pass the task Support set The task-specific parameters obtained by the above update are the global parameters corresponding to the intermediate meta-learning model; is a global parameter, which is a parameter to be learned and optimized in the meta-learning model. The value will affect the performance of the model on the task; For task distribution The i-th specific task extracted from .

[0121] Global update: perform query set evaluation, using the query set of each device to launch the task Evaluate task-specific parameters The performance of the task is calculated and the sum of all loss values is calculated to generate the query set loss value. The query set loss value and the preset learning rate are used to adjust the global parameters according to the query set loss of all devices starting the task. Update. Finally get , that is, the query set loss value and the preset learning rate are used to gradient update the global parameters of the intermediate meta-learning model to generate the target meta-learning model.

[0122] ;

[0123] in, is the learning rate for global update; : query set loss for all tasks; For global parameters Calculate the gradient. The gradient reflects the rate and direction of change of the function at that point and is used to guide the direction of parameter update. N represents the number of tasks, that is, the number of tasks extracted from the task distribution for training. For the task In the query set With the updated parameters The calculated loss function value measures the model's performance in the task The difference between the predicted results and the actual results on the query set.

[0124] Furthermore, step S23 may include the following sub-steps:

[0125] S231 , respectively substituting the startup time data and expected startup time corresponding to the device query sets in the device query set set into a preset startup time loss formula to calculate an initial startup time loss value corresponding to the device query set.

[0126] S232: Calculate the product of the initial startup time loss value and the first preset weight coefficient to generate a target startup time loss value corresponding to the device query set.

[0127] S233: Substitute the load data and preset load distribution data in the device query set into a preset balance loss formula to calculate an initial balance loss value corresponding to the device query set.

[0128] S234: Calculate the product of the initial instantaneous balance value and the second preset weight coefficient to generate a target balance loss value corresponding to the device query set.

[0129] S235 , substituting the device state data corresponding to the device query set into a preset device state loss formula to calculate an initial state loss value corresponding to the device query set.

[0130] S236: Calculate the product of the initial state loss value and the third preset weight coefficient to generate a target state loss value corresponding to the device query set.

[0131] S237: Calculate the sum of the target startup time loss value, the target balance loss value, and the target state loss value to generate a loss value corresponding to the device query set.

[0132] It should be noted that in the scenario of substation startup optimization, the loss function design of the task needs to comprehensively consider the characteristics of substation startup, including startup time, load balancing and equipment safety. Calculate the multiplication between the initial startup time loss value and the first preset weight coefficient to obtain the target startup time loss value; calculate the multiplication between the initial balance instantaneous value and the second preset weight coefficient to generate the target balance loss value corresponding to the equipment query set; calculate the multiplication between the initial state loss value and the third preset weight coefficient to generate the target state loss value corresponding to the equipment query set. The query set loss value is composed of the sum of the target startup time loss value, the target balance loss value and the target state loss value, and the corresponding expression is as follows:

[0133] ;

[0134] in, is the query set loss value; The target startup time loss value refers to the startup time loss minimized when the device starts; is the target balance loss value, which refers to the balance loss of equipment load distribution; is the target state loss value, which refers to the loss related to the health state of the equipment; They are respectively a first preset weight coefficient, a second preset weight coefficient and a third preset weight coefficient, and the values of the preset weight coefficients are adjusted according to the specific requirements of the task.

[0135] Calculate the loss of the minimized startup time, that is, substitute the startup time data and expected startup time corresponding to the device query set in the device query set into the preset startup time loss formula, and calculate the initial startup time loss value corresponding to the device query set. The preset startup time loss formula used is:

[0136] ;

[0137] ;

[0138] in, Start a task for a device The input features of the jth sample (substation topology, etc.) are the basis for model prediction; The expected or optimal startup time can be used as a predefined target value to represent the startup time under ideal circumstances; Minimize the loss function for device boot time, which measures the difference between the device boot time predicted by the model and the expected boot time. The goal is to minimize this loss value as much as possible through training; For the task The number of samples in the support set indicates the number of samples involved in calculating the startup time loss; j is the sample index, which is used to traverse each sample in the support set of task j, from 1 to ; Based on input features and task-specific parameters The predicted device startup time is the predicted value given by the model; For the incumbent The global parameters on the support set of Perform one or more gradient updates to obtain task-specific parameters, which are used for prediction calculations of the model under that task. for The transpose of In the input feature Perform matrix multiplication operations to obtain predicted startup time related values.

[0139] Calculate the balance loss of load distribution by substituting the load data in the device query set and the preset load distribution data into the preset balance loss formula to calculate the initial balance loss value corresponding to the device query set. The preset balance loss formula used is:

[0140] ;

[0141] in, is the initial state loss value, that is, the target balance loss value; Start a task for a device The load of the jth sample in ; is the average value of the grid load, or can be defined as the most ideal load distribution; For the task The number of samples in the support set indicates the number of samples involved in calculating the startup time loss.

[0142] Calculate the loss related to the device health status by substituting the device status data corresponding to the device query set into the preset device status loss formula to calculate the initial state loss value corresponding to the device query set. The preset device status loss formula used is:

[0143] ;

[0144] in, is the initial state loss value; Start a task for a device The load of the jth sample in ; The maximum load that the equipment can safely withstand; For the task The number of samples in the support set indicates the number of samples involved in calculating the startup time loss.

[0145] Step 203: Perform reinforcement learning based on the scenario data corresponding to the target task data through the target meta-learning model to generate an optimal startup plan.

[0146] Furthermore, step 203 may include the following sub-steps:

[0147] S31. Use the startup minimization data, load balancing optimization data and equipment safety data corresponding to the target task data to construct a comprehensive reward function.

[0148] S32. Use the scenario data and comprehensive reward function corresponding to the target task data to enhance and update the target meta-learning model to generate an enhanced model.

[0149] S33. Use the policy gradient method to iteratively optimize the startup plan through the reinforcement model to generate the optimal startup plan.

[0150] In the embodiment of the present invention, the reward function is designed using the startup minimization data, load balancing optimization data, and equipment safety data corresponding to the target task data. The startup minimization data is used to calculate the minimum startup time. The corresponding expression is:

[0151] ;

[0152] in, Indicates the minimum startup time; Indicates the total startup time.

[0153] The load balancing optimization data is used to calculate the load balancing optimization value. The corresponding expression is:

[0154] ;

[0155] in, Indicates the load balancing optimization value; Indicates the degree of imbalance in load distribution.

[0156] Calculate device security using device security data :

[0157] ;

[0158] in, Indicates the health status of the device.

[0159] Therefore, the comprehensive reward function constructed by using the startup minimization data, load balancing optimization data and equipment safety data corresponding to the target task data is:

[0160] ;

[0161] in, They are the minimum startup time, load balancing optimization value and weight coefficient of equipment safety, and their values can be adjusted according to task requirements.

[0162] When reinforcement learning is applied to substation startup optimization, its goal is to learn a strategy that enables the system to automatically generate appropriate startup sequences and load distribution plans under different grid conditions, thereby minimizing startup time, optimizing load balancing, and ensuring equipment safety.

[0163] First, the scene data corresponding to the task data is used to define the environment and state space:

[0164] State Space: State Space Indicates the current status of the substation, including the following possible contents:

[0165] Substation topology: The connection status of each device in the substation.

[0166] Device status: The current operating status of each device (on, off, faulty, etc.).

[0167] Load demand: The current load demand of the substation, that is, the power that needs to be provided.

[0168] Equipment health status: whether the equipment is in good working condition, and possible fault or maintenance information.

[0169] Action space 𝑎: represents the actions that can be taken.

[0170] Start device: whether to start a certain device, such as starting a generator.

[0171] Adjust load distribution: for example, transfer part of the load from one area to another.

[0172] Set device priority: Control the order in which devices are started and which devices should be started first.

[0173] The target meta-learning model is then reinforced and updated using the comprehensive reward function, the defined environment, and the state space to obtain the reinforcement model.

[0174] Furthermore, step S33 may include the following sub-steps:

[0175] S331. Select an action using a policy network through a reinforcement model to generate a first action.

[0176] S332: Calculate the reward using the state data corresponding to the first action to generate a reward value.

[0177] S333: Calculate the cumulative reward corresponding to the state data using the value network and the reward value to generate a cumulative reward value.

[0178] S334. Use the cumulative return value and policy gradient method to update the corresponding global parameters of the reinforcement model and generate the policy network parameters.

[0179] S335: When the policy network parameters do not meet the preset targets, the execution jumps to the step of selecting an action using the policy network through the reinforcement model to generate a first action.

[0180] S336: When the strategic network parameters meet the preset targets, the startup plan corresponding to the strategic network parameters is used as the optimal startup plan.

[0181] In this embodiment of the present invention, the policy gradient method is used to optimize the global parameters by sampling the state-action-reward sequence. .

[0182] ;

[0183] in, To express strategy The objective function J is about the global parameter In the substation startup optimization, by adjusting the global parameters To optimize the startup plan, the gradient is used to guide the update direction of the parameters, so that the objective function changes in a better direction, such as making the startup time shorter and the load distribution more balanced. Is based on strategy In this context, it means that in the strategy Under the action selection method determined, the average expected operation is performed on the subsequent calculation results. In other words, considering the strategy The probability of occurrence of different state-action pairs is calculated and the relevant values are weighted averaged. It's a strategy In state The logarithmic probability of choosing action a with respect to the policy parameter The gradient of the policy parameter That is, global parameters How does a small change in the state affect In substation startup optimization, this gradient determines the degree of influence of adjustment strategy parameters on the probability of action selection such as equipment startup and load distribution.

[0184] represents the return starting from time step t, the formula is ,in, is the reward obtained at time step t, is a discount factor, and its value range is usually between [0,1]. In the substation startup scenario, the reward is determined by factors such as minimizing startup time, optimizing load balancing, and equipment safety. Discount factor This causes future rewards to gradually decay over time when calculating total returns, reflecting that current decisions place greater emphasis on short-term rewards while also taking into account long-term impacts. The policy network is based on the policy parameters , in state The probability distribution of selecting action a under the substation startup optimization process determines the selection probability of operations such as starting equipment, adjusting load distribution, and setting equipment priority under the current state of the substation.

[0185] The policy network is optimized using the following loss function:

[0186] ;

[0187] in, It is the loss function of the policy network (Actor), which is used to measure the gap between the policy network prediction and the ideal situation, and optimize the policy network parameters by minimizing the loss function. Is based on strategy The expectation operator indicates that according to the strategy Sampling, averaging the results of subsequent expressions. For the strategy Under the given state The probability of taking action a when . is the action probability The logarithm is introduced to facilitate the calculation of the gradient and avoid problems such as numerical underflow caused by continuous multiplication.

[0188] is the advantage function, which measures the superiority of action a:

[0189] ;

[0190] in, is the action value function, which represents the reward after performing action a. To estimate the state value function of the current state, based on the parameter The value network, evaluated in the state The long-term cumulative rewards that can be obtained by following the optimal strategy are used as a reference benchmark when calculating the advantage function, eliminating the value impact of the state itself, highlighting the additional value brought by the action, and reducing the variance of the rewards.

[0191] The value network is used to estimate the current state The value of is the expected future reward from this state. Update the value function by minimizing the mean square error.

[0192] ;

[0193] in, It is the loss function of the value network (Critic), which is used to measure the gap between the estimated value of the value network and the true value, and optimize the value network parameters by minimizing it. It is the expectation operator, which means averaging the results of subsequent expressions. Value Network , and adjust these parameters by optimizing the loss function to improve the accuracy of value estimation. The state of the value network output The estimated value of Starting from, the expected estimate of future returns. is the actual return obtained from time step t onwards, which is the true value and is used to calculate the loss compared with the value network estimate.

[0194] Generate an optimized startup plan through the policy network, that is, from the current state Generate action a:

[0195] ;

[0196] in, is the action generated by the policy network at time step t; is the policy network (with parameters ) in the state The probability distribution of taking action a under Indicates that the action is sampled from the probability distribution ; An action sequence is a series of actions. In a substation startup scenario, it represents the order and strategy arrangement of equipment startup.

[0197] The specific process of generating the optimal startup plan is as follows:

[0198] 1. Initialization state: The initial state of the substation, including equipment status, load demand, topology, etc., is to use the scenario data corresponding to the target task data and the comprehensive reward function to strengthen and update the target meta-learning model to generate a reinforcement model.

[0199] 2. Select action: according to the policy network , from the current state Select an action , that is, the strategy network is used to select actions through the reinforcement model to generate the first action.

[0200] 3. Execute the action and observe the new state: Execute the selected action, the state of the substation changes, and a new state is formed .

[0201] 4. Calculate rewards: Calculate rewards based on the new substation status , that is, the state data corresponding to the first action is used to calculate the reward and generate the reward value.

[0202] 5. Reward update: Based on the reward obtained and the current state-action pair , calculate the cumulative return , and use the policy gradient method to update the policy parameters That is, the value network and reward value are used to calculate the cumulative return corresponding to the state data, generate the cumulative return value, and use the cumulative return value and policy gradient method to update the corresponding global parameters of the reinforcement model to generate the policy network parameters.

[0203] 6. Repeat the above steps: Repeat steps 2 to 5 until the substation reaches the predetermined target. That is, when the strategy network parameters do not meet the preset targets, jump to the step of selecting actions using the strategy network through the reinforcement model to generate the first action; when the strategy network parameters meet the preset targets, the startup plan corresponding to the strategy network parameters is used as the optimal startup plan.

[0204] During training, reinforcement learning models continuously sample state-action-reward sequences to optimize their strategies. Through repeated sampling and learning, reinforcement learning generates an optimized launch plan.

[0205] Step 204: Use the startup data corresponding to the optimal startup solution to update the target meta-learning model to generate an optimal meta-learning model.

[0206] In this embodiment of the present invention, the startup strategy generated by reinforcement learning is fed back to the meta-learning model to update the global parameters and generate the optimal meta-learning model, further improving the adaptability and optimization effect of the solution. The specific startup solution feedback process is as follows:

[0207] The following data is collected from the execution of the reinforcement learning process:

[0208] ;

[0209] in, is a collection of data collected during the reinforcement learning process, where each element Represents a sample and records the status , actions taken And the rewards received , ; M is the number of sampled trajectories, that is, the data set The number of samples in .

[0210] Transform the data generated by reinforcement learning into new tasks for meta-learning. The new task set is constructed from reinforcement learning feedback:

[0211] ;

[0212] in, is a new task set constructed by reinforcement learning feedback, containing K new tasks ; Is a new task set The kth task in , each task Include:

[0213] Support Set : Partial state-action pairs from reinforcement learning trajectories.

[0214] Query Set : Used to evaluate the performance of reinforcement learning generation strategies.

[0215] Step 205: Use the optimal meta-learning model as a new initial meta-model, and jump to the step of obtaining initial historical startup data and target task data.

[0216] In an embodiment of the present invention, a new task is added to the task distribution of meta-learning:

[0217] ;

[0218] in, is the original task distribution of the meta-learning task, indicating the probability of different tasks occurring; It is a newly generated set of tasks, constructed by reinforcement learning feedback, which will be added to the meta-learning task distribution.

[0219] Retraining of meta-learning models:

[0220] Perform parameter tuning on the support set of the new task:

[0221] ;

[0222] in, For the new task Model parameters after parameter adjustment on the support set; is the intra-task learning rate, which controls the step size of each parameter update. If the value is too large, the parameter update may skip the optimal solution; if the value is too small, the parameter update speed will be very slow and more training steps will be required to converge; is a global parameter, which is a parameter to be learned and optimized in the meta-learning model. The value will affect the performance of the model on the task; For task distribution The i-th specific task extracted from ; is the loss function About global parameters In the support set The gradient on is used to guide the direction of parameter adjustment.

[0223] Evaluate performance on the query set of the new task and update the global parameters:

[0224] ;

[0225] Integrate the effect feedback of reinforcement learning into the meta-learning loss function and design an improved hybrid optimization objective:

[0226] ;

[0227] in, An improved hybrid optimization objective function that combines meta-learning and reinforcement learning feedback. It is the traditional task loss function for meta-learning, which is used to maintain the generalization ability of the model. It is a task loss function for reinforcement learning feedback, used to adapt to dynamic scenes. and They are the weight coefficients of the traditional task loss function and the task loss function, respectively, balancing the two objectives.

[0228] Meta-learning is used to generate global initialization parameters applicable to multiple scenarios, and reinforcement learning is combined to achieve dynamic startup optimization. The meta-learning model generates initialization parameters through feature extraction of historical startup data, task decomposition, and multi-task training, providing rapid adaptability for deep learning or reinforcement learning models. The reinforcement learning module optimizes the startup strategy for specific scenarios through dynamic interaction with the power grid environment and feeds the optimization results back to the meta-learning model, using the optimal meta-learning model as the new initial meta-model and jumping to the steps of obtaining initial historical startup data and target task data, further improving the generalization performance of the initialization parameters. Multi-task training combined with a task feedback mechanism effectively reduces the convergence time of the model. At the same time, the dynamic interaction process optimizes the practicality of the startup plan, forming a closed-loop optimization system, thereby significantly improving the efficiency and reliability of the substation startup plan.

[0229] Further, if Figure 4 As shown in the figure, a large substation connects power supply to multiple areas, including commercial and industrial loads and residential electricity consumption. The substation houses multiple generators, transformers, and a complex grid topology. The substation is undergoing annual maintenance, and multiple devices need to be restarted, including aging equipment and new equipment. Due to the large number of devices and the stringent requirements for startup sequencing, load balancing, and equipment safety, traditional startup solutions are prone to equipment overload, uneven startup, and premature failure.

[0230] In this complex scenario, by adopting optimization methods based on meta-learning and reinforcement learning, dynamic optimization of the startup plan can be achieved to ensure grid safety, reduce equipment losses, and improve startup efficiency.

[0231] 1) Data collection and preprocessing.

[0232] Obtaining initial historical startup data, or a subset of historical tasks: Collect historical startup scenarios from substation data, including startup time, sequence, load distribution, and success rates. Assume this data includes seasonal grid load variations, equipment failure records, and abnormal startup fluctuations. Then, perform data preprocessing to obtain the target historical startup data.

[0233] Substation topology data: Obtain a grid topology diagram, including the connections between generators, transformers, switches, and lines. Assume that the substation has three main transformers and four backup generators, connected to two grid regions. This topology data will serve as the basis for training the meta-learning model.

[0234] Equipment health and monitoring data: By monitoring equipment health data (such as transformer insulation strength, current, and voltage fluctuations) in real time and historical maintenance records, it is possible to identify an aging transformer in poor health and prioritize its activation.

[0235] Load demand data: This collects real-time load demand data from substations, including changes in power demand and grid load distribution. If power demand surges during a specific period, resulting in significant load volatility, dynamic adjustments to load distribution are necessary.

[0236] 2) Meta-learning model and training.

[0237] After collecting relevant data, the meta-learning model will be trained using this data to quickly adapt to different startup scenarios of the substation.

[0238] Task decomposition and support set: The grid startup task is broken down into multiple subtasks, each representing a specific device startup or load distribution task. For example, one task might be "Start transformer 1 and distribute the load," while another might be "Restart generator 2 and adjust the load." Each task has a corresponding support set and query set.

[0239] Training and global parameter update: Using a meta-learning algorithm, we train a support set of various grid startup tasks to ensure that the meta-learning model can extract common features and provide fast-adapting global initialization parameters for the deep learning model.

[0240] 3) Optimize the startup sequence of reinforcement learning, that is, optimize the reinforcement learning scenario.

[0241] After the meta-learning model is trained, it enters the reinforcement learning phase to generate the best startup plan.

[0242] State space setting: The state space includes the real-time status of substation equipment (for example, whether the equipment is in standby mode, load demand, etc.) and the topology of the power grid (such as whether the equipment is connected, whether the load is evenly distributed, etc.).

[0243] Action space setting: Action space includes starting a device, adjusting load distribution, setting device priority, etc. For example, starting transformer 1, shutting down generator 3, or transferring load from one area to another.

[0244] Reward function design: The reward function is set based on multiple objectives such as startup time, load balancing, and device safety. The optimization goal is to minimize startup time while ensuring load balancing and device safety.

[0245] Policy Optimization: A reinforcement learning algorithm generates an optimal policy to prioritize startup and distribute load, minimizing startup shock and equipment damage. By continuously sampling the state-action-reward sequence, the model gradually adjusts the policy until the optimal startup plan is reached.

[0246] 4) Implementation and feedback mechanism, that is, feeding back the optimization results of reinforcement learning.

[0247] After the startup plan generated by reinforcement learning is executed, the equipment status, load changes, etc. during the startup process are monitored in real time, and these data are fed back to the meta-learning model.

[0248] Real-time feedback: After the startup plan optimized by reinforcement learning is generated, it begins execution. During execution, the startup status of each device is monitored to check whether the load distribution and device status meet expectations.

[0249] Optimization closed loop: Execution results are fed back to the meta-learning model, which updates global parameters based on the support set and query set of the new task and optimizes the launch strategy. Through this closed loop process, the model can improve the adaptability and optimization effect of the solution through continuous feedback and adjustment.

[0250] The present invention aims to improve the intelligence, dynamic optimization, and multi-scenario adaptability of power grid startup plans, while further reducing equipment losses and operational risks during the startup process. The focus is on achieving real-time optimization and feedback loops for startup plans. By collecting historical data, topology, monitoring data, and equipment performance from substation startups, cleaning and feature extraction are performed to ensure data accuracy and integrity. A meta-learning approach is used to train multiple startup scenarios, generating global initialization parameters suitable for different scenarios and providing rapid adaptability for subsequent models. Reinforcement learning is combined to optimize startup sequence and load distribution, and a policy gradient approach is used to generate startup plans in real time, optimizing startup time, load balance, and equipment safety. The startup strategy generated by reinforcement learning is fed back to the meta-learning model to update global parameters, further enhancing the plan's adaptability and optimization effectiveness. Through multi-task training and dynamic optimization mechanisms, it can adapt to changes in the power grid environment in real time, improving startup efficiency and safety. Therefore, the present invention achieves the technical effect of generating a universal parameter framework through meta-learning, providing optimization guidance for deep learning models, and dynamically generating optimal plans based on diverse startup scenarios, enabling rapid adaptation to new tasks while improving the efficiency and safety of system startup.

[0251] See also Figure 5 , Figure 5 This is a structural block diagram of a startup scheme optimization system provided in Example 3 of the present invention.

[0252] The present invention provides a startup scheme optimization system, comprising:

[0253] The data processing module 501 is used to obtain the initial historical startup data and the target task data, perform data preprocessing on the initial historical startup data, and generate the target historical startup data;

[0254] A model training module 502 is configured to perform model training on the global parameters of the initial meta-learning model based on the target historical startup data and the target task data to generate a target meta-learning model;

[0255] The solution generation module 503 is used to perform reinforcement learning based on the scenario data corresponding to the target task data through the target meta-learning model to generate an optimal startup solution.

[0256] Furthermore, the data processing module 501 includes:

[0257] A first processed data generating module, configured to delete duplicate data in the initial historical startup data and generate first processed data;

[0258] A second processed data generating module is used to supplement the missing values in the first processed data by using a mean filling method to generate second processed data;

[0259] a third processed data generating module, configured to remove noise from the second processed data and generate third processed data;

[0260] A startup plan indicator generation module, configured to extract features from the third processed data according to preset extraction indicators to generate startup plan indicators;

[0261] The target historical startup data construction module is used to construct the target historical startup data by using the third processed data and the startup plan indicator.

[0262] Furthermore, the model training module 502 includes:

[0263] An intermediate meta-learning model generation module is used to perform gradient updates on the global parameters of the initial meta-learning model using the support set of the device in the target task data to generate an intermediate meta-learning model;

[0264] A device query set building module is used to build a device query set using the query set of each device in the target task data and the query set of each device in the target historical startup data;

[0265] A loss value calculation module is used to calculate the loss value of the global parameters of the intermediate meta-model using each device query set in the device query set to generate multiple loss values;

[0266] The query set loss value calculation module is used to calculate the sum of all loss values and generate the query set loss value;

[0267] The target meta-learning model generation module is used to perform gradient updates on the global parameters of the intermediate meta-learning model using the query set loss value and the preset learning rate to generate the target meta-learning model.

[0268] Furthermore, the loss value calculation module can perform the following steps:

[0269] Substituting the startup time data and expected startup time corresponding to the device query set in the device query set into the preset startup time loss formula respectively, and calculating the initial startup time loss value corresponding to the device query set;

[0270] Calculate the product of the initial startup time loss value and the first preset weight coefficient to generate a target startup time loss value corresponding to the device query set;

[0271] Substituting the load data and preset load distribution data in the device query set into the preset balance loss formula, and calculating the initial balance loss value corresponding to the device query set;

[0272] Calculating the product of the initial instantaneous balance value and the second preset weight coefficient to generate a target balance loss value corresponding to the device query set;

[0273] Substitute the device state data corresponding to the device query set into the preset device state loss formula to calculate the initial state loss value corresponding to the device query set;

[0274] Calculate the product of the initial state loss value and the third preset weight coefficient to generate a target state loss value corresponding to the device query set;

[0275] Calculate the sum of the target startup time loss, target balance loss, and target state loss to generate the loss value corresponding to the device query set.

[0276] Furthermore, the solution generation module 503 includes:

[0277] A comprehensive reward function construction module is used to construct a comprehensive reward function using startup minimization data, load balancing optimization data, and equipment safety data corresponding to the target task data;

[0278] The reinforcement model generation module is used to enhance and update the target meta-learning model using the scenario data corresponding to the target task data and the comprehensive reward function to generate a reinforcement model;

[0279] The scheme iterative optimization module is used to iteratively optimize the startup scheme using the policy gradient method through the reinforcement model to generate the optimal startup scheme.

[0280] Furthermore, the solution iteration optimization module can perform the following steps:

[0281] The strategy network is used to select actions through the reinforcement model to generate the first action;

[0282] The state data corresponding to the first action is used to calculate the reward and generate a reward value;

[0283] Use the value network and reward value to calculate the cumulative return corresponding to the state data and generate the cumulative return value;

[0284] Use the cumulative return value and policy gradient method to update the corresponding global parameters of the reinforcement model and generate the policy network parameters;

[0285] When the policy network parameters do not meet the preset goals, the execution jumps to the step of selecting an action using the policy network through the reinforcement model to generate the first action;

[0286] When the policy network parameters meet the preset targets, the startup plan corresponding to the policy network parameters is used as the optimal startup plan.

[0287] Furthermore, the system also includes:

[0288] An optimal meta-learning model generation module is used to update the target meta-learning model using the startup data corresponding to the optimal startup scheme to generate an optimal meta-learning model;

[0289] The jump module is used to use the optimal meta-learning model as the new initial meta-model and jump to the steps of obtaining the initial historical startup data and target task data.

[0290] See also Figure 6 , Figure 6 This is a structural block diagram of a computer device provided in Example 4 of the present invention.

[0291] An electronic device according to an embodiment of the present invention includes: a memory 601 and a processor 602, wherein the memory 601 stores a computer program; when the computer program is executed by the processor 602, the processor 602 executes the startup scheme optimization method as described in any of the above embodiments.

[0292] Memory 601 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Memory 601 has storage space 603 for program code 613 for executing any of the method steps described above. For example, storage space 603 for program code may include individual program codes 613 for implementing various steps in the method described above. These program codes may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards, or floppy disks. The program codes may be compressed, for example, in a suitable format. When executed by a processing device, these codes cause the processing device to execute the various steps in the method described above. These program codes may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards, or floppy disks. The program codes may be compressed, for example, in a suitable format. When these codes are executed by a computing and processing device, the computing and processing device is caused to execute the various steps in the startup solution optimization method described above.

[0293] The fifth embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the startup solution optimization method of any of the above embodiments is implemented.

[0294] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0295] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0296] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0297] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0298] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0299] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A startup plan optimization method, characterized in that: include: Acquire initial historical startup data and target task data, perform data preprocessing on the initial historical startup data, and generate target historical startup data; Performing model training on global parameters of an initial meta-learning model based on the target historical startup data and the target task data to generate a target meta-learning model; The target meta-learning model performs reinforcement learning based on the scenario data corresponding to the target task data to generate an optimal startup plan.

2. The startup scheme optimization method according to claim 1, characterized in that: The step of preprocessing the initial historical startup data to generate target historical startup data includes: Deleting duplicate data in the initial historical startup data to generate first processed data; Using a mean filling method to fill in missing values in the first processed data to generate second processed data; removing noise from the second processed data to generate third processed data; Perform feature extraction on the third processed data according to preset extraction indicators to generate a startup plan indicator; Target historical startup data is constructed using the third processed data and the startup plan indicator.

3. The startup scheme optimization method according to claim 1, characterized in that: The step of performing model training on the global parameters of the initial meta-learning model based on the target historical startup data and the target task data to generate a target meta-learning model includes: Performing gradient updates on global parameters of the initial meta-learning model using a support set of the device in the target task data to generate an intermediate meta-learning model; Constructing a device query set set by using a query set of each device in the target task data and a query set of each device in the target historical startup data; Using each device query set in the device query set set to respectively calculate the loss value of the global parameter of the intermediate meta-model to generate multiple loss values; Calculate the sum of all the loss values to generate a query set loss value; The query set loss value and a preset learning rate are used to perform gradient update on the global parameters of the intermediate meta-learning model to generate a target meta-learning model.

4. The startup scheme optimization method according to claim 3, characterized in that: The step of using each device query set in the device query set to calculate the loss value of the global parameter of the intermediate meta-model to generate multiple loss values includes: Substituting the startup time data and expected startup time corresponding to the device query sets in the device query set into a preset startup time loss formula respectively, and calculating an initial startup time loss value corresponding to the device query set; Calculating a multiplication value between the initial startup time loss value and a first preset weight coefficient to generate a target startup time loss value corresponding to the device query set; Substituting the load data and preset load distribution data in the device query set into a preset balance loss formula to calculate an initial balance loss value corresponding to the device query set; Calculating a multiplication value between the initial instantaneous balance value and a second preset weight coefficient to generate a target balance loss value corresponding to the device query set; Substituting the device state data corresponding to the device query set into a preset device state loss formula to calculate the initial state loss value corresponding to the device query set; Calculating a multiplication value between the initial state loss value and a third preset weight coefficient to generate a target state loss value corresponding to the device query set; The sum of the target startup time loss value, the target balance loss value, and the target state loss value is calculated to generate a loss value corresponding to the device query set.

5. The startup scheme optimization method according to claim 1, characterized in that: The step of performing reinforcement learning based on the scenario data corresponding to the target task data by the target meta-learning model to generate an optimal startup plan includes: Constructing a comprehensive reward function using startup minimization data, load balancing optimization data, and equipment safety data corresponding to the target task data; Reinforcing and updating the target meta-learning model using the scenario data corresponding to the target task data and the comprehensive reward function to generate a reinforcement model; The startup plan is iteratively optimized using the policy gradient method through the reinforcement model to generate the optimal startup plan.

6. The startup scheme optimization method according to claim 5, characterized in that: The step of iteratively optimizing the startup plan using the policy gradient method through the reinforcement model to generate the optimal startup plan includes: Perform action selection using a policy network through the reinforcement model to generate a first action; Using the state data corresponding to the first action to perform reward calculation to generate a reward value; Calculating the cumulative reward corresponding to the state data using the value network and the reward value to generate a cumulative reward value; Using the cumulative reward value and policy gradient method to update the corresponding global parameters of the reinforcement model to generate policy network parameters; When the policy network parameters do not meet the preset targets, the execution is skipped to the step of selecting an action using the policy network through the reinforcement model to generate a first action; When the strategic network parameters meet the preset targets, the startup plan corresponding to the strategic network parameters is used as the optimal startup plan.

7. The startup scheme optimization method according to claim 1, characterized in that: Also includes: Using the startup data corresponding to the optimal startup solution to update the target meta-learning model to generate an optimal meta-learning model; The optimal meta-learning model is used as a new initial meta-model, and the step of obtaining the initial historical startup data and target task data is skipped.

8. A startup plan optimization system, characterized in that: include: A data processing module is used to obtain initial historical startup data and target task data, perform data preprocessing on the initial historical startup data, and generate target historical startup data; A model training module, configured to perform model training on global parameters of an initial meta-learning model based on the target historical startup data and the target task data to generate a target meta-learning model; A solution generation module is used to perform reinforcement learning based on the scenario data corresponding to the target task data through the target meta-learning model to generate an optimal startup solution.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the startup scheme optimization method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the startup scheme optimization method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Starting scheme intelligent generation system and method based on deep learning

    CN117540741A

  • Power grid new equipment starting scheme generation method, system, equipment and medium

    CN117609769A

Cited By

  • Insulation performance optimization system of distribution transformer totally-enclosed area

    CN120870783A