Data-driven feedback control hyperparameter tuning methods, systems, equipment, media, and products

CN122569115APending Publication Date: 2026-08-14FOSHAN POWER SUPPLY BUREAU GUANGDONG POWER GRID
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-26
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

然而,数据驱动方法大多存在需要整定的超参数,对收敛性、控制效果等方面影响很大,超参数-控制效果的非线性耦合关系难以明晰,传统依赖经验或穷举搜索的参数整定方式难以适应边缘侧运行场景的动态变化,导致超参数整定效率较低,且精度较差

Benefits of technology

[0040]从以上技术方案可以看出,本发明通过在配电网的云侧结合历史控制策略数据与预设的数据驱动控制模型,获得各运行场景下的超参数组合对应的控制效果评价值,并构建训练样本集,从而利用训练样本集,构建多个运行场景共享的超参数组合与控制效果评价值之间的概率代理模型,从而在算力高的云侧完成先验知识的建模,并基于云侧的知识迁移机制,将概率代理模型和训练样本集下发至边缘侧,从而在算力受限的边缘侧,利用贝叶斯优化算法,对超参数组合进行寻优,得到超参数组合的最优整定结果,从而利用云边知识迁移让边缘侧可直接复用全局先验知识,结合贝叶斯优化快速适配边缘侧运行场景的动态变化,显著提升超参数整定效率与精度,保障量测反馈驱动控制的收敛性与实际控制效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122569115A_ABST
    Figure CN122569115A_ABST
Patent Text Reader

Abstract

This invention relates to the field of data-driven optimization technology, and discloses a data-driven feedback control hyperparameter tuning method, system, device, medium, and product. This method obtains control effect evaluation values ​​corresponding to hyperparameter combinations under various operating scenarios by combining historical control strategy data and a preset data-driven control model on the cloud side of the distribution network, and constructs a training sample set. Using this training sample set, a probabilistic proxy model is built between the shared hyperparameter combinations and control effect evaluation values ​​across multiple operating scenarios. Based on the cloud-side knowledge transfer mechanism, the probabilistic proxy model and the training sample set are distributed to the edge side. A Bayesian optimization algorithm is used to optimize the hyperparameter combinations, obtaining the optimal tuning result. This leverages cloud-edge knowledge transfer to allow the edge side to directly reuse global prior knowledge, and combined with Bayesian optimization, it quickly adapts to the dynamic changes of the edge side's operating scenarios, significantly improving the efficiency and accuracy of hyperparameter tuning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data-driven optimization technology, and in particular to a data-driven feedback control hyperparameter tuning method, system, device, medium, and product. Background Technology

[0002] The widespread integration of new power sources and loads has led to significant fluctuations, time-varying characteristics, and scenario-specific variations in the operation of distribution networks. Traditional control methods that rely on precise physical models and fixed parameters suffer from difficulties in modeling, insufficient adaptability, and degraded control performance in scenarios with incomplete network parameters, frequent changes in operating modes, and high requirements for real-time control at the edge.

[0003] To reduce reliance on precise physical parameters, measurement feedback-driven control methods have gained increasing attention. These methods utilize measurement data to identify the mapping relationship between control inputs and outputs, enabling power distribution network operation control without relying on accurate physical parameters. However, most data-driven methods involve hyperparameters that require tuning, significantly impacting convergence and control performance. The nonlinear coupling relationship between hyperparameters and control effects is difficult to clarify, and traditional parameter tuning methods relying on experience or exhaustive search are ill-suited to the dynamic changes in edge-side operating scenarios, resulting in low hyperparameter tuning efficiency and poor accuracy. Summary of the Invention

[0004] In view of this, in order to solve the above-mentioned technical problems, the present invention provides a data-driven feedback control hyperparameter tuning method, system, device, medium and product.

[0005] The first aspect of this invention provides a data-driven feedback control hyperparameter tuning method, comprising:

[0006] Historical control strategy data of the distribution network under multiple operating scenarios are acquired; on the cloud side of the distribution network, the historical control strategy data and a preset data-driven control model are combined to obtain the control effect evaluation value corresponding to the hyperparameter combination under each operating scenario, and a training sample set is constructed; wherein, the data-driven control model is used to characterize the nonlinear mapping relationship between the control strategy data and the control effect evaluation value of the distribution network under the hyperparameter combination.

[0007] Based on the training sample set, a probabilistic proxy model is constructed between the hyperparameter combination and the control effect evaluation value under multiple shared operating scenarios.

[0008] Based on the knowledge transfer mechanism of the cloud side, the probabilistic proxy model and the training sample set are distributed to the edge side as prior knowledge. The hyperparameter combination is optimized by combining the Bayesian optimization algorithm and the prior knowledge to obtain the optimal tuning result of the hyperparameter combination.

[0009] In one embodiment, the process of constructing the data-driven control model on the cloud side includes:

[0010] Acquire historical control strategy data and multiple control effect measurement data uploaded from the edge side of the target transformer area of ​​the power distribution network to the cloud side;

[0011] The control effect evaluation value is determined by weighting and summing multiple control effect measurement data according to a preset control target weighting factor.

[0012] Based on the historical control strategy data and the preset hyperparameter combination, the control effect evaluation value is nonlinearly predicted to obtain the data-driven control model.

[0013] In one embodiment, the step of nonlinearly predicting the control effect evaluation value based on the historical control strategy data and a preset hyperparameter combination to obtain the data-driven control model includes:

[0014] An initial dynamic mapping matrix is ​​obtained, and linear prediction is performed based on the change values ​​of the historical control strategy data and the initial dynamic mapping matrix to obtain the predicted value of the change in control effect.

[0015] Based on the control effect evaluation value, determine the true value of the control effect change; based on the prediction error between the predicted value of the control effect change and the true value of the control effect change, combine the preset dynamic mapping matrix update step size and the preset dynamic mapping matrix update penalty factor to correct the initial dynamic mapping matrix until the prediction error converges to a preset threshold, and obtain the corrected dynamic mapping matrix.

[0016] Based on the deviation between the control effect evaluation value and the preset control effect reference value, the control strategy data is updated by combining the corrected dynamic mapping matrix, the preset control strategy update step size and the preset control strategy penalty factor, until the deviation reaches the preset deviation threshold, and the updated control strategy data is obtained.

[0017] The hyperparameter combination is formed based on the control target weight factor, the dynamic mapping matrix update step size, the dynamic mapping matrix update penalty factor, the control strategy update step size, and the control strategy penalty factor;

[0018] Based on the implicit coupling relationship between the hyperparameter combination, the updated control strategy data, and the control effect evaluation value, a nonlinear mapping data-driven control model is constructed.

[0019] In one embodiment, the training sample set includes hyperparameter combinations, scene labels, and control effect evaluation values;

[0020] The construction of a probabilistic proxy model based on the training sample set, relating the hyperparameter combination to the control effect evaluation value under multiple shared operating scenarios, includes:

[0021] The hyperparameter combination and the scene label are extracted using a multilayer perceptron. The total loss function is weighted and fused with error term and regularization term. The total loss function is minimized through backpropagation to obtain the latent feature representation.

[0022] Based on the latent feature representation, a probabilistic surrogate model between the hyperparameter combination and the control effect evaluation value is established using a Gaussian process.

[0023] In one embodiment, the hyperparameter combination is optimized by combining the Bayesian optimization algorithm and the prior knowledge to obtain the optimal tuning result of the hyperparameter combination, including:

[0024] Based on the training sample set, and combining the scene labels with the probabilistic proxy model, the modeling error of the candidate scenes under each scene label is determined;

[0025] Based on the training sample set and the probabilistic proxy model, determine the regret reduction rate of the control effect prediction between adjacent iterations;

[0026] Based on the regret reduction rate and the modeling error, the scenario with the largest information gain is selected from the candidate scenarios or the target scenarios as the optimization scenario for this time;

[0027] Candidate data samples for the current optimization scenario are generated using Monte Carlo methods. The optimal hyperparameter combination that satisfies the performance benchmark constraints is selected from all the candidate data samples and used as the optimal hyperparameter combination for this optimization.

[0028] The training sample set is updated based on the candidate data samples used in this optimization to obtain the updated training sample set;

[0029] Based on the updated training sample set, the process of determining the modeling error of candidate scenes under each scene label by combining the scene label and the probabilistic proxy model according to the training sample set is repeated until the preset iteration stopping condition is met, and the optimal hyperparameter combination obtained in the last iteration is determined as the optimal tuning result of the hyperparameter combination.

[0030] In one embodiment, updating the training sample set based on the candidate data samples from the current optimization to obtain the updated training sample set includes:

[0031] Based on the candidate data samples for this optimization and each data sample in the training sample set, determine the sample similarity between the candidate data samples for this optimization and each of the data samples.

[0032] The updated training sample set is obtained by removing data samples whose similarity is higher than a preset tolerance threshold from the training sample set and adding the candidate data samples selected for this optimization.

[0033] Secondly, the present invention also provides a data-driven feedback control hyperparameter tuning system, comprising:

[0034] A sample set construction module is used to acquire historical control strategy data of the distribution network under multiple operating scenarios; on the cloud side of the distribution network, the historical control strategy data is combined with a preset data-driven control model to obtain the control effect evaluation value corresponding to the hyperparameter combination under each operating scenario, and a training sample set is constructed; wherein, the data-driven control model is used to characterize the nonlinear mapping relationship between the control strategy data and the control effect evaluation value of the distribution network under the hyperparameter combination.

[0035] The probabilistic proxy determination module is used to construct a probabilistic proxy model between the hyperparameter combination and the control effect evaluation value under multiple shared operating scenarios based on the training sample set.

[0036] The hyperparameter optimization module is used to distribute the probabilistic proxy model and the training sample set as prior knowledge to the edge side based on the cloud-side knowledge transfer mechanism, and combine the Bayesian optimization algorithm and the prior knowledge to optimize the hyperparameter combination and obtain the optimal tuning result of the hyperparameter combination.

[0037] Thirdly, the present invention also provides an electronic device, the electronic device including a memory and a processor, the memory storing a computer program, the computer program being executed by the processor causing the processor to perform the steps of the data-driven feedback control hyperparameter tuning method as described in the first aspect.

[0038] Fourthly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed, implements the steps of the data-driven feedback control hyperparameter tuning method as described in the first aspect.

[0039] Fifthly, the present invention also provides a computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer performs the steps of the data-driven feedback control hyperparameter tuning method as described in the first aspect.

[0040] As can be seen from the above technical solutions, this invention obtains the control effect evaluation values ​​corresponding to hyperparameter combinations under various operating scenarios by combining historical control strategy data and a preset data-driven control model on the cloud side of the distribution network, and constructs a training sample set. Using this training sample set, a probabilistic proxy model is built between the hyperparameter combinations and control effect evaluation values ​​shared by multiple operating scenarios. This allows for the modeling of prior knowledge on the cloud side with high computing power. Based on the cloud side's knowledge transfer mechanism, the probabilistic proxy model and training sample set are distributed to the edge side. On the edge side, where computing power is limited, a Bayesian optimization algorithm is used to optimize the hyperparameter combinations, obtaining the optimal tuning result. This cloud-edge knowledge transfer allows the edge side to directly reuse global prior knowledge. Combined with Bayesian optimization, it quickly adapts to the dynamic changes of the edge side's operating scenarios, significantly improving the efficiency and accuracy of hyperparameter tuning and ensuring the convergence and actual control effect of measurement feedback-driven control. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a schematic diagram of the structure of a communication system in an application environment of a data-driven feedback control hyperparameter tuning method provided in an embodiment of the present invention;

[0043] Figure 2 A flowchart of a data-driven feedback control hyperparameter tuning method provided in an embodiment of the present invention;

[0044] Figure 3 A schematic diagram of the improved IEEE 33-node structure;

[0045] Figure 4a This is a schematic diagram illustrating the changes in cloud-side solar power fluctuation data during clear weather.

[0046] Figure 4b This is a schematic diagram illustrating the changes in cloud-side solar power fluctuation data on cloudy days.

[0047] Figure 4c This is a schematic diagram illustrating the changes in cloud-side photovoltaic fluctuation data during rainy weather.

[0048] Figure 5a This is a schematic diagram illustrating the changes in cloud-side wind turbine fluctuation data during clear weather.

[0049] Figure 5b This is a schematic diagram illustrating the changes in wind turbine fluctuation data on the cloud side during cloudy weather.

[0050] Figure 5c This is a schematic diagram illustrating the changes in wind turbine fluctuation data on the cloud side during rainy weather.

[0051] Figure 6 This is a schematic diagram illustrating the changes in cloud-side load fluctuation data;

[0052] Figure 7a This is a schematic diagram illustrating the changes in clear-sky fluctuation data on the edge side.

[0053] Figure 7b This is a schematic diagram illustrating the changes in the cloudy fluctuation data on the edge side.

[0054] Figure 7c This is a schematic diagram illustrating the changes in rainy weather fluctuation data on the edge side;

[0055] Figure 8 This is a schematic diagram of the multi-scenario joint optimization process under the voltage control requirements of Scheme II;

[0056] Figure 9a A comparative diagram showing the optimization effect under sunny conditions for voltage control requirements;

[0057] Figure 9b A comparative diagram showing the optimization effect under voltage control requirements on cloudy days;

[0058] Figure 9c A comparative diagram illustrating the optimization effect of voltage control requirements on rainy days;

[0059] Figure 10 This is a schematic diagram of a data-driven feedback control hyperparameter tuning system provided in an embodiment of the present invention;

[0060] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0061] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] The widespread integration of new power sources and loads has led to significant fluctuations, time-varying characteristics, and scenario-specific variations in the operation of distribution networks. Traditional control methods that rely on precise physical models and fixed parameters suffer from difficulties in modeling, insufficient adaptability, and degraded control performance in scenarios with incomplete network parameters, frequent changes in operating modes, and high requirements for real-time control at the edge.

[0063] To reduce reliance on precise physical parameters, measurement feedback-driven control methods have gained increasing attention. These methods utilize measurement data to identify the mapping relationship between control inputs and outputs, enabling power distribution network operation control without relying on accurate physical parameters. However, most data-driven methods involve hyperparameters that require tuning, significantly impacting convergence and control performance. The nonlinear coupling relationship between hyperparameters and control effects is difficult to clarify, and manual parameter tuning is time-consuming and costly. Traditional parameter tuning methods relying on experience or exhaustive search are ill-suited to the dynamic changes in edge-side operating scenarios, and their computational costs and trial-and-error risks are often unacceptable in practical engineering.

[0064] Considering the diverse operating scenarios and control requirements at the edge of the distribution network, the hyperparameter tuning problem of data-driven feedback control needs to be expanded from single-scenario, single-time offline tuning to a cloud-edge collaborative tuning and updating problem oriented towards multiple distribution areas, multiple scenarios, and multiple control requirements. A hyperparameter tuning model capable of representing multi-scenario tuning knowledge and supporting multi-objective adaptation at the edge needs to be established as the foundation for online control optimization. In this problem, there are complex nonlinear coupling relationships between multi-scenario operating characteristics, hyperparameter combinations, and control effects. From an implementation perspective, this is a complex tuning problem involving knowledge transfer, posing significant challenges to engineering applications. To accurately and efficiently solve these problems, a method is needed that can complete multi-scenario tuning experience-based modeling at the cloud side and achieve rapid tuning at the edge side by combining actual data. This will yield hyperparameter configuration strategies adapted to different operating scenarios and control requirements, ensuring the cross-scenario adaptability of data-driven feedback control in the distribution network.

[0065] To address the aforementioned issues, this application proposes a data-driven feedback control hyperparameter tuning method. By combining historical control strategy data with a pre-set data-driven control model on the cloud side of the distribution network, control effect evaluation values ​​corresponding to hyperparameter combinations under various operating scenarios are obtained, and a training sample set is constructed. Using this training sample set, a probabilistic proxy model is built between the shared hyperparameter combinations and control effect evaluation values ​​across multiple operating scenarios. This allows for the modeling of prior knowledge on the cloud side with high computing power. Based on the cloud side's knowledge transfer mechanism, the probabilistic proxy model and training sample set are distributed to the edge side. On the edge side, where computing power is limited, a Bayesian optimization algorithm is used to optimize the hyperparameter combinations, obtaining the optimal tuning result. This leverages cloud-edge knowledge transfer to allow the edge side to directly reuse global prior knowledge. Combined with Bayesian optimization, it rapidly adapts to the dynamic changes of edge side operating scenarios, significantly improving the efficiency and accuracy of hyperparameter tuning and ensuring the convergence and actual control effect of measurement feedback-driven control.

[0066] The data-driven feedback control hyperparameter tuning method provided in this application can be applied to, for example... Figure 1The communication system in the application environment shown is as follows. Edge side 101 communicates with cloud side 102 via a network. The data storage system can store data that cloud side 102 needs to process. The data storage system can be integrated onto cloud side 102, or it can be located in the cloud or on other network servers. The communication system implements a data-driven feedback control hyperparameter tuning method, which includes: acquiring historical control strategy data of the distribution network under multiple operating scenarios; combining the historical control strategy data with a preset data-driven control model on the cloud side 102 of the distribution network to obtain the control effect evaluation value corresponding to the hyperparameter combination under each operating scenario, and constructing a training sample set; wherein, the data-driven control model is used to characterize the nonlinear mapping relationship between the control strategy data and the control effect evaluation value of the distribution network under the hyperparameter combination; based on the training sample set, constructing a probabilistic proxy model between the hyperparameter combination and the control effect evaluation value under multiple operating scenarios; based on the knowledge transfer mechanism of the cloud side 102, distributing the probabilistic proxy model and the training sample set as prior knowledge to the edge side 101, and combining the Bayesian optimization algorithm and prior knowledge to optimize the hyperparameter combination and obtain the optimal tuning result of the hyperparameter combination.

[0067] Edge side 101 is the intelligent terminal on the distribution network area side, responsible for collecting local measurement data in real time, executing control strategies and providing feedback on control effects; Edge side 101 can be, but is not limited to, a smart meter, an RTU (Remote Terminal Unit) or a computer device with edge computing capabilities;

[0068] Cloudside 102 is a cloud-side server that is equipped with high-performance computing units and model training frameworks. Cloudside 102 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.

[0069] like Figure 2 As shown, this application provides a data-driven feedback control hyperparameter tuning method, which is applied to... Figure 1 Taking the communication system in the example, the explanation includes the following steps S1 to S3. Wherein:

[0070] Step S1: Obtain historical control strategy data of the distribution network under multiple operating scenarios; combine the historical control strategy data with the preset data-driven control model on the cloud side of the distribution network to obtain the control effect evaluation value corresponding to the hyperparameter combination under each operating scenario, and construct a training sample set; wherein, the data-driven control model is used to characterize the nonlinear mapping relationship between the control strategy data and the control effect evaluation value of the distribution network under the hyperparameter combination.

[0071] Among them, historical control strategy data of the distribution network under multiple operating scenarios are obtained through the edge side 101. The historical control strategy data includes the distribution network topology, load data, distributed power source operation data, and reactive power output data.

[0072] Historical control strategy data is uploaded to the cloud-side 102. After receiving the data, the cloud-side 102 inputs it into a preset data-driven control model to perform power flow simulation calculations. It outputs the control effect evaluation value corresponding to the hyperparameter combination under each scenario and constructs a structured training sample set containing hyperparameter combinations, scenario labels (category labels of the running scenario, such as sunny, cloudy, and rainy days) and corresponding evaluation values.

[0073] The data-driven control model represents the power distribution network operation control problem as a discrete-time nonlinear mapping relationship between input and output quantities. It also uses a dynamic mapping matrix to characterize the nonlinear mapping relationship between control strategy data and control effect evaluation values. This nonlinear mapping relationship includes the implicit coupling relationship between control strategy data and control effect evaluation values, and incorporates the hyperparameter combination in the implicit coupling relationship.

[0074] Hyperparameter combination is a set of adjustable parameters in a data-driven control model that enables the control strategy data and the control effect evaluation value to achieve the optimal mapping accuracy. It is the variable to be optimized in this application.

[0075] Among them, the control effect evaluation value is a key indicator for quantifying the effectiveness of the control strategy. It is generally obtained through measurement data acquisition devices and includes node voltage deviation rate and network loss.

[0076] Step S2: Based on the training sample set, construct a probabilistic proxy model between hyperparameter combinations and control effect evaluation values ​​under multiple shared operating scenarios.

[0077] The probabilistic surrogate model uses a Gaussian process as its core, integrates neural networks to extract latent feature representations from the training sample set, and characterizes the uncertainty mapping relationship between hyperparameter combinations and control effect evaluation values. Through hyperparameter combinations, the mean and variance of the corresponding control effect evaluation values ​​can be predicted, thereby supporting Bayesian inference of control effects under unknown hyperparameter combinations.

[0078] Step S3: Based on the knowledge transfer mechanism of the cloud side, the probabilistic proxy model and training sample set are sent to the edge side as prior knowledge. Combined with the Bayesian optimization algorithm and prior knowledge, the hyperparameter combination is optimized to obtain the optimal tuning result of the hyperparameter combination.

[0079] Among them, the knowledge transfer mechanism of cloud-side 102 uses model compression and quantization techniques to compress the probabilistic proxy model and training sample set into lightweight parameter representations, which are then embedded into the edge-side local inference engine as the initial knowledge for edge-side 101 local tuning;

[0080] The probabilistic proxy model trained on cloud side 102 and the training sample set are distributed to edge side 101 as follows:

[0081]

[0082] In the formula, This represents the initial probabilistic proxy model for the edge-side platform area. This represents a probabilistic proxy model on the cloud side. This represents the initial data sample set for the edge-side platform area. This is the training sample set for the cloud side.

[0083] By migrating the proxy model and sample set trained on the cloud side 102 to the edge side 101, the prior knowledge in the cloud is deployed in a lightweight manner on the edge side 101. Under the constraint of limited computing power, the edge side 101 can fully inherit the parameter tuning experience of the cloud side 101 without collecting data and training the model again, which significantly improves the efficiency of hyperparameter tuning and cross-scenario generalization ability. At the same time, based on this prior knowledge, the edge side 101 only needs a small amount of local test data to complete the Bayesian optimization iteration, quickly converge to the optimal solution that meets the convergence stopping condition, and accurately find the optimal hyperparameter combination that is suitable for the local area, thereby greatly reducing communication overhead and computation latency.

[0084] This application embodiment obtains the control effect evaluation values ​​corresponding to hyperparameter combinations under various operating scenarios by combining historical control strategy data and a preset data-driven control model on the cloud side of the distribution network, and constructs a training sample set. Using this training sample set, a probabilistic proxy model is built between the hyperparameter combinations and control effect evaluation values ​​shared by multiple operating scenarios. This allows for the modeling of prior knowledge on the cloud side with high computing power. Based on the cloud side's knowledge transfer mechanism, the probabilistic proxy model and training sample set are distributed to the edge side. On the edge side, where computing power is limited, a Bayesian optimization algorithm is used to optimize the hyperparameter combinations, obtaining the optimal tuning result. This cloud-edge knowledge transfer allows the edge side to directly reuse global prior knowledge. Combined with Bayesian optimization, it quickly adapts to the dynamic changes of the edge side's operating scenarios, significantly improving the efficiency and accuracy of hyperparameter tuning, and ensuring the convergence and actual control effect of measurement feedback-driven control.

[0085] In some embodiments, the process of constructing a data-driven control model on the cloud side includes: acquiring historical control strategy data and multiple control effect measurement data uploaded from the edge side of the target distribution area to the cloud side; weighting and summing the multiple control effect measurement data according to a preset control target weight factor to determine the control effect evaluation value; and performing nonlinear prediction on the control effect evaluation value based on the historical control strategy data and a preset hyperparameter combination to obtain the data-driven control model.

[0086] Specifically, based on the selected distribution network edge area, historical control strategy data and multiple control effect measurement data of the area are obtained, and the historical control strategy data and multiple control effect measurement data are input to the cloud side.

[0087] The control effect measurement data can include node voltage deviation and network loss, and can also include multi-dimensional indicators such as voltage fluctuation rate, harmonic distortion rate and equipment temperature rise, depending on actual needs.

[0088] Preferably, the control effect measurement data includes node voltage deviation and network loss. Then, weighting factors are set for each control objective, and the node voltage deviation and network loss are weighted and summed to obtain the control effect evaluation value, i.e.:

[0089]

[0090] In the formula, Let t be the control effect evaluation value. Let be the node voltage deviation vector at time t. To control the target weighting factor, The network loss measurement at time t; where the initial control target weight factor is... Configure voltage deviation and network loss according to their priority in the control objectives.

[0091] Then, by combining historical control strategy data with preset hyperparameters, the control effect evaluation value is nonlinearly fitted to construct a data-driven control model. This model takes control strategy data and hyperparameter combinations as inputs and control effect evaluation values ​​as outputs, thereby achieving accurate prediction and dynamic response of the control effect.

[0092] In some embodiments, a data-driven control model is obtained by nonlinearly predicting the control effect evaluation value based on historical control strategy data and a preset hyperparameter combination. This includes: acquiring an initial dynamic mapping matrix and performing linear prediction based on the changes in historical control strategy data and the initial dynamic mapping matrix to obtain a predicted control effect change value; determining the actual control effect change value based on the control effect evaluation value; and correcting the initial dynamic mapping matrix based on the prediction error between the predicted control effect change value and the actual control effect change value, combined with a preset dynamic mapping matrix update step size and a preset dynamic mapping matrix update penalty factor, until the prediction error converges to a preset threshold. The corrected dynamic mapping matrix is ​​obtained; based on the deviation between the control effect evaluation value and the preset control effect reference value, the control strategy data is updated by combining the corrected dynamic mapping matrix, the preset control strategy update step size, and the preset control strategy penalty factor until the deviation reaches the preset deviation threshold, thus obtaining the updated control strategy data; a hyperparameter combination is constructed based on the control target weight factor, the dynamic mapping matrix update step size, the dynamic mapping matrix update penalty factor, the control strategy update step size, and the control strategy penalty factor; based on the implicit coupling relationship between the hyperparameter combination, the updated control strategy data, and the control effect evaluation value, a nonlinear mapping data-driven control model is constructed.

[0093] The initial dynamic mapping matrix is ​​a sparse matrix that characterizes the initial dynamic local coupling relationship between the change in control input and the change in measurement output, i.e.:

[0094]

[0095] In the formula, Let be the dynamic mapping matrix at time t. express Dynamically map the elements inside the matrix at any time. and These represent row and column indices, respectively. and These represent the dimensions of the output and input variables, respectively.

[0096] The changes in historical control strategy data are the difference sequences of historical control strategy data at different sampling times. This difference sequence reflects the dynamic evolution trend of the control strategy data. Simultaneously, by performing linear prediction using the changes in historical control strategy data and the initial dynamic mapping matrix, the predicted changes in control effect are obtained, i.e.:

[0097]

[0098] In the formula, express Predicted change in control effect at any given time. Let t be the change value of the control strategy data at time t.

[0099] Then, according to The difference between the control effect evaluation value at time t and the control effect evaluation value at time t is used to determine the true value of the change in control effect.

[0100] By exploiting the prediction error between the predicted and actual changes in control effect, and combining this with a preset dynamic mapping matrix update step size and a preset dynamic mapping matrix update penalty factor, the initial dynamic mapping matrix is ​​corrected. This corrected dynamic mapping matrix more accurately characterizes the time-varying coupling characteristics between control input and measurement output, thereby supporting real-time response and dynamic adaptation of the edge side to multi-objective coordinated control of the distribution network. The estimated change value of the dynamic mapping matrix is:

[0101]

[0102] In the formula, for The estimated change of the dynamic mapping matrix at time step. This indicates the step size for updating the dynamic mapping matrix. This indicates that the penalty factor is updated in the dynamic mapping matrix. for The change value of the control strategy data at any given time. for The estimated value of the dynamic mapping matrix at time t. Let be the prediction error between the actual change in control effect at time t and the predicted change in control effect.

[0103] By minimizing the prediction error between the predicted value of the change in control effect and the actual value of the change in control effect, until the prediction error is less than a preset threshold, the dynamic mapping matrix is ​​considered to have fully converged. At this point, the estimated change value of the final dynamic mapping matrix is ​​added to the initial dynamic mapping matrix to obtain the corrected dynamic mapping matrix.

[0104] Subsequently, based on the deviation between the control effect evaluation value and the preset control effect reference value, and in conjunction with the corrected dynamic mapping matrix, the preset control strategy update step size, and the preset control strategy penalty factor, the control strategy data is updated. The updated control strategy data changes are as follows:

[0105]

[0106] In the formula, The updated change value of the control strategy data at time t. Indicates the control policy update step size. This indicates that the control strategy updates the penalty factor. This indicates the preset control effect reference value. The corrected dynamic mapping matrix is... express The dynamic mapping matrix estimates the change value at each moment until the deviation reaches a preset deviation threshold, thus obtaining the final updated change value of the control strategy data. The updated change value of the control strategy data is then added to the current control strategy data to obtain the updated control strategy data.

[0107] Among them, hyperparameter combination .

[0108] Then, through the aforementioned update process of control strategy data, update process of dynamic mapping matrix, and real-time feedback of control effect evaluation value, and with the hyperparameter combination controlling the update rhythm and control target weight of each update process, the dynamic mapping relationship between control strategy data and control effect evaluation value is deeply coupled and abstracted into an implicit coupling function F, thus obtaining the data-driven control model as follows:

[0109]

[0110] The data-driven control model directly characterizes the nonlinear dynamic response characteristics of the control strategy input and multidimensional measurement output under hyperparameter adjustment.

[0111] In some embodiments, the training sample set includes hyperparameter combinations, scene labels, and control effect evaluation values. In this case, based on the training sample set, a probabilistic proxy model between hyperparameter combinations and control effect evaluation values ​​under multiple shared operating scenarios is constructed, including: using a multilayer perceptron to extract features from hyperparameter combinations and scene labels, combining a weighted fusion of error terms and regularization terms into a total loss function, minimizing the total loss function through backpropagation to obtain a latent feature representation; and using a Gaussian process to establish a probabilistic proxy model between hyperparameter combinations and control effect evaluation values ​​based on the latent feature representation.

[0112] This method employs a multilayer perceptron (MLP) as the neural network to extract high-dimensional correlation features between hyperparameter combinations and scene labels, mapping them to a low-dimensional latent space. A total loss function, incorporating error and regularization terms, is set, and backpropagation minimizes this total loss function to obtain stable latent feature representations. Specifically, each sample's hyperparameter combination and scene label are one-hot encoded, normalized, and then concatenated to form the model's original input vector. Simultaneously, the control effect evaluation value corresponding to the sample is used as the model's true label for training. A MLP is constructed with the preprocessed hyperparameters and scene labels as input. The input layer receives the concatenated original features and performs feature learning through multiple hidden layers (using ReLU activation function for nonlinear transformation). The output layer outputs a low-dimensional, highly adaptable latent feature representation, achieving filtering of redundant information from hyperparameters and scene labels, deep extraction of nonlinear correlation features, and reducing subsequent modeling complexity. Based on the latent features output by the network, a total loss function including an error term and a regularization term is constructed. The backpropagation algorithm and gradient optimizer are used to iteratively update the network weights and biases of the multilayer perceptron. In each iteration, the total loss function value is calculated and gradually minimized until the loss function converges and no longer decreases significantly. At this point, training stops, resulting in a trained multilayer perceptron network. The features output by this network are used as the optimal latent feature representation, accurately characterizing the intrinsic coupling relationship between hyperparameters, scene labels, and control effects. The latent features are represented as follows:

[0113]

[0114] In the formula, z is a multilayer perceptron after embedding mapping. The transformed latent feature representation, For the scene The assessment cost is as follows.

[0115] The total loss function is expressed as:

[0116]

[0117]

[0118]

[0119] In the formula, Represents the total loss function. This represents the prediction error term; For regularization terms, Indicates the number of data samples. Indicates the data sample index. Indicates the first Evaluation value of the control effect of the group sample Indicates the first Group of hyperparameter combinations Indicates the first Group input data scene labels, and Let these represent the hyperparameters of the neural network and the Gaussian process, respectively. Indicates the index of the neural network layer. Indicates the number of layers in a neural network. express Layer weight matrix, Represents any non vector.

[0120] The probabilistic surrogate model established using Gaussian processes between hyperparameter combinations and control performance evaluation values ​​is as follows:

[0121]

[0122]

[0123]

[0124]

[0125] In the formula, This indicates the evaluation value of the control effect. This represents a probabilistic proxy model built on the cloud side. This represents the training sample set used by the probabilistic proxy model built on the cloud side. Representing neural networks The weight matrix obtained during training, The hyperparameters of a Gaussian process are represented. Represents a Gaussian process. This represents the mean of the control effectiveness evaluation values. This represents the variance of the control effectiveness evaluation value. Represents the kernel function. and Representing sample sets respectively The Middle and training samples, Represents the index of any sample point to be evaluated. This represents the control effect of any sample point to be evaluated. , Represents the identity matrix. This represents the noise variance of the Gaussian process.

[0126] In some embodiments, the hyperparameter combination is optimized by combining Bayesian optimization algorithm and prior knowledge to obtain the optimal tuning result of the hyperparameter combination, including: determining the modeling error of candidate scenarios under each scene label based on the training sample set, combined with scene labels and probabilistic surrogate model; determining the regret reduction rate of control effect prediction between adjacent iterations based on the training sample set and probabilistic surrogate model; selecting the scene with the largest information gain from candidate scenarios or target scenarios as the scene to be optimized in this iteration based on the regret reduction rate and modeling error; generating candidate data samples under the scene to be optimized in this iteration through Monte Carlo, and selecting the optimal hyperparameter combination that meets the performance benchmark constraint from all candidate data samples as the optimal hyperparameter combination in this iteration; updating the training sample set based on the candidate data samples in this iteration to obtain the updated training sample set; and re-executing the process of determining the modeling error of candidate scenarios under each scene label based on the training sample set, combined with scene labels and probabilistic surrogate model, until the preset iteration stopping condition is met, and obtaining the optimal hyperparameter combination obtained in the last iteration as the optimal tuning result of the hyperparameter combination.

[0127] Among them, modeling error is the fitting deviation of the probabilistic surrogate model to the actual control effect in a specific scenario. The smaller the value, the more accurate the model's characterization of the current scenario. The calculation process of modeling error is as follows:

[0128]

[0129] In the formula, Indicates the first During the second optimization process Modeling error term for each scenario Indicates edge side cutoff Secondary optimization scenario Data samples, for The number of data samples in the middle, express The Middle A combination of hyperparameters, Indicates the first Hyperparameter combinations in various scenarios The predicted value of the control effect This indicates that the model trained without samples has different hyperparameter combinations. The corresponding predicted control effect value.

[0130] Regret reduction rate is a metric that measures the contribution of a scenario to the improvement of the target performance during iterative optimization. The regret reduction rate is calculated as follows:

[0131]

[0132] In the formula, The rate of decline is regrettable. To predict the minimum control effect of scenario s during the k-th optimization, For the first -1 optimization iterations to predict the minimum control effect of scenario s The cost of evaluating scenario s.

[0133] Based on the regret reduction rate and modeling error, the scene with the largest information gain is selected from the candidate or target scenes as the optimization scene for this search, i.e.:

[0134]

[0135] In the formula, For the first The optimization scenario during the second optimization. For the target scenario, Indicates candidate scenarios, This represents the modeling error difference between the candidate scene and the target scene.

[0136] Among them, the target scenario is the ultimate target scenario for system performance optimization, and its hyperparameter combination needs to be dynamically calibrated in real time at the edge; the candidate scenario serves as an exploratory test field to verify the generalization ability and robustness of the control strategy under unknown operating conditions.

[0137] After obtaining the current optimization scenario, candidate data samples are generated for this scenario through Monte Carlo sampling. These candidate data samples include control effect sample values ​​and candidate hyperparameter combinations for the current optimization scenario, and the posterior mean of all control effect sample values ​​is determined.

[0138] Then, the performance constraint sampling domain of the performance benchmark constraint is constructed. By combining the hyperparameters of candidate data samples with the performance-constrained sampling domain Perform intersection filtering, retaining only hyperparameters that satisfy the dual constraints: First constraint: the lower confidence bound corresponding to the hyperparameter. >Original benchmark The first layer ensures that even the worst-case control performance is better than no control; the second layer is the upper confidence bound corresponding to the hyperparameters. > The maximum value of the confidence bound for all hyperparameters guarantees that the parameter has optimal potential; where the performance constraint sampling domain is... for:

[0139]

[0140] Represents the numerical space of hyperparameter combinations. This represents the evaluation value of the distribution network operation control effect without the adoption of control strategies (original baseline). Indicates the lower confidence boundary. It represents the upper confidence boundary.

[0141] Then, construct the acquisition function under performance baseline constraints:

[0142]

[0143] In the formula, Indicates the collected value. Indicates the first Feature variables of the hyperparameter combination of the second optimization. The sampling domain is constrained by the performance benchmark. Indicates an indicator function, when In In the middle, its value is 1, otherwise it is 0. Indicates the index of the collected samples. This indicates the number of Monte Carlo samplings performed on each sample point. Indicates the Monte Carlo sampling index. Indicates the first During the second optimization process, the first The sample point at the th th The posterior mean of the control effect sample values ​​from the Monte Carlo sampling. Indicates the first During the second optimization process, the first The sample point at the th th The control effect sample values ​​of the Monte Carlo sampling This represents the exploration coefficient.

[0144] By selecting the hyperparameter combination with the largest collected value through the acquisition function, the optimal hyperparameter combination for this iteration is chosen. This approach leverages performance benchmark constraints and the acquisition function to ensure that the parameters meet the performance benchmark requirements while accurately identifying the parameters with the greatest optimization potential, thus avoiding blind searching.

[0145] Then, since there is a discrepancy between the sample set distributed by the cloud and the actual operation on the edge, it is necessary to compare the sample used for optimization with the sample distributed by the cloud and remove sample points similar to those on the cloud to improve the efficiency of customized optimization on the edge. Therefore, this application also updates the training sample set based on the candidate data samples used for optimization, and re-executes the modeling error of the candidate scenes under each scene label by combining the scene label and the probabilistic proxy model based on the updated training sample set until the preset iteration stopping condition is met, and the optimal hyperparameter combination obtained in the last iteration is determined as the optimal tuning result of the hyperparameter combination.

[0146] The iteration stopping condition is when the number of iterations reaches the preset maximum number of iterations (e.g., 300) or the total evaluation cost of hyperparameter tuning reaches the upper limit, such as: In the formula, This indicates the upper limit of the assessment cost.

[0147] In some embodiments, updating the training sample set based on the candidate data samples of this optimization to obtain an updated training sample set includes: determining the sample similarity between the candidate data samples of this optimization and each data sample in the training sample set based on the candidate data samples of this optimization and each data sample in the training sample set; removing data samples with sample similarity higher than a preset tolerance threshold from the training sample set and adding the candidate data samples of this optimization to obtain an updated training sample set.

[0148] The calculation process for sample similarity is as follows:

[0149]

[0150] In the formula, For sample similarity, and Indicates the weighting coefficient. The first inherited by the edge side Hyperparameter combinations for a set of samples Indicates the edge side Hyperparameter sample points for secondary optimization The first one representing the edge side inheritance The control effect of the group sample on the sample value Indicates the edge side Sample values ​​of the control effect of the second optimization.

[0151] Then, samples with similarity higher than the preset tolerance threshold are removed from the training sample set. The data samples are used to obtain the discarded sample set:

[0152]

[0153] In the formula, Indicates the first The next best way to eliminate sample sets. Indicates the cloud side Group of sample points.

[0154] After obtaining the discarded sample set, the discarded sample set is filtered out from the training sample set, and the new samples obtained in this optimization are added to form the dynamic updated training sample set on the edge side.

[0155] In addition, the optimal hyperparameter combination is obtained offline for different control requirements, and the corresponding hyperparameter combination is deployed according to the scenario and requirements during actual operation on the edge side.

[0156] For example, if the control requirement is voltage control, the voltage deviation is taken as the control effect. The corresponding voltage deviation value is obtained by running the measurement feedback control model, and the optimal hyperparameter combination under the control requirement is searched using the Bayesian optimization method with a limited number of evaluations.

[0157] For different control requirements, the optimal hyperparameter combinations are stored for different scenarios and matched and called in real time during runtime on the edge side. By performing the above offline optimization process for different control requirements, multiple sets of "control requirements - optimal hyperparameter combinations" correspondences can be obtained. Based on the control requirements determined by the current running status on the edge side, the corresponding optimal hyperparameter combination can be directly called, or a partial online update can be performed on it, thereby realizing adaptive control for different control requirements.

[0158] Next, an example will be used to illustrate the data-driven feedback control hyperparameter tuning method of this application.

[0159] In this example, the impedance values ​​of line elements, the active power reference values ​​and power factors of load elements, and the network topology of the improved IEEE 33-node system are first input. The structure of the IEEE 33-node system is as follows: Figure 3 As shown, detailed parameters are shown in Tables 1 and 2. Six wind turbine systems (WT) with a capacity of 500 kVA are connected at nodes 13, 15, 16, 22, 29, and 30 respectively. Twelve photovoltaic systems (PV) with a capacity of 100 kWp are connected at nodes 11, 12, 17, 18, 20, 21, 23, 24, 25, 31, 32, and 33. The system voltage is 12.66 kV.

[0160] Table 1. Load connection locations and power in the improved IEEE 33-node example.

[0161]

[0162] Table 2. Improved IEEE 33-node example line parameters

[0163]

[0164] Under different operating scenarios, such as typical meteorological conditions like cloudy, sunny, and rainy days, the random fluctuation characteristics of photovoltaic (PV) and wind turbine (WTM) output under each meteorological scenario are modeled as time-correlated fluctuation curves. The cloud-side PV and WTM output fluctuation data under each meteorological scenario are shown below. Figure 4a -4c and Figure 5a As shown in –5c, the load fluctuation data on the cloud side is as follows: Figure 6 As shown, the edge-side data for photovoltaic, wind power, and load fluctuations under various meteorological scenarios are as follows: Figure 7a –7c; This example uses five schemes for comparative analysis:

[0165] Option I (Comparison Option): Without using control methods, the initial operating state of the power distribution system is obtained;

[0166] Scheme II (the scheme proposed in this application): The proposed method is used to perform hyperparameter tuning of data-driven control and to control the distribution network;

[0167] Option III (Comparison Option): Set data-driven control hyperparameters based on experience to control the power distribution network;

[0168] Option IV (Control Option): Use the theoretically optimal physical method to regulate the distributed power generation cluster;

[0169] Option V (Control Option): Hyperparameter tuning is performed using Bayesian optimization based on the fundamental kernel function.

[0170] The computer hardware environment for performing the optimized calculations was an Intel(R) Core(TM) CPU I7-11700 with a clock speed of 3.70GHz and 16GB of memory; the software environment was a Windows 11 operating system.

[0171] The hyperparameter tuning results corresponding to the method proposed in this application are shown in Tables 3 and 4 for voltage control and network loss control requirements, respectively. The results comparison is shown in Tables 5 and 6.

[0172] Table 3 Voltage control parameter configuration:

[0173]

[0174] Table 4. Network loss control parameter configuration:

[0175]

[0176] Table 5 Comparison of Voltage Deviation Control Effects

[0177]

[0178] Table 6 Comparison of Network Loss Control Effects

[0179]

[0180] The optimization process of the proposed method (Scheme II) in the voltage control scenario is described in [link to relevant documentation]. Figure 8 The algorithm's optimization performance in multiple scenarios is as follows: Figures 9a-9c As shown.

[0181] The calculation results show that Scheme II, employing the method proposed in this application, significantly outperforms the baseline Scheme I and the empirical parameter Scheme III under both voltage deviation and network loss control requirements. Regarding voltage deviation control, compared to Scheme I, Scheme II reduces voltage deviation by 77.62%, 68.88%, and 68.84% in sunny, cloudy, and rainy weather scenarios, respectively; compared to Scheme III, it further reduces voltage deviation by 19.91%, 2.51%, and 15.39%. Regarding network loss control, compared to Scheme I, Scheme II reduces network loss by 32.63%, 31.75%, and 39.62% in sunny, cloudy, and rainy weather scenarios, respectively; compared to Scheme III, it further reduces network loss by 37.23%, 22.73%, and 32.43%. Furthermore, the overall results of Scheme II of the proposed method are quite close to the results of the solution based on the physical model in Scheme IV. Moreover, compared with the Bayesian optimization method based on the basic kernel function, the optimization effect is better in multiple scenarios. This shows that by adopting the method proposed in this application, the control effect of the data-driven algorithm can be significantly improved, and the tuning efficiency is high.

[0182] Based on the same inventive concept, this application also provides a data-driven feedback control hyperparameter tuning system for implementing the data-driven feedback control hyperparameter tuning method described above.

[0183] The solution provided by this system is similar to the solution described in the above method. Therefore, the specific limitations of one or more data-driven feedback control hyperparameter tuning system embodiments provided below can be found in the limitations of the data-driven feedback control hyperparameter tuning method described above, and will not be repeated here.

[0184] like Figure 10 As shown in the figure, this application provides a data-driven feedback control hyperparameter tuning system, including:

[0185] The sample set construction module 100 is used to acquire historical control strategy data of the distribution network under multiple operating scenarios; on the cloud side of the distribution network, the historical control strategy data and the preset data-driven control model are combined to obtain the control effect evaluation value corresponding to the hyperparameter combination under each operating scenario, and a training sample set is constructed; among them, the data-driven control model is used to characterize the nonlinear mapping relationship between the control strategy data and the control effect evaluation value of the distribution network under the hyperparameter combination.

[0186] The probabilistic proxy determination module 200 is used to construct a probabilistic proxy model between hyperparameter combinations and control effect evaluation values ​​under multiple shared operating scenarios based on the training sample set;

[0187] The hyperparameter optimization module 300 is used for the knowledge transfer mechanism based on the cloud side. It distributes the probabilistic proxy model and training sample set as prior knowledge to the edge side, and combines the Bayesian optimization algorithm and prior knowledge to optimize the hyperparameter combination and obtain the optimal tuning result of the hyperparameter combination.

[0188] In some embodiments, the data-driven control model is constructed on the cloud side. This system further includes a driver model construction module, used for:

[0189] Acquire historical control strategy data and multiple control effect measurement data uploaded from the edge side of the target transformer area of ​​the distribution network to the cloud side;

[0190] The control effect evaluation value is determined by weighting and summing multiple control effect measurement data according to the preset control target weighting factors.

[0191] Based on historical control strategy data and preset hyperparameter combinations, nonlinear predictions are made on the control effect evaluation value to obtain a data-driven control model.

[0192] In some embodiments, the driving model building module is configured to:

[0193] Obtain the initial dynamic mapping matrix, and perform linear prediction based on the changes in historical control strategy data and the initial dynamic mapping matrix to obtain the predicted value of the change in control effect;

[0194] Based on the control effect evaluation value, determine the actual value of the control effect change; based on the prediction error between the predicted value of the control effect change and the actual value of the control effect change, combine the preset dynamic mapping matrix update step size and the preset dynamic mapping matrix update penalty factor to correct the initial dynamic mapping matrix until the prediction error converges to the preset threshold, and obtain the corrected dynamic mapping matrix.

[0195] Based on the deviation between the control effect evaluation value and the preset control effect reference value, the control strategy data is updated by combining the corrected dynamic mapping matrix, the preset control strategy update step size and the preset control strategy penalty factor until the deviation reaches the preset deviation threshold, and the updated control strategy data is obtained.

[0196] The hyperparameter combination is constructed based on the control objective weight factor, the dynamic mapping matrix update step size, the dynamic mapping matrix update penalty factor, the control strategy update step size, and the control strategy penalty factor.

[0197] Based on the implicit coupling relationship between hyperparameter combinations, updated control strategy data, and control effect evaluation values, a nonlinear mapping data-driven control model is constructed.

[0198] In some embodiments, the training sample set includes hyperparameter combinations, scene labels, and control effect evaluation values;

[0199] The probability agent determination module 200 is used for:

[0200] A multilayer perceptron is used to extract features from hyperparameter combinations and scene labels. The total loss function is weighted and fused with error and regularization terms. The total loss function is minimized through backpropagation to obtain the latent feature representation.

[0201] Based on the latent feature representation, a probabilistic surrogate model between hyperparameter combinations and control effect evaluation values ​​is established using a Gaussian process.

[0202] In some embodiments, the hyperparameter optimization module 300 is used for:

[0203] Based on the training sample set, and combining scene labels with a probabilistic proxy model, the modeling error of candidate scenes under each scene label is determined.

[0204] Based on the training sample set and the probabilistic surrogate model, determine the regret reduction rate of the predicted control effect between adjacent iterations;

[0205] Based on the regret reduction rate and modeling error, the scenario with the largest information gain is selected from the candidate scenarios or target scenarios as the optimization scenario for this time;

[0206] Candidate data samples for this optimization scenario are generated using Monte Carlo simulation. The optimal hyperparameter combination that satisfies the performance benchmark constraints is selected from all candidate data samples and used as the optimal hyperparameter combination for this optimization.

[0207] The training sample set is updated based on the candidate data samples from this optimization, resulting in the updated training sample set.

[0208] Based on the updated training sample set, the modeling error of candidate scenes under each scene label is determined by combining the scene label and the probabilistic proxy model according to the training sample set until the preset iteration stopping condition is met. The optimal hyperparameter combination obtained in the last iteration is determined as the optimal tuning result of the hyperparameter combination.

[0209] In some embodiments, the hyperparameter optimization module 300 is used for:

[0210] Based on the candidate data samples for this optimization and each data sample in the training sample set, determine the sample similarity between the candidate data samples for this optimization and each data sample.

[0211] Remove data samples from the training sample set whose similarity exceeds a preset tolerance threshold, and add candidate data samples for this optimization to obtain an updated training sample set.

[0212] like Figure 11As shown, this application provides an electronic device 10, which includes a memory 20 and a processor 30. The memory 20 stores a computer program. When the computer program is executed by the processor 30, the processor 30 performs the following steps:

[0213] Historical control strategy data of the distribution network under multiple operating scenarios are acquired; on the cloud side of the distribution network, the historical control strategy data and the preset data-driven control model are combined to obtain the control effect evaluation value corresponding to the hyperparameter combination under each operating scenario, and a training sample set is constructed; among them, the data-driven control model is used to characterize the nonlinear mapping relationship between the control strategy data and the control effect evaluation value of the distribution network under the hyperparameter combination.

[0214] Based on the training sample set, a probabilistic proxy model is constructed between hyperparameter combinations and control effect evaluation values ​​under multiple shared operating scenarios.

[0215] Based on the knowledge transfer mechanism of the cloud side, the probabilistic proxy model and training sample set are distributed to the edge side as prior knowledge. Combined with the Bayesian optimization algorithm and prior knowledge, the hyperparameter combination is optimized to obtain the optimal tuning result of the hyperparameter combination.

[0216] In some embodiments, the process of building a data-driven control model on the cloud side includes:

[0217] Acquire historical control strategy data and multiple control effect measurement data uploaded from the edge side of the target transformer area of ​​the distribution network to the cloud side;

[0218] The control effect evaluation value is determined by weighting and summing multiple control effect measurement data according to the preset control target weighting factors.

[0219] Based on historical control strategy data and preset hyperparameter combinations, nonlinear predictions are made on the control effect evaluation value to obtain a data-driven control model.

[0220] In some embodiments, when a computer program is executed by processor 30, causing processor 30 to perform the following steps:

[0221] Obtain the initial dynamic mapping matrix, and perform linear prediction based on the changes in historical control strategy data and the initial dynamic mapping matrix to obtain the predicted value of the change in control effect;

[0222] Based on the control effect evaluation value, determine the actual value of the control effect change; based on the prediction error between the predicted value of the control effect change and the actual value of the control effect change, combine the preset dynamic mapping matrix update step size and the preset dynamic mapping matrix update penalty factor to correct the initial dynamic mapping matrix until the prediction error converges to the preset threshold, and obtain the corrected dynamic mapping matrix.

[0223] Based on the deviation between the control effect evaluation value and the preset control effect reference value, the control strategy data is updated by combining the corrected dynamic mapping matrix, the preset control strategy update step size and the preset control strategy penalty factor until the deviation reaches the preset deviation threshold, and the updated control strategy data is obtained.

[0224] The hyperparameter combination is constructed based on the control objective weight factor, the dynamic mapping matrix update step size, the dynamic mapping matrix update penalty factor, the control strategy update step size, and the control strategy penalty factor.

[0225] Based on the implicit coupling relationship between hyperparameter combinations, updated control strategy data, and control effect evaluation values, a nonlinear mapping data-driven control model is constructed.

[0226] In some embodiments, the training sample set includes hyperparameter combinations, scene labels, and control effect evaluation values;

[0227] When a computer program is executed by processor 30, the processor 30 performs the following steps:

[0228] A multilayer perceptron is used to extract features from hyperparameter combinations and scene labels. The total loss function is weighted and fused with error and regularization terms. The total loss function is minimized through backpropagation to obtain the latent feature representation.

[0229] Based on the latent feature representation, a probabilistic surrogate model between hyperparameter combinations and control effect evaluation values ​​is established using a Gaussian process.

[0230] In some embodiments, when a computer program is executed by processor 30, causing processor 30 to perform the following steps:

[0231] Based on the training sample set, and combining scene labels with a probabilistic proxy model, the modeling error of candidate scenes under each scene label is determined.

[0232] Based on the training sample set and the probabilistic surrogate model, determine the regret reduction rate of the predicted control effect between adjacent iterations;

[0233] Based on the regret reduction rate and modeling error, the scenario with the largest information gain is selected from the candidate scenarios or target scenarios as the optimization scenario for this time;

[0234] Candidate data samples for this optimization scenario are generated using Monte Carlo simulation. The optimal hyperparameter combination that satisfies the performance benchmark constraints is selected from all candidate data samples and used as the optimal hyperparameter combination for this optimization.

[0235] The training sample set is updated based on the candidate data samples from this optimization, resulting in the updated training sample set.

[0236] Based on the updated training sample set, the modeling error of candidate scenes under each scene label is determined by combining the scene label and the probabilistic proxy model according to the training sample set until the preset iteration stopping condition is met. The optimal hyperparameter combination obtained in the last iteration is determined as the optimal tuning result of the hyperparameter combination.

[0237] In some embodiments, when a computer program is executed by processor 30, causing processor 30 to perform the following steps:

[0238] Based on the candidate data samples for this optimization and each data sample in the training sample set, determine the sample similarity between the candidate data samples for this optimization and each data sample.

[0239] Remove data samples from the training sample set whose similarity exceeds a preset tolerance threshold, and add candidate data samples for this optimization to obtain an updated training sample set.

[0240] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed, performs the following steps:

[0241] Historical control strategy data of the distribution network under multiple operating scenarios are acquired; on the cloud side of the distribution network, the historical control strategy data and the preset data-driven control model are combined to obtain the control effect evaluation value corresponding to the hyperparameter combination under each operating scenario, and a training sample set is constructed; among them, the data-driven control model is used to characterize the nonlinear mapping relationship between the control strategy data and the control effect evaluation value of the distribution network under the hyperparameter combination.

[0242] Based on the training sample set, a probabilistic proxy model is constructed between hyperparameter combinations and control effect evaluation values ​​under multiple shared operating scenarios.

[0243] Based on the knowledge transfer mechanism of the cloud side, the probabilistic proxy model and training sample set are distributed to the edge side as prior knowledge. Combined with the Bayesian optimization algorithm and prior knowledge, the hyperparameter combination is optimized to obtain the optimal tuning result of the hyperparameter combination.

[0244] In some embodiments, the process of building a data-driven control model on the cloud side includes:

[0245] Acquire historical control strategy data and multiple control effect measurement data uploaded from the edge side of the target transformer area of ​​the distribution network to the cloud side;

[0246] The control effect evaluation value is determined by weighting and summing multiple control effect measurement data according to the preset control target weighting factors.

[0247] Based on historical control strategy data and preset hyperparameter combinations, nonlinear predictions are made on the control effect evaluation value to obtain a data-driven control model.

[0248] In some embodiments, when a computer program is executed, it also performs the following steps:

[0249] Obtain the initial dynamic mapping matrix, and perform linear prediction based on the changes in historical control strategy data and the initial dynamic mapping matrix to obtain the predicted value of the change in control effect;

[0250] Based on the control effect evaluation value, determine the actual value of the control effect change; based on the prediction error between the predicted value of the control effect change and the actual value of the control effect change, combine the preset dynamic mapping matrix update step size and the preset dynamic mapping matrix update penalty factor to correct the initial dynamic mapping matrix until the prediction error converges to the preset threshold, and obtain the corrected dynamic mapping matrix.

[0251] Based on the deviation between the control effect evaluation value and the preset control effect reference value, the control strategy data is updated by combining the corrected dynamic mapping matrix, the preset control strategy update step size and the preset control strategy penalty factor until the deviation reaches the preset deviation threshold, and the updated control strategy data is obtained.

[0252] The hyperparameter combination is constructed based on the control objective weight factor, the dynamic mapping matrix update step size, the dynamic mapping matrix update penalty factor, the control strategy update step size, and the control strategy penalty factor.

[0253] Based on the implicit coupling relationship between hyperparameter combinations, updated control strategy data, and control effect evaluation values, a nonlinear mapping data-driven control model is constructed.

[0254] In some embodiments, the training sample set includes hyperparameter combinations, scene labels, and control effect evaluation values;

[0255] In some embodiments, when a computer program is executed, it also performs the following steps:

[0256] A multilayer perceptron is used to extract features from hyperparameter combinations and scene labels. The total loss function is weighted and fused with error and regularization terms. The total loss function is minimized through backpropagation to obtain the latent feature representation.

[0257] Based on the latent feature representation, a probabilistic surrogate model between hyperparameter combinations and control effect evaluation values ​​is established using a Gaussian process.

[0258] In some embodiments, when a computer program is executed, it also performs the following steps:

[0259] Based on the training sample set, and combining scene labels with a probabilistic proxy model, the modeling error of candidate scenes under each scene label is determined.

[0260] Based on the training sample set and the probabilistic surrogate model, determine the regret reduction rate of the predicted control effect between adjacent iterations;

[0261] Based on the regret reduction rate and modeling error, the scenario with the largest information gain is selected from the candidate scenarios or target scenarios as the optimization scenario for this time;

[0262] Candidate data samples for this optimization scenario are generated using Monte Carlo simulation. The optimal hyperparameter combination that satisfies the performance benchmark constraints is selected from all candidate data samples and used as the optimal hyperparameter combination for this optimization.

[0263] The training sample set is updated based on the candidate data samples from this optimization, resulting in the updated training sample set.

[0264] Based on the updated training sample set, the modeling error of candidate scenes under each scene label is determined by combining the scene label and the probabilistic proxy model according to the training sample set until the preset iteration stopping condition is met. The optimal hyperparameter combination obtained in the last iteration is determined as the optimal tuning result of the hyperparameter combination.

[0265] In some embodiments, when a computer program is executed, it also performs the following steps:

[0266] Based on the candidate data samples for this optimization and each data sample in the training sample set, determine the sample similarity between the candidate data samples for this optimization and each data sample.

[0267] Remove data samples from the training sample set whose similarity exceeds a preset tolerance threshold, and add candidate data samples for this optimization to obtain an updated training sample set.

[0268] This application provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer performs the following steps:

[0269] Historical control strategy data of the distribution network under multiple operating scenarios are acquired; on the cloud side of the distribution network, the historical control strategy data and the preset data-driven control model are combined to obtain the control effect evaluation value corresponding to the hyperparameter combination under each operating scenario, and a training sample set is constructed; among them, the data-driven control model is used to characterize the nonlinear mapping relationship between the control strategy data and the control effect evaluation value of the distribution network under the hyperparameter combination.

[0270] Based on the training sample set, a probabilistic proxy model is constructed between hyperparameter combinations and control effect evaluation values ​​under multiple shared operating scenarios.

[0271] Based on the knowledge transfer mechanism of the cloud side, the probabilistic proxy model and training sample set are distributed to the edge side as prior knowledge. Combined with the Bayesian optimization algorithm and prior knowledge, the hyperparameter combination is optimized to obtain the optimal tuning result of the hyperparameter combination.

[0272] In some embodiments, the process of building a data-driven control model on the cloud side includes:

[0273] Acquire historical control strategy data and multiple control effect measurement data uploaded from the edge side of the target transformer area of ​​the distribution network to the cloud side;

[0274] The control effect evaluation value is determined by weighting and summing multiple control effect measurement data according to the preset control target weighting factors.

[0275] Based on historical control strategy data and preset hyperparameter combinations, nonlinear predictions are made on the control effect evaluation value to obtain a data-driven control model.

[0276] In some embodiments, when program instructions are executed by a computer, the computer also performs the following steps:

[0277] Obtain the initial dynamic mapping matrix, and perform linear prediction based on the changes in historical control strategy data and the initial dynamic mapping matrix to obtain the predicted value of the change in control effect;

[0278] Based on the control effect evaluation value, determine the actual value of the control effect change; based on the prediction error between the predicted value of the control effect change and the actual value of the control effect change, combine the preset dynamic mapping matrix update step size and the preset dynamic mapping matrix update penalty factor to correct the initial dynamic mapping matrix until the prediction error converges to the preset threshold, and obtain the corrected dynamic mapping matrix.

[0279] Based on the deviation between the control effect evaluation value and the preset control effect reference value, the control strategy data is updated by combining the corrected dynamic mapping matrix, the preset control strategy update step size and the preset control strategy penalty factor until the deviation reaches the preset deviation threshold, and the updated control strategy data is obtained.

[0280] The hyperparameter combination is constructed based on the control objective weight factor, the dynamic mapping matrix update step size, the dynamic mapping matrix update penalty factor, the control strategy update step size, and the control strategy penalty factor.

[0281] Based on the implicit coupling relationship between hyperparameter combinations, updated control strategy data, and control effect evaluation values, a nonlinear mapping data-driven control model is constructed.

[0282] In some embodiments, the training sample set includes hyperparameter combinations, scene labels, and control effect evaluation values;

[0283] When program instructions are executed by the computer, the computer also performs the following steps:

[0284] A multilayer perceptron is used to extract features from hyperparameter combinations and scene labels. The total loss function is weighted and fused with error and regularization terms. The total loss function is minimized through backpropagation to obtain the latent feature representation.

[0285] Based on the latent feature representation, a probabilistic surrogate model between hyperparameter combinations and control effect evaluation values ​​is established using a Gaussian process.

[0286] In some embodiments, when program instructions are executed by a computer, the computer also performs the following steps:

[0287] Based on the training sample set, and combining scene labels with a probabilistic proxy model, the modeling error of candidate scenes under each scene label is determined.

[0288] Based on the training sample set and the probabilistic surrogate model, determine the regret reduction rate of the predicted control effect between adjacent iterations;

[0289] Based on the regret reduction rate and modeling error, the scenario with the largest information gain is selected from the candidate scenarios or target scenarios as the optimization scenario for this time;

[0290] Candidate data samples for this optimization scenario are generated using Monte Carlo simulation. The optimal hyperparameter combination that satisfies the performance benchmark constraints is selected from all candidate data samples and used as the optimal hyperparameter combination for this optimization.

[0291] The training sample set is updated based on the candidate data samples from this optimization, resulting in the updated training sample set.

[0292] Based on the updated training sample set, the modeling error of candidate scenes under each scene label is determined by combining the scene label and the probabilistic proxy model according to the training sample set until the preset iteration stopping condition is met. The optimal hyperparameter combination obtained in the last iteration is determined as the optimal tuning result of the hyperparameter combination.

[0293] In some embodiments, when program instructions are executed by a computer, the computer also performs the following steps:

[0294] Based on the candidate data samples for this optimization and each data sample in the training sample set, determine the sample similarity between the candidate data samples for this optimization and each data sample.

[0295] Remove data samples from the training sample set whose similarity exceeds a preset tolerance threshold, and add candidate data samples for this optimization to obtain an updated training sample set.

[0296] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, electronic devices, computer storage media, and computer program products described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0297] It should be noted that the terms "comprising" and "having" and any variations thereof in the specification, claims and accompanying drawings of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such processes, methods, products or devices.

[0298] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0299] In the several embodiments provided by this invention, it should be understood that the disclosed systems, electronic devices, computer storage media, computer program products, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.

[0300] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0301] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0302] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for executing all or part of the steps of the methods described in the various embodiments of the present invention through a computer device (which may be a personal computer, a server, or a network device, etc.). The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.

[0303] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data-driven feedback control hyperparameter tuning method, characterized in that, include: Historical control strategy data of the distribution network under multiple operating scenarios are acquired; on the cloud side of the distribution network, the historical control strategy data and a preset data-driven control model are combined to obtain the control effect evaluation value corresponding to the hyperparameter combination under each operating scenario, and a training sample set is constructed; wherein, the data-driven control model is used to characterize the nonlinear mapping relationship between the control strategy data and the control effect evaluation value of the distribution network under the hyperparameter combination. Based on the training sample set, a probabilistic proxy model is constructed between the hyperparameter combination and the control effect evaluation value under multiple shared operating scenarios. Based on the knowledge transfer mechanism of the cloud side, the probabilistic proxy model and the training sample set are distributed to the edge side as prior knowledge. The hyperparameter combination is optimized by combining the Bayesian optimization algorithm and the prior knowledge to obtain the optimal tuning result of the hyperparameter combination.

2. The data-driven feedback control hyperparameter tuning method according to claim 1, characterized in that, The process of constructing the data-driven control model on the cloud side includes: Acquire historical control strategy data and multiple control effect measurement data uploaded from the edge side of the target transformer area of ​​the power distribution network to the cloud side; The control effect evaluation value is determined by weighting and summing multiple control effect measurement data according to a preset control target weighting factor. Based on the historical control strategy data and the preset hyperparameter combination, the control effect evaluation value is nonlinearly predicted to obtain the data-driven control model.

3. The data-driven feedback control hyperparameter tuning method according to claim 2, characterized in that, The step of performing nonlinear prediction on the control effect evaluation value based on the historical control strategy data and a preset hyperparameter combination to obtain the data-driven control model includes: An initial dynamic mapping matrix is ​​obtained, and linear prediction is performed based on the change values ​​of the historical control strategy data and the initial dynamic mapping matrix to obtain the predicted value of the change in control effect. Based on the control effect evaluation value, determine the true value of the control effect change; based on the prediction error between the predicted value of the control effect change and the true value of the control effect change, combine the preset dynamic mapping matrix update step size and the preset dynamic mapping matrix update penalty factor to correct the initial dynamic mapping matrix until the prediction error converges to a preset threshold, and obtain the corrected dynamic mapping matrix. Based on the deviation between the control effect evaluation value and the preset control effect reference value, the control strategy data is updated by combining the corrected dynamic mapping matrix, the preset control strategy update step size and the preset control strategy penalty factor, until the deviation reaches the preset deviation threshold, and the updated control strategy data is obtained. The hyperparameter combination is formed based on the control target weight factor, the dynamic mapping matrix update step size, the dynamic mapping matrix update penalty factor, the control strategy update step size, and the control strategy penalty factor; Based on the implicit coupling relationship between the hyperparameter combination, the updated control strategy data, and the control effect evaluation value, a nonlinear mapping data-driven control model is constructed.

4. The data-driven feedback control hyperparameter tuning method according to claim 1, characterized in that, The training sample set includes hyperparameter combinations, scene labels, and control effect evaluation values; The construction of a probabilistic proxy model based on the training sample set, relating the hyperparameter combination to the control effect evaluation value under multiple shared operating scenarios, includes: The hyperparameter combination and the scene label are extracted using a multilayer perceptron. The total loss function is weighted and fused with error term and regularization term. The total loss function is minimized through backpropagation to obtain the latent feature representation. Based on the latent feature representation, a probabilistic surrogate model between the hyperparameter combination and the control effect evaluation value is established using a Gaussian process.

5. The data-driven feedback control hyperparameter tuning method according to claim 1, characterized in that, Combining the Bayesian optimization algorithm and the prior knowledge, the hyperparameter combination is optimized to obtain the optimal tuning result of the hyperparameter combination, including: Based on the training sample set, and combining the scene labels with the probabilistic proxy model, the modeling error of the candidate scenes under each scene label is determined; Based on the training sample set and the probabilistic proxy model, determine the regret reduction rate of the predicted control effect between adjacent iterations; Based on the regret reduction rate and the modeling error, the scenario with the largest information gain is selected from the candidate scenarios or the target scenarios as the optimization scenario for this time; Candidate data samples for the current optimization scenario are generated using Monte Carlo methods. The optimal hyperparameter combination that satisfies the performance benchmark constraints is selected from all the candidate data samples and used as the optimal hyperparameter combination for this optimization. The training sample set is updated based on the candidate data samples used in this optimization to obtain the updated training sample set; Based on the updated training sample set, the process of determining the modeling error of candidate scenes under each scene label by combining the scene label and the probabilistic proxy model according to the training sample set is repeated until the preset iteration stopping condition is met, and the optimal hyperparameter combination obtained in the last iteration is determined as the optimal tuning result of the hyperparameter combination.

6. The data-driven feedback control hyperparameter tuning method according to claim 5, characterized in that, The step of updating the training sample set based on the candidate data samples selected in this optimization to obtain the updated training sample set includes: Based on the candidate data samples for this optimization and each data sample in the training sample set, determine the sample similarity between the candidate data samples for this optimization and each of the data samples. The updated training sample set is obtained by removing data samples whose similarity is higher than a preset tolerance threshold from the training sample set and adding the candidate data samples selected for this optimization.

7. A data-driven feedback control hyperparameter tuning system, characterized in that, include: A sample set construction module is used to acquire historical control strategy data of the distribution network under multiple operating scenarios; on the cloud side of the distribution network, the historical control strategy data is combined with a preset data-driven control model to obtain the control effect evaluation value corresponding to the hyperparameter combination under each operating scenario, and a training sample set is constructed; wherein, the data-driven control model is used to characterize the nonlinear mapping relationship between the control strategy data and the control effect evaluation value of the distribution network under the hyperparameter combination. The probabilistic proxy determination module is used to construct a probabilistic proxy model between the hyperparameter combination and the control effect evaluation value under multiple shared operating scenarios based on the training sample set. The hyperparameter optimization module is used to distribute the probabilistic proxy model and the training sample set as prior knowledge to the edge side based on the cloud-side knowledge transfer mechanism, and combine the Bayesian optimization algorithm and the prior knowledge to optimize the hyperparameter combination and obtain the optimal tuning result of the hyperparameter combination.

8. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the data-driven feedback control hyperparameter tuning method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the steps of the data-driven feedback control hyperparameter tuning method as described in any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, wherein when the program instructions are executed by a computer, the computer performs the steps of the data-driven feedback control hyperparameter tuning method as described in any one of claims 1-6.