Offshore wind plant DC current collection system operation performance comprehensive evaluation weight adaptive adjustment method and system based on reinforcement learning

By using a reinforcement learning-based approach, the weights of the DC power collection system in offshore wind farms are adaptively adjusted, solving the problem of statically fixed weights in existing technologies. This enables more accurate comprehensive evaluation and optimized decision support, thereby improving the operational efficiency and safety of offshore wind farms.

CN121998624APending Publication Date: 2026-05-08TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2026-02-03
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In existing technologies, the weights in the comprehensive evaluation methods for DC collection systems of offshore wind farms are statically fixed, which cannot adapt to complex and ever-changing operating conditions and equipment health status, resulting in distorted evaluation results and poor decision-making guidance.

Method used

We employ a reinforcement learning-based approach to construct a Markov decision process by collecting and preprocessing multi-source operational data. By leveraging the interaction between the agent and the Markov decision process environment, we achieve adaptive weight adjustment. Furthermore, we optimize the weight adjustment strategy by combining the Actor-Critic architecture and the improved PPO algorithm.

Benefits of technology

It achieves adaptive and intelligent adjustment of weights, improves the accuracy of comprehensive evaluation and the effectiveness of decision support, and enhances the operational efficiency and safety of offshore wind farms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998624A_ABST
    Figure CN121998624A_ABST
Patent Text Reader

Abstract

The invention specifically provides an offshore wind plant DC current collection system operation performance comprehensive evaluation weight adaptive adjustment method and system based on reinforcement learning, and the method comprises the steps: collecting and preprocessing multi-source operation data of an offshore wind plant DC current collection system, and obtaining a real-time performance index; based on the real-time performance index and the current evaluation weight vector, constructing a Markov decision process for weight adjustment; the intelligent agent based on reinforcement learning driving interacts with the constructed Markov decision process environment to obtain an optimal weight adjustment strategy; and based on the optimal weight adjustment strategy, obtaining an optimal weight vector, and based on the optimal weight vector, carrying out real-time comprehensive evaluation on the operation performance of the DC current collection system of the offshore wind plant. The invention provides an innovative technical approach for solving the problem of dynamic evaluation of a complex industrial system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent operation and maintenance and dynamic decision-making technology of offshore wind power DC collection systems, specifically involving a method and system for adaptive adjustment of weights for comprehensive evaluation of the operation performance of offshore wind farm DC collection systems based on reinforcement learning. Background Technology

[0002] With the growing global demand for clean energy, offshore wind power, as an important form of renewable energy, is developing rapidly. Offshore wind farms typically use DC collection systems to gather the electricity generated by numerous wind turbines and transmit it over long distances. The safe and stable operation of this system is crucial to the overall performance and power supply reliability of the wind farm.

[0003] To comprehensively evaluate the operating status of DC collector systems, the industry typically employs a comprehensive evaluation method. This method is achieved by constructing a comprehensive evaluation system covering multiple dimensions such as stability, reliability, and energy efficiency. Specifically, a series of key performance indicators need to be selected for each dimension, and a weighted summation model needs to be constructed. The most common model is the weighted summation method, i.e.: Comprehensive Score = w1Indicator1 + w2Indicator2 + ... + w n *Indices n, where w1, w2, ..., w n The weights of each indicator.

[0004] However, existing comprehensive evaluation methods have the following significant drawbacks:

[0005] 1. Static fixed weights: Once set, the weights remain unchanged for a long period of time. These weights are usually determined once during the design phase using subjective or semi-subjective methods based on expert experience (such as the Analytic Hierarchy Process, AHP).

[0006] 2. Inability to adapt to time-varying operating conditions: The operating environment of offshore wind farms is highly dynamic and uncertain. Natural conditions such as wind speed and sea state change drastically, and equipment gradually ages or experiences random failures over time. Static weights cannot be adjusted according to the current actual operating state of the system and the external environment, leading to distorted evaluation results. For example, when equipment is severely aged, the weight of reliability indicators should be increased accordingly, but the static weight model cannot reflect this change.

[0007] 3. Delayed evaluation results and poor decision-making guidance: Because the weights cannot be adaptively adjusted, the evaluation results often fail to accurately and timely reflect the system's shortcomings and potential risks. This makes it impossible for operation and maintenance personnel to make optimal control strategies or maintenance decisions based on the evaluation results, which may lead to missing the best intervention opportunity or even causing safety accidents.

[0008] A series of national and industry standards have been established in the field of power system and wind farm evaluation. For example, GB / T 2900.13-2008 "Electrical Engineering Terminology: Reliability and Service Quality" provides reliability terms such as Expected Energy Deficiency (EENS); DL / T 793.1-2017 "Reliability Evaluation Code for Power Generation Equipment Part 1: General Rules" specifies the classification of power generation equipment into operating, available, and unavailable states, and the statistical methods for reliability indicators such as availability factor and outage factor; DL / T686-2018 "Guidelines for Calculating Power Losses in Power Grids" and GB / T 40267-2021 "Guidelines for Calculating Power Losses in Power Systems" unify the calculation methods for transmission line losses and line loss rates; and NB / T 31117-2017, NB / T 11603-2024, and NB / T11599-2024... Engineering criteria were given for submarine cable current carrying capacity, thermal stability limit, and voltage deviation and recovery time of offshore converter stations. However, the above standards mainly serve static planning and design and periodic evaluation. Existing work generally does not combine these standardized indicators with the dynamic adjustment mechanism of online multidimensional performance evaluation weights, and still mainly relies on fixed weight models, which makes it difficult to achieve adaptive optimization of comprehensive evaluation under complex time-varying operating conditions.

[0009] Therefore, there is an urgent need for a technical solution that can overcome the shortcomings of static weights and enable the evaluation weights to be adaptively adjusted according to real-time operating conditions, so as to improve the accuracy of comprehensive evaluation and the effectiveness of decision support. Summary of the Invention

[0010] The present invention aims to solve the technical problem that in the prior art, when comprehensively evaluating the DC collection system of offshore wind farms, the weights of the evaluation indicators are mostly statically set, which cannot adapt to the complex and ever-changing operating conditions and equipment health status, resulting in distorted evaluation results and poor decision-making guidance.

[0011] To achieve the above objectives, the present invention provides the following solution:

[0012] An adaptive adjustment method for weights in the comprehensive performance evaluation of DC collector systems for offshore wind farms based on reinforcement learning includes:

[0013] Collect and preprocess multi-source operation data of the DC collection system of offshore wind farm to obtain real-time performance indicators;

[0014] Based on the real-time performance metrics and the current evaluation weight vector, a Markov decision process for weight adjustment is constructed.

[0015] The agent, driven by reinforcement learning, interacts with the constructed Markov decision process environment to obtain the optimal weight adjustment strategy.

[0016] Based on the optimal weight adjustment strategy, the optimal weight vector is obtained, and the operating performance of the offshore wind farm DC collection system is evaluated in real time based on the optimal weight vector.

[0017] Preferably, the multi-source operational data includes stability data, reliability data, and energy efficiency data; the real-time performance indicators include aggregated stability indicators, aggregated reliability indicators, and aggregated energy efficiency indicators.

[0018] The stability data includes DC bus voltage fluctuation rate, converter station current fluctuation rate, and system frequency deviation.

[0019] The reliability data includes the health index of key equipment, the predicted mean time between failures, and the equipment utilization rate.

[0020] The energy efficiency data includes grid loss rate and transmission efficiency, which are used to characterize the loss level and energy utilization efficiency of DC power collection systems during energy transmission.

[0021] Preferably, the Markov decision process includes: a state space, an action space, and a reward function;

[0022] The elements of the state space include the current evaluation weight vector and the real-time performance index vector;

[0023] The elements of the action space are the adjustment actions of each weight parameter in the current evaluation weight vector using a preset adjustment step size;

[0024] The reward function adopts a multi-objective composite form, including a first sub-reward term, a second sub-reward term, a third sub-reward term, and a stability penalty term.

[0025] Preferably, in the reward function,

[0026] The first sub-reward item is the relative improvement rate of the voltage stability index between two adjacent time points;

[0027] The second sub-reward item is the relative improvement rate of the health index of key equipment between two adjacent time points;

[0028] The third sub-reward item is the relative improvement rate of the energy efficiency aggregate index between two adjacent time points;

[0029] The stability penalty term is determined by the Euclidean distance between the evaluation weight vector at the current moment and the evaluation weight vector at the previous moment.

[0030] The value of the reward function is obtained by weighting and summing the first sub-reward item, the second sub-reward item, and the third sub-reward item according to preset first, second, and third weight coefficients, and then subtracting the stability penalty item.

[0031] Preferably, the agent adopts an Actor-Critic architecture; both the Actor policy network and the Critic value network include an input layer, a first hidden layer, a second hidden layer, and an output layer;

[0032] Actor networks are used to provide the probability distribution of each action in the current state and to adjust the decision evaluation weights.

[0033] Critic networks are used to calculate state value estimates and obtain the advantage function.

[0034] This invention also provides a weighted adaptive adjustment system for the comprehensive evaluation of the operating performance of offshore wind farm DC collector systems based on reinforcement learning, used to implement the method, comprising:

[0035] The data acquisition and preprocessing module is used to acquire and preprocess multi-source operating data of the DC collection system of offshore wind farms to obtain real-time performance indicators.

[0036] The reinforcement learning intelligent decision-making module is used to construct a Markov decision process for weight adjustment based on the real-time performance indicators and the current evaluation weight vector.

[0037] The weight dynamic execution module is used to enable the reinforcement learning-driven agent to interact with the constructed Markov decision process environment to obtain the optimal weight adjustment strategy.

[0038] The comprehensive evaluation output module is used to obtain the optimal weight vector based on the optimal weight adjustment strategy, and to perform a real-time comprehensive evaluation of the operating performance of the offshore wind farm DC collection system based on the optimal weight vector.

[0039] Preferably, the data acquisition and preprocessing module establishes a connection with the SCADA system and CMS system of the offshore wind farm DC power collection system through an industrial communication protocol to realize the real-time acquisition of multi-source operating data.

[0040] Preferably, the comprehensive evaluation output module also provides a Web service interface and a visualization interface to support real-time display of evaluation results and historical data query.

[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0042] (1) The adaptive and intelligent adjustment of weights is realized: This invention overcomes the defects of subjective and fixed weight setting in traditional evaluation methods. By utilizing the self-exploration and learning ability of reinforcement learning, the weights can be automatically adjusted according to real-time working conditions without manual intervention.

[0043] (2) Improved the accuracy and objectivity of the comprehensive evaluation: Since the weights are data-driven and dynamically generated, they can more accurately reflect the system's operational priorities and potential risks at a specific moment, making the comprehensive evaluation results closer to the system's true state.

[0044] (3) Enhanced effectiveness of decision support: Accurate real-time evaluation results provide reliable data support for the optimization control, predictive maintenance and operation and maintenance decisions of wind farms, which helps to improve the long-term operating efficiency and resource utilization of power generation systems while ensuring safety.

[0045] (4) Robust and widely applicable: The framework of this invention has good versatility. By adjusting the design of the state space, action space and reward function, it can be applied to DC collection systems of offshore wind farms of different types and scales. Attached Figure Description

[0046] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 This is a structural block diagram of the weight adaptive adjustment system according to an embodiment of the present invention;

[0048] Figure 2 This is a flowchart of the weight adaptive adjustment method according to an embodiment of the present invention;

[0049] Figure 3 This is a schematic diagram of the Markov decision process construction process according to an embodiment of the present invention;

[0050] Figure 4 This is a block diagram of the Actor-Critic network structure according to an embodiment of the present invention;

[0051] Figure 5 This is a schematic diagram of the improved PPO training process according to an embodiment of the present invention;

[0052] Figure 6 This is a flowchart illustrating the online dynamic weight adjustment and comprehensive evaluation process in an embodiment of the present invention.

[0053] Figure 7 The following is a comparison chart of the time series of key indicators between the traditional fixed weight method and the adaptive weight method of the present invention in the embodiments of the present invention; wherein, (a) is a comparison chart of voltage stability indicators; (b) is a comparison chart of equipment health index; and (c) is a comparison chart of energy efficiency indicators.

[0054] Figure 8This is an overview of the performance improvement of the present invention compared to the traditional fixed-weight method in the embodiments of the present invention;

[0055] Figure 9 This is a radar chart showing the multi-dimensional comprehensive performance of the traditional method compared to the present invention in an embodiment of the present invention. Detailed Implementation

[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0058] The national and industry standards mentioned in the embodiments include, but are not limited to:

[0059] 1) GB / T 2900.13-2008 "Electrical Engineering Terminology: Reliability and Quality of Service", used to define reliability terms such as Expected Energy Deficiency (EENS);

[0060] 2) DL / T 793.1-2017 "Reliability Evaluation Code for Power Generation Equipment Part 1: General Rules" is used to classify equipment into use, available, and unavailable states and to count downtime and power loss. Based on this, the present invention expands the evaluation object to key components such as wind turbines, DC submarine cables, transformers, and converter stations.

[0061] 3) DL / T 686-2018 and GB / T 40267-2021 are used for calculating energy efficiency indicators such as line I²R loss, resistance temperature correction and line loss rate;

[0062] 4) GB 38755-2019, used to determine the criteria for N-1 safety, voltage stability and transient recovery time;

[0063] 5) NB / T 31117-2017 and NB / T 11603-2024 are used for the verification of the current carrying capacity and thermal stability limit of submarine cables;

[0064] 6) NB / T 11599-2024, Engineering criteria for stability indicators such as DC bus voltage deviation, voltage stability margin and voltage recovery time after fault in offshore converter stations.

[0065] By introducing the above standards, this invention ensures that all performance metrics used in reinforcement learning models have unified engineering significance and comparability.

[0066] Example 1:

[0067] like Figure 2 As shown, a reinforcement learning-based adaptive adjustment method for weights in the comprehensive performance evaluation of offshore wind farm DC collector systems includes:

[0068] S1: Collect and preprocess multi-source operational data of the offshore wind farm's DC power collection system to obtain real-time performance indicators. A further implementation method involves including stability data, reliability data, and energy efficiency data in the multi-source operational data; and real-time performance indicators including aggregated stability indicators, aggregated reliability indicators, and aggregated energy efficiency indicators.

[0069] Stability data include DC bus voltage fluctuation (calculated as standard deviation over 1 minute), converter station current fluctuation (peak to mean deviation), and system frequency deviation (relative to rated frequency).

[0070] Reliability data includes the health index of key equipment (HI, calculated based on the fusion of data from multiple sensors such as temperature and vibration), the predicted mean time between failures (MTBF) (based on a reliability model of Weibull distribution), and equipment utilization (the ratio of actual operating time to statistical period).

[0071] Energy efficiency data include grid loss rate (based on power flow calculation) and transmission efficiency (the ratio of output power to input power), which are used to characterize the loss level and energy utilization efficiency of DC collector systems during energy transmission.

[0072] The data preprocessing process includes:

[0073] ① Data cleaning: Sliding window detection method is used to identify abnormal data, and time series prediction model is used for data repair.

[0074] ② Data alignment: Timestamp synchronization is performed on heterogeneous data sources with a sampling period of 1 minute.

[0075] ③ Normalization: The min-max standardization method is used to map each index to the [0,1] interval to eliminate the influence of dimensions.

[0076] The key features obtained after preprocessing include:

[0077] Stability-related raw data: DC bus voltage sequence Vdc(i,t), current fluctuation rate, frequency deviation, etc.

[0078] Reliability-related raw data: basic quantities of equipment health index (temperature, vibration, insulation parameters, etc.), predicted mean time between failures, equipment utilization rate, etc.

[0079] Based on this, this embodiment first performs feature aggregation on the above-mentioned original data to obtain two aggregation performance indicators: stability aggregation index p. stab (t):

[0080] For example, it can be constructed primarily using the voltage stability index (VSI), supplemented by factors such as frequency deviation:

[0081] (1)

[0082] in Let i represent the set of lines connected to node i. The coupling coefficient is calculated based on the DC network parameters. This represents the magnitude of the DC bus voltage at node j, which is connected to node i. This represents the magnitude of the DC bus voltage at node i. Represents the set of nodes (or lines) connected to node i.

[0083] Take the worst value from the entire network:

[0084] (2)

[0085] To unify the direction, it is linearly mapped to the [0,1] interval and combined with frequency deviation Δf, for example:

[0086] (3)

[0087] Where ⋅^⋅ represents the Min–Max normalization result, + =1. , This represents the weighting coefficient.

[0088] Stability Polymerization Performance Indicators The voltage stability index (VSI) and its related sub-indices are calculated based at least on the voltage stability requirements in GB 38755 2019 "Guidelines for the Safety and Stability of Power Systems", NB / T 11599 2024 "Design Code for Offshore Converter Stations in Wind Farm Projects", NB / T31117 2017 "Technical Guidelines for the Selection and Laying of AC Submarine Cables for Offshore Wind Farms", and NB / T 11603 2024 "Technical Guidelines for the Selection and Laying of DC Submarine Cables for Offshore Wind Farms".

[0089] Reliability aggregate index prel (t) Further, at least in part, the expected unsupplied energy (EENS) is calculated based on the definition of term 191-30-01 "Expected energy notsupplied; expected unsupplied energy (EENS / EUE)" in GB / T 2900.13-2008 "Electrical Engineering Terminology Reliability and Service Quality". The statistical scope of the EENS is preferably based on the provisions of DL / T 793.1-2017 "Power Generation Equipment Reliability Evaluation Code Part 1: General Rules" regarding the classification of use, availability, and unavailability states, as well as the statistics of downtime and power loss. Fault and downtime data are statistically analyzed for the components constituting the DC collection system, such as wind turbines, transformers, submarine cables, and converter stations.

[0090] Reliability aggregate performance index Build as follows:

[0091]

[0092] in For reliability aggregation performance indicators, f(⋅) is a preset normalization function, whose independent variables include at least HI, EENS, AF, and UOF, and can further include other reliability indicators such as Mean Time To Repair (MTTR) and Unplanned Outage Rate (UOR) as needed. HI is the equipment health index, and AF and UOF are the availability coefficient and unplanned outage coefficient respectively, calculated according to DL / T 793.1-2017.

[0093] With the Health Index (HI) of critical equipment as the core, HI can be calculated using existing health assessment methods:

[0094] (4)

[0095] in Let m be the m-th normalized health characteristic (temperature, vibration, etc.) of device k. As weight.

[0096] Take the weighted average of the HI values ​​of all key devices:

[0097] (5)

[0098] This is a coefficient representing the degree of influence of the k-th device on the overall health of the system. Typically, in the calculation, the sum of the weights of all devices is 1 (i.e., ∑...). =1).

[0099] Similarly, after normalization, we obtain... It can also be combined with Mean Time Between Failures (MTBF), etc.

[0100] (6)

[0101] in + =1. , This represents the weighting coefficient.

[0102] Energy efficiency aggregate index :

[0103] It is constructed based on the system network loss rate and transmission efficiency. Let the average network loss rate over a certain statistical period be Loss. rate (t), transmission efficiency is η trans (t), after being normalized to the interval [0,1], is denoted as and Then we can define:

[0104] (7)

[0105] in, + = 1, , These are the weighting coefficients for each sub-indicator under the energy efficiency dimension. The above definition guarantees... A higher value indicates lower energy loss and higher transmission efficiency. The calculation of network loss rate and transmission efficiency should preferably refer to the I²R loss and line loss rate calculation methods in DL / T 686-2018 "Guidelines for Calculation of Power Grid Energy Loss" and GB / T 40267-2021 "Guidelines for Calculation of Power System Energy Loss", and be combined with the submarine cable resistance and temperature correction methods based on IEC 60287 in NB / T 31117-2017 and NB / T 11603-2024.

[0106] Network loss rate is calculated based on the amount of electricity lost along the line. Electricity transmitted in the same period Calculation of the ratio:

[0107] (8)

[0108] (9)

[0109] Among them, line resistance The 20 ℃ resistor is temperature corrected according to the IEC 60287 series standards or equivalent domestic standards.

[0110] S2: Based on real-time performance metrics and the current evaluation weight vector, construct a Markov decision process for weight adjustment. For example... Figure 3 As shown, a further implementation method is that the Markov decision process includes: a state space, an action space, and a reward function;

[0111] The elements of the state space include the current evaluation weight vector and the real-time performance index vector; specifically, at any decision time t, the evaluation weight vector is constructed as follows:

[0112] (10)

[0113] Real-time performance metric vector:

[0114] (11)

[0115] Among them, the stability polymerization performance index It is calculated at least based on the Voltage Stability Index (VSI), which represents the voltage stability of a DC collector system. Reliability aggregate performance index It is calculated based at least on the Equipment Health Index (HI), which represents the health status of critical equipment. Energy efficiency aggregate performance index. It is calculated based on at least one of the following operational efficiency indicators that represent the system's energy utilization efficiency: network loss rate and / or transmission efficiency.

[0116] The state vector is defined as:

[0117] (12)

[0118] in: This represents the weight vector at time t. This represents the performance index vector at time t.

[0119] Performance metrics are extracted from raw data through feature engineering, such as It integrates multiple sub-indicators such as voltage stability and frequency stability.

[0120] The elements of the action space are the adjustment actions performed on each weight parameter in the current evaluation weight vector using a preset adjustment step size; specifically, when performing a certain adjustment action in the action space, the current evaluation weight vector is adjusted. The weight parameters are increased or decreased accordingly to obtain the intermediate evaluation weight vector. The intermediate evaluation weight vector is then normalized using the softmax function. Normalization is performed to obtain the evaluation weight vector for the next time step. The softmax normalization function is configured to make... All weight parameters are positive and their sum is 1, thus suppressing drastic fluctuations in the weight vector; Action space design:

[0121] ①. Definition of Action Vector

[0122] Using a three-dimensional discrete action space, the two weights are increased or decreased in small increments:

[0123] (13)

[0124] in =0.01 is the adjustment step size. After the action is executed, the weights are normalized using the softmax function.

[0125] (14)

[0126] ②. Weight Update and Softmax Normalization

[0127] For the current weight vector Perform an action The intermediate weights are obtained as follows:

[0128] (15)

[0129] Then normalization is performed using the softmax function:

[0130] (16)

[0131] The intermediate weight vector representing time t The component value on the i-th evaluation dimension, The intermediate weight vector representing time t The component value on the j-th evaluation dimension.

[0132] Let the vector form be:

[0133] (17)

[0134] This process ensures that each component is positive and the sum of the weights is 1, while also providing a smoothing and suppression effect for large step length movements.

[0135] reward function It adopts a multi-objective composite form, including the first sub-reward term ΔS, the second sub-reward term ΔR, the third sub-reward term ΔE, and the stability penalty term D.

[0136] A further implementation method is that, in the reward function,

[0137] The first sub-reward item is the relative improvement rate of the voltage stability index between two adjacent time points;

[0138] The second sub-reward item is the relative improvement rate of the health index of key equipment between two adjacent time points;

[0139] The third sub-reward item is the relative improvement rate of the energy efficiency aggregate index between two adjacent time points;

[0140] The stability penalty term is determined by the evaluation weight vector at the current time step. The evaluation weight vector from the previous time step The Euclidean distance between them is determined;

[0141] The value of the reward function is obtained by weighting and summing the first, second, and third sub-reward items according to preset first, second, and third weight coefficients, and then subtracting the stability penalty term.

[0142] Specifically, the reward function design:

[0143] ①. Definition of Sub-reward Items

[0144] Stability improvement rate:

[0145] (18)

[0146] When VSI approaches a better state, ΔS t >0. Reliability improvement rate:

[0147] (19)

[0148] ②. Penalty for weight changes

[0149] (20)

[0150] ③. Energy efficiency improvement rate

[0151] Let the aggregate energy efficiency indices for the previous cycle and the current cycle be respectively , Then the definition is:

[0152] (twenty one)

[0153] in To prevent extremely small positive numbers with a denominator of 0; when the system network loss rate decreases and transmission efficiency increases, Increase, thus > 0.

[0154] ③. Total Reward Function

[0155] (twenty two)

[0156] in: , , These are weighting coefficients. This is a penalty term for changes in weight.

[0157] Indicates the rate of improvement in stability. Indicates time Voltage stability index.

[0158] For energy efficiency improvement rate;

[0159] molecular This indicates improved stability (if VSI decreases, the improvement rate is positive).

[0160] The hyperparameters are set to [α,β, δ] = [0.4,0.3,0.2,0.1] The reward function encourages improvements in voltage stability and equipment reliability; it penalizes excessive weight jumps to ensure smooth weight adjustment.

[0161] Formalization of Markov Decision Processes (MDPs)

[0162] In summary, an MDP can be represented as a quintuple:

[0163] (twenty three)

[0164] Where: S: is the state space, and all possible S t Action space A: The set of discrete actions A mentioned above;

[0165] P(S t+1 |S t , ): This represents the state transition probability, which is determined by the evolution of the physical system and the softmax update.

[0166] R(S t , ,S t+1 )=R t ; is the discount factor γ∈(0,1) for the reward function, such as γ=0.99.

[0167] The MDP construction method adopts the reinforcement learning theoretical framework. This invention combines the characteristics of DC collector systems and independently designed the state, action and reward components.

[0168] S3: The reinforcement learning-driven agent interacts with the constructed Markov decision process environment to obtain the optimal weight adjustment policy. An improved proximal policy optimization (PPO) algorithm based on constrained state space and a composite reward function is used for model training, and the agent adopts an Actor-Critic architecture. Policy training is performed using the MDP defined in this invention. Figure 4 As shown.

[0169] A further implementation method is that the agent adopts an Actor-Critic architecture; both the Actor policy network and the Critic value network include an input layer, a first hidden layer, a second hidden layer, and an output layer.

[0170] Actor Network (Policy Network):

[0171] Input layer: 6-dimensional, corresponding to the state vector S t :

[0172] Hidden layer 1: Fully connected, 256 neurons, ReLU activation function:

[0173] (twenty four)

[0174] Hidden layer 2: Fully connected, 128 neurons, ReLU activation function:

[0175] (25)

[0176] Output layer: Dimension |A|=6, each dimension corresponds to the unnormalized score of an action.

[0177] (26)

[0178] The action probability distribution is obtained through softmax:

[0179] (27)

[0180] Where θ={W1,W2,W3,b1,b2,b3} are the policy network parameters.

[0181] Critic Network (Value Network):

[0182] Input layer: 6-dimensional, also a state vector. ;

[0183] The hidden layer structure can be similar to that of the Actor (two fully connected layers, 256 and 128, ReLU).

[0184] Output layer: 1-dimensional scalar, providing state value estimates:

[0185] (28)

[0186] Where ϕ represents the value network parameters. This is the output of Critic's second hidden layer.

[0187] Actor networks are used to provide the probability distribution of each action in the current state. The adjustment of decision evaluation weights; the Critic network is used to calculate state value estimation. To obtain the dominant function A t This stabilizes the policy gradient update.

[0188] like Figure 5 As shown, the objective function and update process of the improved PPO algorithm of this invention are as follows:

[0189] ①. Dominance function estimation:

[0190] Generalized dominance estimation (GAE) is used:

[0191] (29)

[0192] (30)

[0193] Where: λ is the GAE decay coefficient (0 < λ ≤ 1, such as 0.95). The superscript l represents the l-th time step after the current time step t.

[0194] The TD error (or single-step advantage estimate) at time t is used to measure the deviation of the sum of the immediate reward and expected value generated by the current action relative to the value of the current state.

[0195] Let be the estimate of the advantage function at time step t obtained using generalized advantage estimation (GAE), used to measure the advantage in state S. t Execute action The degree of superiority or inferiority compared to the benchmark strategy.

[0196] ②. The truncation objective function of the improved PPO:

[0197] For each time step, define the policy ratio:

[0198] (31)

[0199] Cutting objective function:

[0200] (32)

[0201] Where ϵ is the trimming parameter (e.g., 0.2). [] indicates that the empirical average is taken from all time step samples within a sampling batch.

[0202] Value function loss:

[0203] (33)

[0204] in The target value is the target state value at time step t. Represents the value network parameters obtained from the previous training round. State The value estimate, Used as a supervisory signal for the Critic network.

[0205] (34)

[0206] in This is the policy entropy, used to encourage the policy to maintain sufficient exploration early in training; , , which is a weighting coefficient used to balance the effects of the loss term and the entropy regularization term in the value function.

[0207] ③. Training Process

[0208] 1) Initialize the policy parameters θ and the value network parameters ϕ;

[0209] 2) In a simulation environment or historical playback environment, use the current strategy π θ Run several steps to collect sequences (S) t ,a t ,R t ,S t+1 );

[0210] 3) Calculate δ according to the above formula. t Advantage A t Target Value ;

[0211] 4) Use mini-batch gradient descent (such as the Adam optimizer) to minimize the loss L(θ,ϕ) and update θ,ϕ;

[0212] 5) Repeat steps 2 to 4 until the loss converges or the set number of training rounds is reached.

[0213] The aforementioned objective function and update process together constitute the core update mechanism of the improved PPO algorithm of this invention, which is used to train an adaptive adjustment strategy for evaluation weights under the constraints of the state-action space and the composite reward function.

[0214] While PPO, Actor-Critic, and GAE are publicly available reinforcement learning techniques, this invention does not simply apply them directly. Instead, it makes targeted improvements to them, forming an improved PPO algorithm for adaptive weight adjustment in offshore wind farm DC collector systems. The improvements include at least:

[0215] 1) The evaluation weight vector and the aggregated performance index vector are combined to form a joint state space. The constrained discrete weight increase and decrease actions are combined with the softmax normalization update mechanism to ensure that each weight parameter is always positive and the weight sum is always 1, thereby explicitly satisfying the engineering constraints of the comprehensive evaluation model.

[0216] 2) Construct a multi-objective composite reward function that includes stability improvement rate, reliability improvement rate, energy efficiency improvement rate and weight change penalty term, and embed the weight change penalty term as a regularization term into the policy update process to suppress drastic weight oscillations and improve learning stability.

[0217] 3) Based on the aforementioned constrained action space and composite reward function, the hyperparameters of the Actor-Critic structure and PPO update process are specifically tuned to ensure that the obtained optimal policy can achieve synergistic optimization of stability, reliability, and energy efficiency while satisfying engineering constraints. Through the above improvements, the reinforcement learning model constructed in this invention differs from existing general PPO / Actor-Critic / GAE algorithms in terms of state modeling method, action constraint mechanism, and reward design.

[0218] S4: Based on the optimal weight adjustment strategy, obtain the optimal weight vector, and perform a real-time comprehensive evaluation of the operating performance of the DC collection system of the offshore wind farm based on the optimal weight vector.

[0219] like Figure 6 As shown, the online weight update process is as follows:

[0220] ①. Real-time status awareness:

[0221] Output the aggregation performance metrics at the current moment. The weight W of the previous period t Together they form the state vector S t .

[0222] (35)

[0223] (36)

[0224] ②. Strategic Reasoning (Actor Forward Propagation)

[0225] S t Input the Actor network and calculate the probability distribution of each action:

[0226] (37)

[0227] in This represents the forward computation process of the policy network.

[0228] Choose the optimal action from among them:

[0229] In engineering, the action with the highest probability should be taken:

[0230] (38)

[0231] ③. Dynamic weighting:

[0232] The weighted dynamic execution module receives actions. As mentioned above:

[0233] (39)

[0234] (40)

[0235] The updated The results were then distributed to the comprehensive evaluation team.

[0236] (2) Comprehensive evaluation calculation

[0237] First, obtain the new weights. and the latest performance indicators :

[0238] (41)

[0239] Then calculate:

[0240] ① Overall score

[0241]

[0242] (42)

[0243] ② Dimensional Scoring

[0244] (43)

[0245] (44)

[0246] (45)

[0247] These results are presented to operations and maintenance personnel through a visual interface to assist in decision-making.

[0248] Example 2:

[0249] like Figure 1 As shown, the present invention also provides a weight adaptive adjustment system for comprehensive evaluation of the operation performance of offshore wind farm DC collection systems based on reinforcement learning, used to implement the method of Embodiment 1, including:

[0250] The data acquisition and preprocessing module is used to acquire and preprocess multi-source operating data of the offshore wind farm DC collection system to obtain real-time performance indicators. In a further implementation, the data acquisition and preprocessing module establishes a connection with the offshore wind farm DC collection system's monitoring and data acquisition system SCADA and equipment condition monitoring system CMS through an industrial communication protocol to realize the real-time acquisition of multi-source operating data.

[0251] The reinforcement learning intelligent decision-making module is used to construct a Markov decision process for weight adjustment based on real-time performance indicators and the current evaluation weight vector. As the core of the system's computing, it is deployed on an industrial server equipped with a GPU accelerator and has a built-in reinforcement learning algorithm based on the PyTorch 2.x series framework, which is responsible for the generation and optimization of the weight adjustment strategy.

[0252] The dynamic weight execution module is used to interact with the constructed Markov decision process environment based on reinforcement learning-driven intelligence to obtain the optimal weight adjustment strategy. Specifically, the weight adjustment strategy is parsed as adjustment instructions to increase, decrease, or keep unchanged each weight parameter in the current weight vector by a preset step size. As the control instruction execution unit, a high-reliability PLC or industrial computer is used to ensure the accurate execution of the weight adjustment instructions.

[0253] The comprehensive evaluation output module is used to obtain the optimal weight vector based on the optimal weight adjustment strategy, and to perform a real-time comprehensive evaluation of the operating performance of the DC collection system of the offshore wind farm based on the optimal weight vector.

[0254] A further implementation method is that the comprehensive evaluation output module, as a human-computer interaction interface, also provides a Web service interface and a visualization interface, supporting real-time display of evaluation results and historical data query.

[0255] The system in this embodiment is deployed in a simulated 400MW offshore wind farm environment (modified to enhance versatility), and the specific configuration is as follows:

[0256] 1. Hardware environment:

[0257] Server: Dell PowerEdge R750, 2 x Intel Xeon Gold 6330

[0258] Accelerator Card: NVIDIA A100 40GB

[0259] Memory: 256GB DDR4

[0260] Storage: 4TB NVMe SSD

[0261] 2. Software environment:

[0262] Operating System: Ubuntu 20.04 LTS

[0263] Deep learning frameworks: PyTorch 2.x series

[0264] Database: TimescaleDB 2.8

[0265] This system achieves significant improvements in several key performance aspects through dynamic weight adaptive optimization. The system operates efficiently, possessing the capability for periodic dynamic updates and rapid evaluation response. The core algorithm model boasts high inference accuracy, ensuring the system's continuous high availability.

[0266] Compared with traditional fixed-weight methods, this system has achieved comprehensive performance optimization in multiple core dimensions of power grid operation, specifically reflected in: a significant improvement in the accuracy of comprehensive evaluation, an effective improvement in system operating efficiency, a significant enhancement in equipment operating reliability, and a significant optimization of power grid voltage stability.

[0267] Figure 7 The following is a time series comparison chart of key indicators between the traditional fixed weight method and the adaptive weight method of the present invention in the embodiments of the present invention; wherein, (a) is a comparison chart of voltage stability indicators; (b) is a comparison chart of equipment health index; and (c) is a comparison chart of energy efficiency indicators.

[0268] Figure 8 This is an overview of the performance improvement of the present invention compared to the traditional fixed-weight method in the embodiments of the present invention.

[0269] Figure 9 This is a radar chart showing the multi-dimensional comprehensive performance of the traditional method compared to the present invention in an embodiment of the present invention.

[0270] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for adaptively adjusting weights in the comprehensive performance evaluation of offshore wind farm DC collector systems based on reinforcement learning, characterized in that, include: Collect and preprocess multi-source operation data of the DC collection system of offshore wind farm to obtain real-time performance indicators; Based on the real-time performance metrics and the current evaluation weight vector, a Markov decision process for weight adjustment is constructed. The agent, driven by reinforcement learning, interacts with the constructed Markov decision process environment to obtain the optimal weight adjustment strategy. Based on the optimal weight adjustment strategy, the optimal weight vector is obtained, and the operating performance of the offshore wind farm DC collection system is evaluated in real time based on the optimal weight vector.

2. The method according to claim 1, characterized in that, The multi-source operational data includes stability data, reliability data, and energy efficiency data; the real-time performance indicators include aggregated stability indicators, aggregated reliability indicators, and aggregated energy efficiency indicators. The stability data includes DC bus voltage fluctuation rate, converter station current fluctuation rate, and system frequency deviation. The reliability data includes the health index of key equipment, the predicted mean time between failures, and the equipment utilization rate. The energy efficiency data includes grid loss rate and transmission efficiency, which are used to characterize the loss level and energy utilization efficiency of DC power collection systems during energy transmission.

3. The method according to claim 2, characterized in that, The Markov decision process includes: a state space, an action space, and a reward function; The elements of the state space include the current evaluation weight vector and the real-time performance index vector; The elements of the action space are the adjustment actions of each weight parameter in the current evaluation weight vector using a preset adjustment step size; The reward function adopts a multi-objective composite form, including a first sub-reward term, a second sub-reward term, a third sub-reward term, and a stability penalty term.

4. The method according to claim 3, characterized in that, In the reward function, The first sub-reward item is the relative improvement rate of the voltage stability index between two adjacent time points; The second sub-reward item is the relative improvement rate of the health index of key equipment between two adjacent time points; The third sub-reward item is the relative improvement rate of the energy efficiency aggregate index between two adjacent time points; The stability penalty term is determined by the Euclidean distance between the evaluation weight vector at the current moment and the evaluation weight vector at the previous moment. The value of the reward function is obtained by weighting and summing the first sub-reward item, the second sub-reward item, and the third sub-reward item according to preset first, second, and third weight coefficients, and then subtracting the stability penalty item.

5. The method according to claim 1, characterized in that, The agent adopts an Actor-Critic architecture; both the Actor policy network and the Critic value network include an input layer, a first hidden layer, a second hidden layer, and an output layer. Actor networks are used to provide the probability distribution of each action in the current state and to adjust the decision evaluation weights. Critic networks are used to calculate state value estimates and obtain the advantage function.

6. A weighted adaptive adjustment system for comprehensive performance evaluation of offshore wind farm DC collector systems based on reinforcement learning, used to implement the method described in any one of claims 1-5, characterized in that, include: The data acquisition and preprocessing module is used to acquire and preprocess multi-source operating data of the DC collection system of offshore wind farms to obtain real-time performance indicators. The reinforcement learning intelligent decision-making module is used to construct a Markov decision process for weight adjustment based on the real-time performance indicators and the current evaluation weight vector. The weight dynamic execution module is used to enable the reinforcement learning-driven agent to interact with the constructed Markov decision process environment to obtain the optimal weight adjustment strategy. The comprehensive evaluation output module is used to obtain the optimal weight vector based on the optimal weight adjustment strategy, and to perform a real-time comprehensive evaluation of the operating performance of the offshore wind farm DC collection system based on the optimal weight vector.

7. The system according to claim 6, characterized in that, The data acquisition and preprocessing module establishes a connection with the SCADA system and CMS system of the offshore wind farm DC collection system through the industrial communication protocol, so as to realize the real-time acquisition of multi-source operation data.

8. The system according to claim 6, characterized in that, The comprehensive evaluation output module also provides a web service interface and a visualization interface, supporting real-time display of evaluation results and historical data query.