Intelligent fault diagnosis method for renewable energy equipment based on deep learning

By combining fractional-order physical modeling and dynamic statistical learning, the problems of difficult mechanism modeling and poor environmental adaptability in fault diagnosis of renewable energy equipment are solved, efficient identification and stable prediction of complex dynamic behaviors are achieved, and the accuracy of fault diagnosis and the efficiency of strategy optimization are improved.

CN120336789BActive Publication Date: 2025-10-03TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510814092.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-03
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

Existing technologies in renewable energy equipment fault diagnosis have difficulties in mechanism modeling, poor environmental adaptability, slow strategy updates, and are unable to effectively identify sudden faults. In addition, traditional methods are insufficient in modeling nonlinear and non-integer-order dynamic behaviors, resulting in low diagnostic efficiency and poor accuracy.

Method used

By adopting the methods of fractional-order physical modeling, dynamic statistical learning and reinforcement learning closed-loop optimization, a weighted shift Grünwald-Letnikov format fractional-order physical information neural network (WSGD-FPINNs) and a piecewise Poisson regression model were constructed. Combined with the relative reward regression mechanism, the fault prediction strategy was optimized to achieve accurate identification and dynamic adaptation of equipment failures.

Benefits of technology

It improves the fitting accuracy and robustness of complex non-integer-order dynamic behaviors, dynamically adapts to non-stationary environments, improves fault recognition rate and strategy stability, reduces predictive maintenance costs, and achieves a technological leap from passive maintenance to active prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336789B_ABST
    Figure CN120336789B_ABST
Patent Text Reader

Abstract

The present invention provides a method for intelligent diagnosis of renewable energy equipment faults based on deep learning. It relates to the field of renewable energy fault diagnosis and includes: S1: collecting data such as photovoltaic systems and performing data preprocessing to construct a training sample set with fault labels; S2: constructing a fractional-order physical information neural network in a weighted shift Grünwald‑Letnikov format to output a predicted value of the equipment failure probability; S3: establishing a piecewise Poisson regression model, using time variables and environmental virtual variables as independent variables, performing dynamic trend modeling on the predicted value of the equipment failure probability, and outputting a statistically corrected failure probability; S4: fusing the predicted value of the equipment failure probability with the corrected failure probability to generate a final failure probability; S5: optimizing the prediction strategy parameters based on a reinforcement learning method based on a relative reward regression mechanism. The present invention improves the ability to identify implicit fault features and significantly improves the accuracy of fault diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of renewable energy fault diagnosis, and in particular to a deep learning-based intelligent fault diagnosis method for renewable energy equipment. Background Art

[0002] With the rapid development of renewable energy technologies, photovoltaic, wind power, and energy storage systems are increasingly playing a role in the modern energy mix. However, these devices are susceptible to various types of failures during long-term operation, such as hot spots on photovoltaic modules, mechanical wear of wind turbines, and battery degradation. If not promptly identified and addressed, these problems can severely impact the system's power generation efficiency and safety.

[0003] Traditional fault diagnosis methods rely on manual experience and feature engineering, often requiring large amounts of labeled data and complex signal processing pipelines, and are insufficiently capable of modeling complex nonlinear and non-integer-order physical processes. To this end, researchers have attempted to introduce deep learning methods into the field of fault diagnosis. In particular, physical information neural networks (PINNs), which embed physical priors into deep models, offer an effective approach to overcoming the problems of small sample sizes and weak supervision. However, standard PINNs still have limited performance when faced with the fractional-order dynamic behaviors commonly found in equipment operation (such as memory effects and dissipative delays), and are unable to effectively model non-integer-order control characteristics. Traditional statistical models ignore physical mechanisms and have a delayed response to sudden faults (such as photovoltaic hot spots).

[0004] At the same time, identifying equipment faults often involves extracting and modeling multi-scale, segmented features. Accurately characterizing structural changes in the data poses a major challenge for diagnostic systems. Furthermore, with the widespread application of large models in sequential tasks, more sophisticated modeling mechanisms are urgently needed to effectively evaluate the quality of generated outputs. Furthermore, when optimizing generative model strategies, traditional reinforcement learning methods suffer from low sample efficiency and poor convergence. Consequently, there is a lack of an efficient, theoretically robust, and practically deployable intelligent diagnosis method for renewable energy equipment faults. Summary of the Invention

[0005] The present invention aims to solve the above-mentioned problems and proposes a deep learning-based intelligent diagnosis method for renewable energy equipment faults. Through fractional-order physical modeling, dynamic statistical learning and reinforcement learning closed-loop optimization, it solves the technical difficulties in renewable energy equipment fault diagnosis, such as difficult mechanism modeling, poor environmental adaptability, and slow strategy updates. It also overcomes the technical problems of existing technologies such as delayed response to sudden faults (such as photovoltaic hot spots), low strategy optimization sample efficiency, unstable convergence, and low fault recognition rate.

[0006] To achieve the above objectives, the following technical solutions are adopted:

[0007] A deep learning-based intelligent fault diagnosis method for renewable energy equipment includes the following steps:

[0008] S1: Collect sensor data from photovoltaic systems, battery energy storage systems, and wind turbines, perform data preprocessing, and construct a training sample set with fault labels;

[0009] S2: Construct a fractional-order physical information neural network in the weighted shift Grünwald-Letnikov format, approximate the fractional-order derivative terms through discretization and introduce a physical constraint loss function to output the physical-driven equipment failure probability prediction value. ;

[0010] S3: Establish a segmented Poisson regression model, using time variables and environmental dummy variables as independent variables to predict the probability of equipment failure. Perform dynamic trend modeling and output statistically corrected failure probability ;

[0011] S4: Use weighted fusion strategy to predict the probability of failure of the equipment The statistically corrected probability of failure Fusion is performed to generate the final failure probability ;

[0012] S5: Reinforcement learning method based on relative reward regression mechanism, and environmental context features as input, and optimize the prediction strategy parameters by minimizing the relative reward difference loss function.

[0013] Furthermore, the construction of the weighted shift Grünwald-Letnikov format fractional-order physical information neural network includes:

[0014] Establish fractional-order governing equations;

[0015] Using deep neural network-like state response variables, fractional derivatives are approximated by weighted shift Grünwald-Letnikov coefficients.

[0016] Designing a comprehensive loss function ,in, is the physical residual loss term, is the observation data error term, is the initial boundary condition loss term; : Loss term weighting coefficient, used to balance the impact of different error terms.

[0017] Furthermore, the construction of the weighted shift Grünwald-Letnikov format fractional-order physical information neural network further includes:

[0018] The fault parameter θ is optimized through the parameter inversion mechanism, specifically including:

[0019] The fault parameter θ is used as the reinforcement learning action output; the reward function is constructed based on the neural network prediction error and physical consistency; the improved deep reinforcement learning method is used to optimize the fault parameter θ. The intelligent agent continuously adjusts the parameter through trial and error learning to minimize the prediction error, achieve accurate inversion of renewable energy equipment faults, and output the optimal fault parameter after inversion. , optimal network weight parameters and predicted values ​​of state variables , where x is the spatial position variable and t is the time variable.

[0020] Furthermore, the equipment failure probability prediction value By responding to the network output state variable The nonlinear transformation is obtained, that is, the normalized mapping is achieved through a Sigmoid function:

[0021]

[0022] in, : represents the predicted value of the failure probability of a certain device; : The state variable finally output by the neural network is the response of the trained network under the inverted parameters; : Mapping function from state response to fault risk score; : Sigmoid function, used to normalize the mapping result to interval.

[0023] Furthermore, the piecewise Poisson regression model expression is:

[0024]

[0025] in, is the equipment failure probability output by the piecewise Poisson regression model; T represents the observation time; Indicates the time node set in the model, that is, the turning point of the fault occurrence trend change; and dummy variables representing three external operating environments: high temperature, strong wind, and high humidity; is the regression coefficient to be estimated; is a random error term, which is subject to the Poisson distribution assumption.

[0026] Furthermore, wherein, said S4: adopting weighted fusion strategy to predict the probability of failure of said equipment The statistically corrected probability of failure Fusion is performed to generate the final failure probability :

[0027]

[0028] in, and are the weight coefficients of the prediction results of the two models, satisfying + =1; and, the weight coefficient and Determined by any of the following methods: dynamic adjustment based on historical forecast accuracy; allocation based on model confidence intervals; or adaptive optimization using cross-validation.

[0029] Furthermore, the reinforcement learning method based on the relative reward regression mechanism achieves strategy optimization by minimizing the following loss function:

[0030]

[0031] in, : The current strategy parameters to be optimized, representing the weight set of the strategy model; : The strategy parameters at the nth iteration are the reference benchmark for the current strategy; : New strategy parameters obtained through this optimization update; : Strategy parameter space, representing the set of all possible parameter values; : Data sample triplet, indicating the The state-response pairs collected in rounds, : Status, including sensor observations, environmental parameters, and historical prediction results during device operation; : An action generated under the current policy, including control instructions and maintenance suggestions; : Another action for comparison, used to construct the "relative reward" difference; : The training dataset collected in round n, i.e., the set of state and response pairs; : In state Under the current strategy parameters Output action The probability value of represents the strategy's preference for action selection; : In state Next, the last iteration strategy Output Action The probability of , which is used to measure the difference between the new and old strategies; : In state Next action The immediate rewards received; : The learning rate for policy updates, which controls the step size during policy iteration.

[0032] Furthermore, the method further comprises:

[0033] S6: Probabilistically model the policy actions through the output token probability estimation method to generate executable maintenance instructions or scheduling plans.

[0034] Furthermore, the output token probability estimation method adopts the following probability modeling formula:

[0035]

[0036] in, is a complete action token sequence consisting of the middle idea token Token with action execution spliced ​​together, ρ is the scaling factor and 0.2≤ρ≤0.5; Input to the current environment, serving as background information for generating actions; For the system at time Contextual state information, including current observation state, historical interactions, and system feedback.

[0037] Furthermore, in said S1, wherein, the photovoltaic system sensor data includes output voltage, current, component temperature and ambient temperature and humidity; the battery energy storage system data includes cell voltage, battery pack temperature, current fluctuation, SOC and SOH; the wind turbine generator set data includes blade speed, vibration signal, main shaft temperature, current, voltage and wind speed;

[0038] Data preprocessing operations include: using bandpass filtering or wavelet denoising to remove noise from current, voltage, and vibration signals; eliminating outliers through Z-score analysis and sliding median method; slicing the time series signal according to a fixed time length to construct a sample window and associating it with a fault type label.

[0039] Compared with the prior art, the present invention achieves the following beneficial effects:

[0040] 1. The present invention constructs a fractional-order physical information neural network (WSGD-FPINNs) in a weighted shift Grünwald-Letnikov format. By introducing a weighted discrete format to accurately approximate fractional-order derivatives, the physical information neural network has higher fitting accuracy and robustness in modeling complex non-integer-order dynamic behaviors such as equipment aging and energy dissipation, thereby improving the ability to identify hidden fault characteristics.

[0041] 2. The present invention constructs a piecewise Poisson regression model, which can automatically identify structural change points in operating data. The model is suitable for dynamic modeling of changes in fault frequency over time or equipment status, helps to discover non-stationary failure trends and mutation signal characteristics, and dynamically adapts to non-stationary environments.

[0042] 3. This paper proposes a reinforcement learning method based on a relative reward regression mechanism, which transforms the policy optimization problem into a least squares regression problem of relative rewards, avoids dependence on partition functions, significantly improves training efficiency and policy stability, and is particularly suitable for policy fine-tuning in sequence generation tasks.

[0043] 4. The present invention proposes a method for estimating the probability of output tokens, which can perform probabilistic modeling on the tokens output by a language model or a sequence generation model, and accurately characterize the uncertainty of the output, thereby realizing dynamic evaluation and error correction of the model output results, and improving the credibility of the generated diagnostic recommendations or prediction results.

[0044] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The above and other features, advantages and aspects of the embodiments of the present invention will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are provided for a better understanding of the present invention and do not constitute a limitation of the present invention. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:

[0046] Figure 1 This is a flowchart of a method for intelligent fault diagnosis of renewable energy equipment based on deep learning according to an embodiment of the present invention;

[0047] Figure 2 1 is a flowchart of a method for intelligent fault diagnosis of renewable energy equipment based on deep learning according to another embodiment of the present invention;

[0048] Figure 3 This is a schematic diagram of the architecture of a deep learning-based intelligent fault diagnosis method for renewable energy equipment according to an embodiment of the present invention. DETAILED DESCRIPTION

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0050] In this document, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.

[0051] Figure 1 A schematic flow chart of a method for intelligent fault diagnosis of renewable energy equipment based on deep learning according to an embodiment of the present invention is shown; Figure 3 This is a schematic diagram of the architecture of a deep learning-based intelligent fault diagnosis method for renewable energy equipment according to an embodiment of the present invention. Figure 1 As shown, a deep learning-based intelligent fault diagnosis method for renewable energy equipment includes the following steps:

[0052] S1: Collect sensor data from photovoltaic systems, battery energy storage systems, and wind turbines, perform data preprocessing, and construct a training sample set with fault labels;

[0053] This step S1, as the foundation of the entire fault diagnosis system, aims to build a structured, high-quality data sample set that can be used for deep model training. By integrating operating data from different types of renewable energy equipment, it comprehensively reflects their physical state and operating conditions, supporting subsequent feature modeling and classification tasks. Furthermore, this step 1 includes the following steps:

[0054] S11: Sensor Data Acquisition

[0055] The present invention collects data sources from the following typical renewable energy devices:

[0056] Photovoltaic system: collects parameters such as output voltage, current, component temperature, ambient temperature and humidity;

[0057] Battery energy storage system: including cell voltage, battery pack temperature, current fluctuation, state of charge (SOC), state of health (SOH), etc.

[0058] Wind turbine generators: Collect operating parameters such as blade rotation speed, vibration signals, main shaft temperature, current, voltage, and wind speed. Data preprocessing operations include: using bandpass filtering or wavelet denoising to remove noise from current, voltage, and vibration signals; eliminating outliers through Z-score analysis and the sliding median method; slicing the time series signal according to a fixed duration to construct a sample window and associating fault type labels.

[0059] S12: Data Preprocessing

[0060] To ensure the consistency of sensor data, the system introduces a unified timestamp mechanism to align the timing of various types of collected data and fill in gaps. In addition, the following preprocessing operations are performed:

[0061] Noise filtering: Use bandpass filtering, wavelet denoising and other algorithms to remove random noise from current, voltage and vibration signals;

[0062] Outlier detection and elimination: Z-score analysis and sliding median method are used to eliminate non-physical outliers (such as measurement mutations caused by instrument failure);

[0063] Sliding window slicing to construct sample sets: Slice the continuous time series signal according to a fixed length to construct a sample window.

[0064] Label and sample organization: Use known equipment failure events as labels for training samples, and mark the corresponding failure type or normal state for each window.

[0065] S2: Construct a fractional-order physical information neural network in the weighted shift Grünwald-Letnikov format, approximate the fractional-order derivative terms through discretization and introduce a physical constraint loss function to output the physical-driven equipment failure probability prediction value. ;

[0066] The step S2 specifically includes the following steps:

[0067] S2.1: Constructing Weighted Shift Grünwald-Letnikov Fractional-Order Physical Information Neural Networks (WSGD-FPINNs)

[0068] S2.1.1: Problem Modeling

[0069] The fractional-order control equation is established based on the response data of the actual equipment. The equation can be expressed as follows:

[0070]

[0071] : State response variables (such as voltage, current, temperature, vibration intensity, etc.) are the target physical quantities to be predicted; :right The fractional derivative of is used to describe the historical memory effect and complex degradation behavior within the device; : fractional order, , reflecting the non-local time response characteristics of equipment degradation; : Spatial position variable, which can represent device number, sensor location, component area, etc. : Time variable, used to model the evolution of faults over time; : The fault parameter vector to be identified (such as fault location, type, severity level, etc.); : The governing equations that contain the physical constraints of the device can describe coupled physical mechanisms such as thermal, electrical, and mechanical. For example:

[0072] (1) Photovoltaic system hot spot failure model (thermal-electric coupling)

[0073] Governing equations:

[0074] in, : PV module temperature (state variable); : 1.5-order fractional derivative (α=1.5), describing the non-local thermal memory effect of silicon materials; : Fault parameters (R: Hot spots cause local resistance to increase, : blocks additional heat source); κ is the effective thermal conductivity, h is the convective heat transfer coefficient; : Joule heating term (strongly correlated with the fault resistance R); ρ is the density of the photovoltaic material; is the specific heat capacity of the material; is the ambient temperature; σ is the conductivity; I is the working current; is the fault resistance.

[0075] The hot spot area generates abnormal Joule heating due to increased resistance, and the fractional order operator More accurately characterizes the superdiffusion behavior of heat in inhomogeneous materials (traditional integer-order models underestimate the heat accumulation rate). , R rises, indicating an abnormal increase in resistance (hot spot sign); >0 indicates that there is an additional heat source blocked.

[0076] (2) Battery energy storage system aging model (electrochemical-thermal coupling)

[0077] A. Active lithium loss kinetic equation

[0078]

[0079] B. State of Charge (SOC) Evolution Equation

[0080]

[0081] in, : Active lithium loss (key state variable); : fractional derivatives (α1=1.2, α2=0.8), describing the sublinear kinetics of lithium ion diffusion; : Fault parameters ( : Capacity attenuation coefficient, : The activation energy barrier increases, : side reaction rate); L: active lithium loss (positively correlated with aging); I: charge and discharge current; is the rated capacity; is the lithium loss-capacity attenuation coefficient;

[0082] The fractional order α < 1 reflects the slow kinetics of ion diffusion (diffusion is hindered due to aging), and α > 1 represents the delayed response of charge transfer. middle, Rising, indicating that the side reaction is accelerating (electrolyte decomposition); A decrease indicates a decrease in the barrier to ion migration (electrode damage); A decrease indicates capacity degradation (decreasing state of health SOH).

[0083] (3) Wind turbine bearing wear model (mechanical vibration)

[0084] Governing equations:

[0085] in, : bearing vibration displacement (state variable); : fractional derivative (α=1.6), simulating the history dependence of friction damping; : Fault parameters ( : Defect additional impact force, : defect location); δ(·): Dirac function (locating local defects); m is the equivalent mass; y is the vibration displacement; c is the damping coefficient; k is the stiffness coefficient; is the basic exciting force; w is the rotation angular frequency; Defect impact force.

[0086] Fractional damping term More accurate description of nonlinear friction dissipation (traditional integer-order models cannot capture the energy decay hysteresis under high-frequency impact). middle, >0, indicating the presence of local impact (wear / crack), Able to accurately locate defects.

[0087] S2.1.2: Network structure design

[0088] Building a deep neural network As an approximation of the state variables, is the network weight parameter. Automatic differentiation is used to calculate integer-order derivatives, and the weighted shift Grünwald-Letnikov (WSGD) method is used to approximate fractional-order derivatives:

[0089]

[0090] H: spatial discrete step length; N: number of truncation terms, used to control the accuracy and computational complexity of derivative approximation; : Weighted Grünwald-Letnikov coefficient, specifically defined as follows:

[0091]

[0092] : Grünwald coefficient, defined as:

[0093]

[0094] S2.1.3: Loss Function Design

[0095] The comprehensive loss function is defined to include physical residual loss, boundary condition loss and observation data error term:

[0096]

[0097] : Total loss function, used for the optimization objective of neural network training; : Physical residual loss, which measures the fitting error of the prediction results to the fractional-order physical equation; : The difference between the actual fault monitoring data (such as sensor data); : Initial boundary condition loss (such as the initial state of the device, boundary temperature or voltage, etc.); : Loss term weighting coefficient, used to balance the impact of different error terms.

[0098] S2.1.4: Introducing parameter inversion mechanism

[0099] The fault parameters As the action output of the intelligent agent, the state space is defined as the neural network prediction error and response behavior. A reward mechanism based on loss function and physical consistency is designed. The improved deep reinforcement learning method is used to optimize the fault parameter θ. The intelligent agent continuously adjusts the parameters through trial and error learning to achieve accurate inversion of renewable energy equipment faults and output the inverted parameters. , optimal network weight parameters and predicted values ​​of state variables This method does not rely on the differentiability of the objective function and is applicable to complex, multi-source physical models.

[0100] S2.1.5: Output results

[0101] The probability value of equipment failure is calculated by the network output state response variable The nonlinear transformation is obtained, that is, the normalized mapping is achieved through a Sigmoid function:

[0102]

[0103] : represents the predicted value of the failure probability of a certain device; : The state variables (such as temperature, voltage, vibration intensity, etc.) ultimately output by the neural network are the responses of the trained network under the inverted parameters; : The mapping function from the state response to the fault risk score, usually a linear combination or a fully connected layer, for example: , where u is the input state vector, for example, in a battery scenario: u = [voltage fluctuation, temperature gradient, lithium loss]; w is the weight vector, reflecting the contribution of each physical quantity to the fault; b is the bias term, the benchmark threshold for fault judgment.

[0104] : Sigmoid function, used to normalize the mapping result to The interval is expressed as a probability:

[0105]

[0106] S3: Establish a segmented Poisson regression model, using time variables and environmental dummy variables as independent variables to predict the probability of equipment failure. Perform dynamic trend modeling and output statistically corrected failure probability ;

[0107] Step S3 further characterizes the dynamic trend of fault probability over time and environmental changes. Based on the physical driver prediction results, a piecewise Poisson regression model is introduced to model the predicted probability. This model uses time as the primary independent variable and the predicted probability value as the dependent variable. It also introduces multiple environmental conditions as dummy variables to enhance the model's adaptability to external disturbances.

[0108] A piecewise Poisson regression model is established, and its mathematical expression is as follows:

[0109]

[0110] in, The probability of equipment failure output by the piecewise Poisson regression model; Indicates the observation time; Indicates the time node set in the model (such as the turning point of the fault trend change); and dummy variables representing three external operating environments (such as high temperature, strong wind, high humidity, etc.); is the regression coefficient to be estimated; is a random error term, which is subject to the Poisson distribution assumption.

[0111] Through this model, we can achieve a quantitative description of the failure rate change trend under different environmental scenarios, and improve the model's sensitivity and generalization ability to changes in boundary conditions.

[0112] S4: Use weighted fusion strategy to predict the probability of failure of the equipment The statistically corrected probability of failure Fusion is performed to generate the final failure probability ;

[0113] In order to take into account the high interpretability of physical modeling and the generalization ability of statistical learning models, this step S4 uses a weighted fusion strategy to integrate the prediction results of the fractional-order physical information neural network (step S2) and the piecewise Poisson regression model (step S3). The final prediction probability after fusion is Expressed as:

[0114]

[0115] in, The failure probability predicted by WSGD-FPINNs; is the failure probability output by the piecewise Poisson regression model; and are the weight coefficients of the prediction results of the two models, satisfying + =1.

[0116] Among them, the weight coefficient and This weighting strategy can be set based on historical prediction accuracy, model confidence intervals, or expert experience, and can also be adaptively optimized through cross-validation. This weighting strategy effectively integrates the physical drive model's ability to deeply characterize degradation mechanisms with the statistical model's ability to respond to multi-source interference, improving the accuracy and robustness of the overall fault prediction system.

[0117] S5: Reinforcement learning method based on relative reward regression mechanism, and environmental context features as input, and optimize the prediction strategy parameters by minimizing the relative reward difference loss function.

[0118] In order to achieve dynamic adjustment and closed-loop optimization of the prediction strategy, a reinforcement learning method based on the relative reward regression mechanism is introduced on the basis of integrating physical modeling and statistical regression prediction results (see steps S2-S4). The purpose of this algorithm is to provide an efficient, scalable and conservative strategy optimization path in the context of constantly changing equipment operating environment and the need for continuous updating of prediction mechanisms. Unlike traditional reinforcement learning that relies on global reward signals, the reinforcement learning method based on the relative reward regression mechanism reconstructs the strategy optimization problem into a supervised regression problem of relative rewards, so as to reduce variance fluctuations in strategy updates and enhance the ability to robustly respond to changes in prediction errors. Combined with the predicted probability value generated in step 2 Combined with the environmental context features, a reinforcement learning method based on the relative reward regression mechanism is optimized and modeled in the contextual bandit setting.

[0119] The reinforcement learning objective is achieved by minimizing a regression loss function of the following form:

[0120]

[0121] in, : The current strategy parameters to be optimized, representing the weight set of the strategy model; : The strategy parameter at the nth iteration is the reference benchmark for the current strategy, where n represents the number of iterations; : New strategy parameters obtained through this optimization update; : Strategy parameter space, representing the set of all possible parameter values; : Data sample triplet, representing the state-response pair collected in round n; : Status, such as sensor observation values, environmental parameters, historical prediction results, etc. during device operation; : An action generated under the current policy, such as control instructions, maintenance suggestions, etc. : Another action for comparison, used to construct the "relative reward" difference. : The training dataset collected in round n (a set of state and response pairs). : In state Under the current strategy parameters Output action The probability value of , which indicates the strategy's preference for action selection. : In state Next, the last iteration strategy Output Action , which is used to measure the difference between the new and old strategies. : In state Next action The immediate rewards obtained can be defined by indicators such as prediction accuracy, system stability or execution effect. : The learning rate for policy updates, which controls the step size during policy iteration.

[0122] The objective function compares two responses under the same state. and The relative value of This improves the computational feasibility of the algorithm.

[0123] Figure 2 FIG. 1 shows a flow chart of a method for intelligent fault diagnosis of renewable energy equipment based on deep learning according to another embodiment of the present invention; FIG. Figure 2 As shown, the method further includes:

[0124] S6: Probabilistically model the policy actions through the output token probability estimation method to generate executable maintenance instructions or scheduling plans.

[0125] To achieve a closed-loop feedback loop between policy generation and environment interaction, we further introduce an output token probability estimation method. This method is particularly suitable for text-driven or natural language interaction policy models, where the policy output is structured text. Post-processing functions are required to parse text actions into executable actions, interact with the external environment, and obtain reward signals and state transitions.

[0126] To estimate the action probability, the output token probability estimation method adopts the following probabilistic modeling formula:

[0127]

[0128] in, is a complete action token sequence consisting of the middle "idea" token Token with action execution Spliced ​​together, the splicing form is usually ; : The intermediate “thought” token in the strategy generation process reflects the intermediate thinking or explanatory content in the strategy reasoning process; : The final "action" token, which represents the executable response given by the policy network, such as maintenance instructions, scheduling plans, etc. Input for the current environment, which serves as background information for generating actions, such as the current operating status of renewable energy equipment, sensor data, diagnostic context, etc. For the system at time Contextual state information, including current observation state, historical interactions, system feedback, etc. : Policy network, composed of parameters The probability model of control is used to generate the conditional distribution of action tokens. ρ is a scaling factor used to adjust the importance of the "idea" part in the overall action probability modeling. ρ has a significant impact on the final reinforcement learning performance. Extreme values ​​(such as close to 0 or 1) will cause the policy learning bias to be too narrow or the distribution to be too dispersed. In order to maintain the stability and expressiveness of the policy update, it is recommended to select .

[0129] The deployment and application of the deep learning-based intelligent fault diagnosis method for renewable energy equipment according to the above embodiment of the present invention are as follows:

[0130] The overall system adopts a modular deployment architecture with containerization and microservices as the core to achieve rapid delivery and cross-platform migration. The deployment architecture supports elastic scaling through Kubernetes+Docker, and supports isolated deployment and parallel management of policy models in multi-tenant environments.

[0131] This system has broad cross-industry applicability, with three typical deployment scenarios listed below: energy storage unit scheduling in smart grid systems: predicting power fluctuation trends, using reinforcement learning to dynamically optimize charging and discharging strategies, balancing loads and reducing energy costs; industrial equipment fault prediction and proactive maintenance: integrating stress-time physical models with monitoring data to identify potential faults in advance and generate scheduling or maintenance recommendations; and pipe network risk warning in urban water systems: predicting pipe burst risks based on flow-pressure distribution and historical event statistics, and dynamically adjusting water valve control strategies using reinforcement learning methods based on a relative reward regression mechanism.

[0132] According to the above embodiments of the present invention, by constructing WSGD-FPINNs and a piecewise Poisson regression model, a relative reward regression mechanism is designed to avoid partition function calculation, improve strategy stability, accurately model non-integer-order physical processes of equipment, improve fault feature recognition rate, obtain integrated physical and statistical prediction results, improve the accuracy of full life cycle diagnosis, and reduce predictive maintenance costs. Through the deep coupling of physical mechanisms and data-driven, a unified diagnostic framework for three types of equipment, photovoltaic / wind power / energy storage, is supported, and a diagnostic chain of fractional-order modeling → dynamic statistical correction → closed-loop decision optimization is constructed, which effectively solves the problems of complex mechanisms, time-varying environment, and strategy lag in fault diagnosis of renewable energy equipment, and realizes a technological leap from passive maintenance to active prediction.

[0133] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of each step described above can refer to the corresponding process in the aforementioned system embodiment and will not be repeated here.

[0134] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the methods.

[0135] It should also be noted that, in the embodiments of the present application, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device comprising the elements.

[0136] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined in the embodiments of the present application may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown in the embodiments of the present application, but rather will conform to the widest scope consistent with the principles and novel features disclosed in the embodiments of the present application.

Claims

1. A deep learning-based intelligent fault diagnosis method for renewable energy equipment, characterized in that: The following steps are involved: S1: Collect sensor data from photovoltaic systems, battery energy storage systems, and wind turbines, perform data preprocessing, and construct a training sample set with fault labels; S2: Construct a fractional-order physical information neural network in the weighted shift Grünwald-Letnikov format, approximate the fractional-order derivative terms through discretization and introduce a physical constraint loss function to output the physical-driven equipment failure probability prediction value. ; S3: Establish a segmented Poisson regression model, using time variables and environmental dummy variables as independent variables to predict the probability of equipment failure. Perform dynamic trend modeling and output statistically corrected failure probability ; S4: Use weighted fusion strategy to predict the probability of failure of the equipment The statistically corrected probability of failure Fusion is performed to generate the final failure probability ; S5: Reinforcement learning method based on relative reward regression mechanism, with the final failure probability and environmental context features as input, and optimize the prediction strategy parameters by minimizing the relative reward difference loss function.

2. The method according to claim 1, characterized in that in, The construction of the fractional-order physical information neural network of the weighted shift Grünwald-Letnikov format includes: Establish fractional-order governing equations; Using deep neural network-like state response variables, fractional derivatives are approximated by weighted shift Grünwald-Letnikov coefficients. Designing a comprehensive loss function ,in, is the physical residual loss term, is the observation data error term, is the initial boundary condition loss term; : Loss term weighting coefficient, used to balance the impact of different error terms.

3. The method according to claim 2, characterized in that in, The construction of the fractional-order physical information neural network of the weighted shift Grünwald-Letnikov format also includes: The fault parameter θ is optimized through the parameter inversion mechanism, specifically including: The fault parameter θ is used as the reinforcement learning action output; the reward function is constructed based on the neural network prediction error and physical consistency; the improved deep reinforcement learning method is used to optimize the fault parameter θ. The intelligent agent continuously adjusts the parameter through trial and error learning to minimize the prediction error, achieve accurate inversion of renewable energy equipment faults, and output the optimal fault parameter after inversion. , optimal network weight parameters and predicted values ​​of state variables , where x is the spatial position variable and t is the time variable.

4. The method according to claim 3, characterized in that in, The predicted value of the equipment failure probability By responding to the network output state variable The nonlinear transformation is obtained, that is, the normalized mapping is achieved through a Sigmoid function: ; in, : represents the predicted value of the failure probability of a certain device; : The state variable finally output by the neural network is the response of the trained network under the inverted parameters; : Mapping function from state response to fault risk score; : Sigmoid function, used to normalize the mapping result to interval.

5. The method according to claim 1, wherein in, The piecewise Poisson regression model expression is: ; in, is the equipment failure probability output by the piecewise Poisson regression model; T represents the observation time; Indicates the time node set in the model, that is, the turning point of the fault occurrence trend change; and dummy variables representing three external operating environments: high temperature, strong wind, and high humidity; is the regression coefficient to be estimated; is a random error term, which is subject to the Poisson distribution assumption.

6. The method according to claim 1, wherein in, S4: using a weighted fusion strategy to predict the equipment failure probability The statistically corrected probability of failure Fusion is performed to generate the final failure probability : ; in, and are the weight coefficients of the prediction results of the two models, satisfying + =1; and, the weight coefficient and Determined by any of the following methods: dynamic adjustment based on historical forecast accuracy; allocation based on model confidence intervals; or adaptive optimization using cross-validation.

7. The method according to claim 4, characterized in that The reinforcement learning method based on the relative reward regression mechanism achieves policy optimization by minimizing the following loss function: ; in, : The current strategy parameters to be optimized, representing the weight set of the strategy model; : The strategy parameters at the nth iteration are the reference benchmark for the current strategy; : New strategy parameters obtained through this optimization update; : Strategy parameter space, representing the set of all possible parameter values; : Data sample triplet, indicating the The state-response pairs collected in rounds, : Status, including sensor observations, environmental parameters, and historical prediction results during device operation; : An action generated under the current policy, including control instructions and maintenance suggestions; : Another action for comparison, used to construct the "relative reward" difference; : The training dataset collected in round n, i.e., the set of state and response pairs; : In state Under the current strategy parameters Output action The probability value of represents the strategy's preference for action selection; : In state Next, the last iteration strategy Output Action The probability of , which is used to measure the difference between the new and old strategies; : In state Next action The immediate rewards received; : The learning rate for policy updates, which controls the step size during policy iteration.

8. The method according to claim 7, characterized in that The method further comprises: S6: Probabilistically model the policy actions through the output token probability estimation method to generate executable maintenance instructions or scheduling plans.

9. The method according to claim 8, characterized in that in, The output token probability estimation method adopts the following probability modeling formula: ; in, is a complete action token sequence consisting of the middle idea token Token with action execution spliced ​​together, ρ is the scaling factor and 0.2≤ρ≤0.5; Input to the current environment, serving as background information for generating actions; For the system at time Contextual state information, including current observation state, historical interactions, and system feedback.

10. The method according to claim 8, characterized in that in, In S1, the photovoltaic system sensor data includes output voltage, current, component temperature and ambient temperature and humidity; the battery energy storage system data includes cell voltage, battery pack temperature, current fluctuation, SOC and SOH; the wind turbine generator set data includes blade speed, vibration signal, main shaft temperature, current, voltage and wind speed; Data preprocessing operations include: using bandpass filtering or wavelet denoising to remove noise from current, voltage, and vibration signals; eliminating outliers through Z-score analysis and sliding median method; slicing the time series signal according to a fixed time length to construct a sample window and associating it with a fault type label.

Citation Information

Patent Citations

  • Cross-working-condition reinforcement learning fault diagnosis method

    CN117668488A

  • Wind turbine generator fault diagnosis method and system based on fractional order neural network

    CN119288783A