Renewable energy equipment fault intelligent diagnosis method based on deep learning
By constructing the combination of WSGD-FPINNs and segmented Poisson regression models, the problems of difficult mechanism modeling and poor environmental adaptability in the fault diagnosis of renewable energy equipment are solved, efficient and accurate fault identification and diagnosis are achieved, and the operation stability and prediction capabilities of the equipment are improved.
Patent Information
- Application Number
- CN202510814092.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-18
AI Technical Summary
In the fault diagnosis of renewable energy equipment, the existing technology has difficulty in mechanism modeling, poor environmental adaptability, and slow strategy updates, and the inability to effectively identify mutation faults. In addition, traditional methods lack modeling of nonlinear and non-integer order dynamic behaviors, resulting in inaccuracy and inefficient diagnostic accuracy.
Fractional order physical information neural networks (WSGD-FPINNs) in the weighted shifted Grünwald-Letnikov format combined with segmented Poisson regression model are constructed through dynamic statistical learning and reinforcement learning optimization, and fault probability prediction is carried out in combination with physical constraints and environment variables, and strategy parameters are optimized through relative reward regression mechanism.
It improves the fitting accuracy and robustness of complex non-integer order dynamic behaviors, can dynamically adapt to environmental changes, improves fault recognition rate and diagnostic accuracy, and reduces the complexity and cost of strategy optimization.
Smart Images

Figure CN120336789A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of renewable energy fault diagnosis, and particularly to an intelligent fault diagnosis method for renewable energy devices based on deep learning. Background Art
[0002] With the rapid development of renewable energy technologies, the proportion of photovoltaic, wind power and energy storage systems in the modern energy structure is increasing continuously. However, during long-term operation, these devices are prone to various types of faults, such as hot spots in photovoltaic modules, mechanical wear in wind turbines, battery aging degradation, etc. If not detected and handled in a timely manner, it will seriously affect the power generation efficiency and safety of the system.
[0003] Traditional fault diagnosis methods rely on manual experience and feature engineering, often requiring a large amount of labeled data and complex signal processing procedures, and having insufficient modeling capabilities for complex physical processes with nonlinear and non-integer order. Therefore, researchers have tried to introduce deep learning methods into the field of fault diagnosis. In particular, the physics-informed neural network (PINNs) that embeds physical priors into deep models provides an effective way to solve the problems of small samples and weak supervision. However, standard PINNs still have limitations when facing the fractional-order dynamic behaviors (such as memory effects, dissipative delays, etc.) commonly existing in device operation and cannot effectively model non-integer order control characteristics. Traditional statistical models ignore physical mechanisms and have a lag in response to sudden faults (such as photovoltaic hot spots).
[0004] Meanwhile, the identification of device faults often involves the extraction and modeling of multi-scale and segmented features. How to accurately characterize the structural changes in data has become a major challenge for the diagnosis system. In addition, with the wide application of large models in sequence tasks, how to effectively evaluate the quality of generated outputs also urgently requires a more refined modeling mechanism. Moreover, when optimizing the generation model strategy, traditional reinforcement learning methods have problems such as low sample efficiency and poor convergence, lacking an efficient, theoretically robust and easy-to-deploy intelligent fault diagnosis method for renewable energy devices. Summary of the Invention
[0005] The present invention aims to solve the above problems and proposes an intelligent fault diagnosis method for renewable energy devices based on deep learning. Through fractional-order physical modeling, dynamic statistical learning and closed-loop optimization of reinforcement learning, it solves the technical problems such as difficult mechanism modeling, poor environmental adaptability, and slow strategy update in the fault diagnosis of renewable energy devices, and further overcomes the technical problems of the prior art, such as lag in response to sudden faults (such as photovoltaic hot spots), low sample efficiency in strategy optimization, unstable convergence, and low fault recognition rate.
[0006] To achieve the above object, it is realized through the following technical solutions: An intelligent fault diagnosis method for renewable energy equipment based on deep learning, comprising the following steps: S1: Collect sensor data of a photovoltaic system, a battery energy storage system, and a wind turbine generator set, and after data preprocessing, construct a training sample set with fault labels; S2: Construct a fractional-order physics-informed neural network in the weighted shifted Grünwald-Letnikov format, approximate the fractional-order derivative term through discretization, and introduce a physical constraint loss function to output a physically-driven device fault probability prediction value ; S3: Establish a piecewise Poisson regression model, with the time variable and environmental dummy variables as independent variables, and perform dynamic trend modeling on the device fault probability prediction value to output a statistically corrected fault probability ; S4: Adopt a weighted fusion strategy to fuse the device fault probability prediction value and the statistically corrected fault probability to generate a final fault probability ; S5: Based on a reinforcement learning method with a relative reward regression mechanism, use and environmental context features as inputs, and optimize the prediction strategy parameters by minimizing the relative reward difference loss function.
[0007] Furthermore, the construction of the fractional-order physics-informed neural network in the weighted shifted Grünwald-Letnikov format includes: Establish a fractional-order control equation; Adopt a deep neural network-like state response variable, and approximate the fractional-order derivative through the weighted shifted Grünwald-Letnikov coefficient; Design a comprehensive loss function , where is a physical residual loss term, is an observation data error term, is an initial boundary condition loss term; : Loss term weighting coefficient, used to balance the influence of different error terms.
[0008] Furthermore, the construction of the fractional-order physics-informed neural network in the weighted shifted Grünwald-Letnikov format also includes: Optimize the fault parameter θ through a parameter inversion mechanism, specifically including: The fault parameter θ is used as the output of the reinforcement learning action; a reward function is constructed based on the neural network prediction error and physical consistency; an improved deep reinforcement learning method is used to optimize the policy of the fault parameter θ, and the agent continuously adjusts the parameters through trial-and-error learning to minimize the prediction error, realizing the accurate inversion of the renewable energy equipment fault and outputting the optimal fault parameter after inversion. , the optimal network weight parameters and the predicted values of the state variables , where x is the spatial position variable and t is the time variable.
[0009] Furthermore, among them, the predicted value of the equipment failure probability is obtained through the non-linear transformation of the network output state response variable , that is, a normalization mapping is realized through a Sigmoid function:
[0010] Among them, : represents the predicted value of the failure probability of a certain device; : the state variable finally output by the neural network, which is the response of the trained network under the inverted parameters; : the mapping function from the state response to the failure risk score; : the Sigmoid function, used to normalize the mapping result to the interval.
[0011] Furthermore, among them, the expression of the piecewise Poisson regression model is:
[0012] Among them, is the device failure probability output by the piecewise Poisson regression model; T represents the observation time; represents the time node set in the model, that is, the turning point of the failure occurrence trend change; and respectively represent the dummy variables of three external operating environments: high temperature, strong wind, and high humidity; is the regression coefficient to be estimated; is the random error term, subject to the assumption of Poisson distribution.
[0013] Furthermore, among them, in S4: a weighted fusion strategy is adopted to fuse the predicted value of the equipment failure probability and the statistically corrected failure probability to generate the final failure probability :
[0014] Among them, and are the weight coefficients of the prediction results of two models respectively, satisfying + = 1; and, the weight coefficients and are determined by any of the following methods: dynamically adjusted based on historical prediction accuracy; allocated according to the model confidence interval; adaptively optimized using cross-validation.
[0015] Furthermore, the reinforcement learning method based on the relative reward regression mechanism realizes policy optimization by minimizing the following loss function:
[0016] Among them, : The policy parameters to be optimized currently, representing the weight set of the policy model; : The policy parameters at the n-th iteration, which are the reference benchmarks of the current policy; : The new policy parameters updated after this optimization; : The policy parameter space, representing the set of all possible parameter values; : The data sample triple, representing the state-response pair collected in the -th round, : The state, including sensor observations, environmental parameters, and historical prediction results during device operation; : An action generated under the current policy, including control instructions and maintenance suggestions; : Another action for comparison, used to construct the "relative reward" difference; : The training data set collected in the n-th round, that is, the set of state-response pairs; : Under the state , the probability value of the action output by the current policy parameter , representing the policy's preference for the action; : Under the state , the probability of the action output by the previous iteration policy , used to measure the difference between the new and old policies; : The immediate reward obtained by executing the action under the state ; : The learning rate of policy update, controlling the step size of policy iteration.
[0017] Furthermore, the method further includes: S6: Probability model the policy actions through an output token probability estimation method to generate executable maintenance instructions or scheduling plans.
[0018] Further, in the output token probability estimation method, the following probability modeling formula is adopted:
[0019] where, is the complete action token sequence, which is composed of the intermediate idea token and the execution action token concatenated together, ρ is a scaling factor and 0.2 ≤ ρ ≤ 0.5; is the current environmental input, serving as the background information for generating actions; is the context state information of the system at time , including the current observation state, historical interactions, and system feedback.
[0020] Further, in the S1, the photovoltaic system sensor data includes output voltage, current, component temperature, and environmental temperature and humidity; the battery energy storage system data includes cell voltage, battery pack temperature, current fluctuation, SOC, and SOH; the wind turbine generator data includes wind blade speed, vibration signal, main shaft temperature, current, voltage, and wind speed; where, the data preprocessing operations include: using band-pass filtering or wavelet denoising to remove the noise of current, voltage, and vibration signals; removing outliers through Z-score analysis and sliding median method; slicing the time series signal according to a fixed duration to construct a sample window, and associating with the fault type label.
[0021] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention constructs a weighted shifted Grünwald-Letnikov fractional-order physics-informed neural network (WSGD-FPINNs). By introducing a weighted discrete format to accurately approximate the fractional-order derivative, the physics-informed neural network has higher fitting accuracy and robustness in modeling complex non-integer-order dynamic behaviors such as equipment aging and energy dissipation, and improves the ability to identify implicit fault features.
[0022] 2. The present invention constructs a piecewise Poisson regression model. This model can automatically identify the structural change points in the operation data, is suitable for dynamic modeling of the change of fault occurrence frequency with time or equipment state, helps to discover non-stationary failure trends and mutation signal features, and dynamically adapts to non-stationary environments.
[0023] 3. The present invention proposes a reinforcement learning method based on a relative reward regression mechanism, which transforms the policy optimization problem into a least squares regression problem of relative rewards, avoids the dependence on the partition function, significantly improves the training efficiency and policy stability, and is particularly suitable for policy fine-tuning in sequence generation tasks.
[0024] 4. The present invention proposes an output token probability estimation method, which can perform probability modeling on the tokens output by a language model or a sequence generation model, finely characterize the uncertainty of the output, so as to realize the dynamic evaluation and error correction of the model output results, and improve the credibility of the generated diagnostic suggestions or prediction results.
[0025] It should be understood that the content described in the summary of the invention is not intended to limit the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages and aspects of the embodiments of the present invention will become more obvious. The drawings are used to better understand the solution and do not constitute a limitation to the present invention. In the drawings, the same or similar reference numerals represent the same or similar elements, where: Figure 1 is a schematic flowchart of a method for intelligent fault diagnosis of renewable energy equipment based on deep learning according to an embodiment of the present invention; Figure 2 is a schematic flowchart of a method for intelligent fault diagnosis of renewable energy equipment based on deep learning according to another embodiment of the present invention; Figure 3 is a schematic architecture diagram of a method for intelligent fault diagnosis of renewable energy equipment based on deep learning according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0028] In addition, the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0029] Figure 1 Shows the schematic flow chart of a method for intelligent fault diagnosis of renewable energy equipment based on deep learning according to an embodiment of the present invention; Figure 3 Is the schematic architecture diagram of a method for intelligent fault diagnosis of renewable energy equipment based on deep learning according to an embodiment of the present invention. As Figure 1 Shown, a method for intelligent fault diagnosis of renewable energy equipment based on deep learning includes the following steps: S1: Collect sensor data of photovoltaic systems, battery energy storage systems and wind turbine generators, and after data preprocessing, construct a training sample set with fault labels; This step S1 serves as the basis of the entire fault diagnosis system, aiming to construct a structured, high-quality data sample set that can be used for deep model training. By integrating the operation data from different types of renewable energy equipment, it comprehensively reflects their physical states and operating conditions, supporting subsequent feature modeling and classification tasks. Further, this step 1 includes the following steps: S11: Sensor data collection The present invention collects data sources from the following typical renewable energy equipment: Photovoltaic system: Collect parameters such as its output voltage, current, component temperature, ambient temperature and humidity, etc.; Battery energy storage system: including cell voltage, battery pack temperature, current fluctuation, state of charge (SOC), state of health (SOH), etc.; Wind turbine generator: Collect operating parameters such as wind blade rotation speed, vibration signal, main shaft temperature, current, voltage, wind speed, etc.; Among them, the data preprocessing operations include: using band-pass filtering or wavelet denoising to remove current, voltage and vibration signal noise; removing outliers through Z-score analysis and sliding median method; slicing the time series signal according to a fixed duration to construct a sample window and associating the fault type label.
[0030] S12: Data preprocessing To ensure the consistency of sensor data, the system introduces a unified timestamp mechanism to perform time series alignment and missing value filling on various types of collected data. In addition, the following preprocessing operations are also carried out: Noise filtering: Use algorithms such as band-pass filtering and wavelet denoising to remove random noise in current, voltage and vibration signals; Outlier detection and removal: Use Z-score analysis and sliding median method to remove non-physical outliers (such as measurement mutations caused by instrument failures); Constructing a sample set by sliding window slicing: Slice the continuous time series signal according to a fixed length to construct a sample window.
[0031] Label and sample arrangement: Use known equipment fault events as labels for training samples, and label the fault type or normal state corresponding to each window.
[0032] S2: Construct a fractional-order physics-informed neural network in the weighted shifted Grünwald-Letnikov form. Approximate the fractional-order derivative term through discretization and introduce a physical constraint loss function to output the physically driven device fault probability prediction value. ; This step S2 specifically includes the following steps: S2.1: Construct a fractional-order physics-informed neural network in the weighted shifted Grünwald-Letnikov form (WSGD-FPINNs). S2.1.1: Problem modeling Establish a fractional-order control equation based on the response data of the actual equipment. The equation form can be:
[0033] : State response variables (such as voltage, current, temperature, vibration intensity, etc.), which are the target physical quantities to be predicted. : The fractional-order derivative of , used to describe the historical memory effect and complex degradation behavior inside the equipment. : Fractional-order, , reflecting the non-local time response characteristics of equipment degradation. : Spatial position variable, which can represent equipment number, sensor position, component area, etc. : Time variable, used to model the evolution process of faults over time. : Fault parameter vector to be identified (such as fault location, type, severity level, etc.). : Control equation containing equipment physical constraints, which can describe coupled physical mechanisms such as heat, electricity, and mechanics. For example: (1) Thermal spot fault model of photovoltaic system (thermal-electric coupling) Control equation:
[0034] Among them, : Photovoltaic module temperature (state variable); : 1.5-order fractional derivative (α = 1.5), describing the non-local thermal memory effect of silicon material. : Fault parameter (R: local resistance increase caused by thermal spot, : additional heat source due to occlusion); κ is the effective thermal conductivity, and h is the convective heat transfer coefficient. : Joule heat term (strongly related to the fault resistance R); ρ is the density of photovoltaic material. is the specific heat capacity of the material. is the ambient temperature; σ is the conductivity; I is the working current; is the fault resistance.
[0035] In the hot spot area, abnormal Joule heat is generated due to the increased resistance, and the fractional-order operator can more accurately characterize the superdiffusion behavior of heat in inhomogeneous materials (the traditional integer-order model underestimates the heat accumulation rate). The fault parameter , when R rises, it indicates an abnormal increase in resistance (a sign of hot spots); >0 indicates the existence of an additional occluding heat source.
[0036] (2) Battery energy storage system aging model (electrochemical-thermal coupling) A. Kinetic equation for active lithium loss
[0037] B. State of charge (SOC) evolution equation
[0038] Among them, : The amount of active lithium loss (key state variable); : Fractional derivative (α1 = 1.2, α2 = 0.8), describing the sublinear kinetics of lithium-ion diffusion; : Fault parameter ( : Capacity attenuation coefficient, : Activation energy barrier increase, : Side reaction rate); L: The amount of active lithium loss (positively correlated with aging); I is the charge and discharge current; is the rated capacity; is the lithium loss-capacity attenuation coefficient; The fractional order α < 1 reflects the slow kinetics process of ion diffusion (aging causes diffusion hindrance), and α > 1 characterizes the delayed response of charge transfer. In the fault parameter , When it rises, it indicates an acceleration of side reactions (electrolyte decomposition); When it drops, it indicates a decrease in the ion migration barrier (electrode damage); When it drops, it indicates capacity attenuation (a decrease in the state of health SOH).
[0039] (3) Wind turbine bearing wear model (mechanical vibration) Control equation:
[0040] Among them, : Bearing vibration displacement (state variable); : Fractional derivative (α = 1.6), simulating the historical dependence of friction damping; : Fault parameters ( : Defect additional impact force, : Defect location); δ(·): Dirac function (to locate local defects); m is the equivalent mass; y is the vibration displacement; c is the damping coefficient; k is the stiffness coefficient; is the base excitation force; w is the rotational angular frequency; is the defect impact force.
[0041] Fractional-order damping term can more accurately describe the non-linear frictional dissipation (the traditional integer-order model cannot capture the energy decay lag under high-frequency impacts). In the fault parameters , > 0 indicates the existence of local impact forces (wear / cracks), and can accurately locate the defects.
[0042] S2.1.2: Network structure design Construct a deep neural network as an approximation of the state variables, where are the network weight parameters. Use automatic differentiation to calculate the integer-order derivative and combine the weighted shifted Grünwald-Letnikov (WSGD) method to approximate the fractional-order derivative:
[0043] H: Spatial discretization step size; N: Number of truncation terms, used to control the accuracy and computational cost of the derivative approximation; : Weighted Grünwald-Letnikov coefficient, specifically defined as follows:
[0044] : Grünwald coefficient, defined as:
[0045] S2.1.3: Loss function design Define a comprehensive loss function including physical residual loss, boundary condition loss, and observed data error term:
[0046] : Total loss function, the optimization objective for neural network training; : Physical residual loss, measuring the fitting error of the prediction result to the fractional-order physical equation; : The difference from the actual fault monitoring data (such as sensor data); : Initial boundary condition loss (such as the initial state of equipment startup, boundary temperature, or voltage, etc.); : The loss term weighting coefficient is used to balance the influence of different error terms.
[0047] S2.1.4: Introduce a parameter inversion mechanism Take the fault parameter as the action output of the agent. Define the state space as the neural network prediction error and response behavior. Design a reward mechanism based on the loss function and physical consistency. Use an improved deep reinforcement learning method to optimize the policy for the fault parameter θ. The agent continuously adjusts the parameters through trial-and-error learning to achieve accurate inversion of the renewable energy device fault and output the inverted parameters , the optimal network weight parameters and the predicted values of the state variables . This method does not rely on the differentiability of the objective function and is applicable to complex, multi-source physical models.
[0048] S2.1.5: Output results The probability value of the device having a fault is obtained through non-linear transformation of the network output state response variable , that is, through a Sigmoid function to achieve normalized mapping:
[0049] : Represents the predicted fault probability value of a certain device; : The state variable finally output by the neural network (such as temperature, voltage, vibration intensity, etc.), which is the response of the trained network under the inverted parameters; : The mapping function from the state response to the fault risk score, usually a linear combination or a fully connected layer, for example: , where u is the input state vector, for example, in the battery scenario: u = [voltage fluctuation, temperature gradient, lithium loss]; w is the weight vector, reflecting the contribution degree of each physical quantity to the fault; b is the bias term, the reference threshold for fault judgment.
[0050] : The Sigmoid function is used to normalize the mapping result to the interval, so as to represent it as a probability:
[0051] S3: Establish a piecewise Poisson regression model, with the time variable and the environmental dummy variable as independent variables, to perform dynamic trend modeling on the predicted value of the device fault probability and output the statistically corrected fault probability ; In this step S3, to further characterize the dynamic trend of the failure probability varying with time and environment, based on the physically-driven prediction results, a piecewise Poisson regression model is introduced to model the predicted probability. This model takes the time variable as the main independent variable, the predicted probability value as the dependent variable, and introduces multiple environmental conditions as dummy variables for control, enhancing the adaptability of the model to external disturbances.
[0052] A piecewise Poisson regression model is established, and its mathematical expression is as follows:
[0053] Where, is the equipment failure probability output by the piecewise Poisson regression model; represents the observation time; represents the time node set in the model (such as the turning point of the failure occurrence trend change); and respectively represent the dummy variables of three external operating environments (such as external operating environments like high temperature, strong wind, high humidity, etc.); are the regression coefficients to be estimated; is the random error term, assuming it follows a Poisson distribution.
[0054] Through this model, the quantitative description of the change trend of the failure rate under different environmental scenarios can be realized, enhancing the sensitivity and generalization ability of the model to boundary condition changes.
[0055] S4: Adopt a weighted fusion strategy to fuse the predicted value of the equipment failure probability and the statistically corrected failure probability to generate the final failure probability ; To balance the high interpretability of the physics-based modeling and the generalization ability of the statistical learning model, in this step S4, a weighted fusion strategy is adopted to integrate the prediction results of the fractional-order physical information neural network (step S2) and the piecewise Poisson regression model (step S3). Then the finally fused predicted probability is expressed as:
[0056] Where, is the failure probability predicted by WSGD-FPINNs; is the failure probability output by the piecewise Poisson regression model; and are the weight coefficients of the prediction results of the two models respectively, satisfying + = 1.
[0057] Among them, the weight coefficients and can be set according to historical prediction accuracy, model confidence interval or expert experience, or can be adaptively optimized through cross-validation. This weighted strategy effectively integrates the deep characterization ability of the physics-driven model for degradation mechanisms and the response ability of the statistical model to multi-source interference, improving the accuracy and robustness of the overall fault prediction system.
[0058] S5: The reinforcement learning method based on the relative reward regression mechanism takes and the environmental context features as inputs, and optimizes the prediction policy parameters by minimizing the relative reward difference loss function.
[0059] To achieve the dynamic adjustment and closed-loop optimization of the prediction policy, based on the fusion of the physical modeling and statistical regression prediction results (see steps S2 - S4), a reinforcement learning method based on the relative reward regression mechanism is introduced. The purpose of this algorithm is to provide an efficient, scalable and highly conservative policy optimization path in the context of the continuously changing device operation environment and the need for continuous update of the prediction mechanism. Different from the traditional reinforcement learning that relies on the global reward signal, the reinforcement learning method based on the relative reward regression mechanism reconstructs the policy optimization problem into a supervised regression problem for relative rewards, in order to reduce the variance fluctuation in policy updates and enhance the robust response ability to prediction error changes. Combining the prediction probability values generated in step two and the environmental context features, the reinforcement learning method based on the relative reward regression mechanism conducts optimization modeling under the context bandit setting.
[0060] The reinforcement learning objective is achieved by minimizing the regression loss function in the following form:
[0061] Among them, : The policy parameters to be optimized currently, representing the weight set of the policy model; : The policy parameters at the nth iteration, which are the reference benchmarks for the current policy, and n represents the number of iteration rounds; : The new policy parameters updated after this optimization; : The policy parameter space, representing the set of all possible parameter values; : The data sample triple, representing the state-response pair collected at the nth round; : The state, such as sensor observation values, environmental parameters, historical prediction results, etc. when the device is running; : An action generated under the current policy, such as a control instruction, maintenance suggestion, etc.; : Another action for comparison, used to construct the "relative reward" difference. : The training data set collected in the nth round (a set of state and response pairs). : In state , the probability value of the action output by the current policy parameter , indicating the preference of the policy for action selection. : In state , the probability of the action output by the previous iteration policy , used to measure the difference between the new and old policies. : The immediate reward obtained by executing the action in state , which can be defined by indicators such as prediction accuracy, system stability, or execution effect. : The learning rate for policy update, controlling the step size of policy iteration.
[0062] This objective function eliminates the dependence on the non - scoreable region function and by comparing the relative values of two responses in the same state, improving the computational feasibility of the algorithm.
[0063] Figure 2 shows a schematic flow diagram of an intelligent fault diagnosis method for renewable energy equipment based on deep learning according to another embodiment of the present invention; as Figure 2 shown, the method further includes: S6: Probability - model the policy actions through an output token probability estimation method to generate executable maintenance instructions or scheduling schemes.
[0064] To achieve the closed - loop feedback between policy generation and environment interaction, an output token probability estimation method is further introduced. This method is particularly suitable for text - driven or natural - language interaction - type policy models, where the policy output is structured text, and the text actions need to be parsed into executable actions through a post - processing function to interact with the external environment, thereby obtaining reward signals and state transitions.
[0065] To estimate the action probability, the output token probability estimation method adopts the following probability modeling formula:
[0066] Wherein, is the complete action token sequence, composed of the intermediate "thought" token and the execution action token spliced together, and the splicing form is usually ; : The intermediate "thought" token in the policy generation process, reflecting the intermediate thinking or explanatory content in the policy reasoning process; : The final "execution action" token, representing the executable response given by the policy network, such as maintenance instructions, scheduling plans, etc.; is the current environmental input, serving as the background information for generating actions, such as the current operating status of renewable energy devices, sensor data, diagnostic context, etc.; is the context state information of the system at time , including the current observation state, historical interactions, system feedback, etc. : The policy network, a probability model controlled by the parameter , used to generate the conditional distribution of action tokens. ρ is a scaling factor used to adjust the importance of the "thought" part in the overall action probability modeling. ρ has a significant impact on the final reinforcement learning performance. Extreme values (such as approaching 0 or 1) will cause the policy learning to be too narrow or the distribution to be too scattered. To maintain the stability and expressiveness of policy updates, it is recommended to select .
[0067] For a method for intelligent fault diagnosis of renewable energy devices based on deep learning in the above embodiments of the present invention, the deployment and application are as follows: The overall system adopts a modular deployment architecture, with containerization and microservices as the core, to achieve rapid delivery and cross-platform migration. The deployment architecture supports elastic scaling through Kubernetes + Docker and supports isolated deployment and parallel management of policy models in a multi-tenant environment.
[0068] This system has wide cross-industry applicability. The following lists three typical deployment scenarios: Energy storage unit scheduling in the smart grid system: predicting the trend of power fluctuations, dynamically optimizing the charging and discharging strategies through reinforcement learning, balancing the load and reducing the energy consumption cost; Fault prediction and proactive maintenance of industrial equipment: fusing the stress-time physical model and monitoring data to identify potential faults in advance and generating scheduling or maintenance suggestions; Pipe network risk warning in the urban water supply system: predicting the risk of pipe bursts based on the flow-pressure distribution and historical event statistical data, and dynamically adjusting the water valve control strategy based on the reinforcement learning method of the relative reward regression mechanism.
[0069] According to the above embodiments of the present invention, by constructing WSGD-FPINNs and piecewise Poisson regression models, designing a relative reward regression mechanism to avoid partition function calculation, improving policy stability, accurately modeling the non-integer order physical process of equipment, improving the fault feature recognition rate, obtaining the fusion of physical and statistical prediction results, improving the full life cycle diagnosis accuracy, reducing the cost of predictive maintenance, through the deep coupling of physical mechanism and data-driven, supporting the unified diagnosis framework for three types of equipment, namely photovoltaic / wind power / energy storage, constructing a diagnostic chain of fractional order modeling → dynamic statistical correction → closed-loop decision optimization, effectively solving the problems of complex mechanism, time-varying environment, and lagging strategy in the fault diagnosis of renewable energy equipment, and realizing the technical leap from passive maintenance to active prediction.
[0070] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the described steps can refer to the corresponding processes in the foregoing system embodiments and will not be elaborated herein.
[0071] It should be noted that the embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0072] It should also be noted that in the embodiments of the present application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0073] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined in the embodiments of the present application can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown in the embodiments of the present application, but will be accorded the widest scope consistent with the principles and novel features disclosed in the embodiments of the present application.
Claims
1. An intelligent fault diagnosis method for renewable energy equipment based on deep learning, characterized in that, It includes the following steps: S1: Collect the sensor data of the photovoltaic system, battery energy storage system and wind turbine generator set, and after data preprocessing, construct a training sample set with fault labels; S2: Construct a fractional-order physics-informed neural network with a weighted shift Grünwald-Letnikov scheme, approximate the fractional-order derivative term by discretization, and introduce a physical constraint loss function to output the physically-driven device failure probability prediction value ; S3: Establish a segmented Poisson regression model, with the time variable and environmental dummy variables as independent variables, for the predicted value of the equipment failure probability to perform dynamic trend modeling and output the statistically corrected failure probability ; S4: Use a weighted fusion strategy for the predicted value of the device failure probability and the failure probability after statistical correction to perform fusion and generate the final failure probability ; S5: A reinforcement learning method based on a relative reward regression mechanism, using the final failure probability and environmental context features as inputs, and optimizing the prediction policy parameters by minimizing the relative reward difference loss function.
2. The method according to claim 1, wherein Among them, The construction of the weighted shifted Grünwald-Letnikov fractional-order physics-informed neural network includes: Establish a fractional-order control equation; Adopt a deep neural network-like state response variable to approximate the fractional-order derivative through the weighted shifted Grünwald-Letnikov coefficient; Design comprehensive loss function , where is the physical residual loss term, is the observation data error term, is the initial boundary condition loss term; : Loss term weighting coefficient, used to balance the influence of different error terms.
3. The method according to claim 2, wherein Among them, The construction of the weighted shifted Grünwald-Letnikov fractional-order physics-informed neural network also includes: Optimize the fault parameter θ through a parameter inversion mechanism, specifically including: The fault parameter θ is output as the reinforcement learning action; a reward function is constructed based on the neural network prediction error and physical consistency; an improved deep reinforcement learning method is used to optimize the policy of the fault parameter θ. The agent continuously adjusts the parameters through trial-and-error learning to minimize the prediction error, achieving accurate inversion of the renewable energy equipment fault and outputting the optimal fault parameter after inversion. , the optimal network weight parameter and the predicted value of the state variable , where x is the spatial position variable and t is the time variable.
4. The method according to claim 3, wherein Among them, The predicted value of the device failure probability is obtained through the non-linear transformation of the network output state response variable That is, a normalization mapping is achieved through a Sigmoid function: ; Among them, : represents the predicted value of the failure probability of a certain device; : the state variable finally output by the neural network, which is the response of the trained network under the inverted parameters; : the mapping function from the state response to the failure risk score; : the Sigmoid function, which is used to normalize the mapping result to interval.
5. The method according to claim 1, characterized in that Among them, The expression of the piecewise Poisson regression model is: ; Among them, is the equipment failure probability output by the piecewise Poisson regression model; T represents the observation time; represents the time node set in the model, that is, the turning point of the change trend of the occurrence of the failure; and respectively represent the dummy variables of three external operating environments: high temperature, strong wind, and high humidity; are the regression coefficients to be estimated; is the random error term, assuming a Poisson distribution.
6. The method according to claim 1, characterized in that, Among them, Step S4: Using a weighted fusion strategy for the predicted value of the device failure probability and the statistically corrected failure probability to perform fusion and generate the final failure probability : ; Among them, and are the weight coefficients of the prediction results of the two models respectively, satisfying + = 1; and the weight coefficients and are determined by any of the following methods: dynamically adjusted based on historical prediction accuracy; allocated according to the model confidence interval; adaptively optimized using cross-validation.
7. The method according to claim 4, wherein The reinforcement learning method based on the relative reward regression mechanism realizes policy optimization by minimizing the following loss function: ; Among them, : The policy parameters to be optimized currently, representing the weight set of the policy model; : The policy parameters at the n-th iteration, which are the reference benchmarks of the current policy; : The new policy parameters updated through this optimization; : The policy parameter space, representing the set of all possible parameter values; : The data sample triple, representing the state-response pair collected in the -th round, : The state, including sensor observations, environmental parameters, and historical prediction results during device operation; : An action generated under the current policy, including control instructions and maintenance suggestions; : Another action for comparison, used to construct the "relative reward" difference; : The training data set collected in the n-th round, that is, the set of state-response pairs; : Under the state The probability value of the action output by the current policy parameters represents the policy's preference for action selection; : Under the state The probability of the action output by the previous iteration policy is used to measure the difference between the new and old policies; : The immediate reward obtained by executing the action under the state ; : The learning rate of policy update, which controls the step size of policy iteration.
8. The method according to claim 7, wherein The method also includes: S6: Perform probability modeling on the policy actions through an output token probability estimation method to generate executable maintenance instructions or scheduling plans.
9. The method according to claim 8, wherein Among them, The output token probability estimation method adopts the following probability modeling formula: ; Among them, is the complete action token sequence, which is composed of the intermediate thought token and the execution action token concatenated together, ρ is the scaling factor and 0.2 ≤ ρ ≤ 0.5; is the current environment input, serving as the background information for generating actions; is the context state information of the system at time , including the current observation state, historical interaction, and system feedback.
10. The method according to claim 8, wherein Among them, In the S1, among them, the photovoltaic system sensor data includes output voltage, current, component temperature and ambient temperature and humidity; the battery energy storage system data includes cell voltage, battery pack temperature, current fluctuation, SOC and SOH; the wind turbine generator set data includes wind blade speed, vibration signal, main shaft temperature, current, voltage and wind speed; Among them, the data preprocessing operations include: using band-pass filtering or wavelet denoising to remove the noise of current, voltage and vibration signals; removing outliers through Z-score analysis and sliding median method; slicing the time series signal according to a fixed duration to construct a sample window, and associating with the fault type label.
Citation Information
Patent Citations
Cross-working-condition reinforcement learning fault diagnosis method
CN117668488A
Wind turbine generator fault diagnosis method and system based on fractional order neural network
CN119288783A
Fractional-order model predictive control for neurophysiological cyber-physical systems
US20210315527A1
Cited By
Explosion-proof intelligent temperature control method and system for hydrogen peroxide storage tank based on multi-mode monitoring
CN120578242A
SiC power supply optimization method and system based on reinforcement learning
CN121092946A
A SiC power supply optimization method and system based on reinforcement learning
CN121092946B
Self-adaptive evaluation method for health degree of electrolytic cell
CN121561291A
Rail transit vehicle fault diagnosis method based on machine learning
CN122130397A