Fault Diagnosis Method, System, Device and Medium of Photovoltaic Inverter Based on Improved DQN Algorithm
Through the improvement of DQN algorithm and reinforcement learning method, the problems of cumbersome data collection and unstable model training in photovoltaic inverter fault diagnosis are solved, and efficient and accurate fault diagnosis is achieved, reducing cost and time consumption.
Patent Information
- Application Number
- CN202310282846.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-21
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-03-21
AI Technical Summary
In the prior art, the photovoltaic inverter fault diagnosis method has high requirements for related physical theories and feature extraction techniques, and the data collection work is cumbersome, and the results are not universal. Traditional neural network algorithms have the problem that training results converge to local minimum values and model overfitting.
The improved DQN algorithm is adopted to collect state and environment data of the photovoltaic inverter, feature extraction is used using the CNN model, and combined with reinforcement learning methods, the loss function of the DQN algorithm is optimized, and the training set, test set and verification set are obtained using the Hold-Out data division method, and hyperparameters are adjusted to improve the convergence speed and generalization of the model.
It improves the efficiency and accuracy of photovoltaic inverter fault diagnosis, reduces computing power consumption and debugging time, reduces costs, and improves the generalization ability of the model.
Smart Images

Figure CN116401612B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of photovoltaic inverter fault diagnosis, and relates to a photovoltaic inverter fault diagnosis method, system, device and medium based on an improved DQN algorithm. Background Art
[0002] As an important component of photovoltaic power generation, due to the unique characteristics of photovoltaics, the power generation efficiency of inverters is greatly affected by changes in the external environment, and traditional fault diagnosis experience and technologies cannot be directly transferred in this field. IGBT is the main power device in photovoltaic inverters, and its state directly affects the health status of the inverters to which it belongs. Being able to instantaneously and accurately detect the fault conditions and types of IGBTs is crucial for the efficiency and effectiveness of inverter fault diagnosis.
[0003] Currently, the fault diagnosis for photovoltaics is mainly divided into three methods: based on parameter identification, based on physical signal analysis, and artificial intelligence. The methods of parameter identification and physical signal analysis have high requirements for relevant physical theories and feature extraction technologies. The data collection work required by the methods is cumbersome, and the results are not universal. In the field of artificial intelligence, current research mostly focuses on simple judgments using similarity measurement methods such as Euclidean and Mahalanobis distances, or uses shallow neural networks for feature learning. In recent years, there have also been some studies using machine learning models such as convolutional neural networks. However, although the above methods have good accuracy recognition capabilities, the input data of the network, such as Figure 1 shown, are feature vectors or pictures, and most of them have been processed manually. Many of the input pixels of the pictures are very low, which results in the algorithm requiring good prior knowledge to effectively denoise. On the other hand, traditional neural network algorithms (including but not limited to CNN models) mostly have problems such as the training results converging to local minima and model overfitting. For the CNN model, the pooling layer in the structure will also lose a large amount of valuable information, thereby causing the network to ignore the correlation between the local and the whole. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems in the prior art that the methods of parameter identification and physical signal analysis have high requirements for relevant physical theories and feature extraction technologies, the data collection work required by the methods is cumbersome, the results are not universal, and traditional neural network algorithms mostly have problems such as the training results converging to local minima and model overfitting, and to provide a photovoltaic inverter fault diagnosis method, system, device and medium based on an improved DQN algorithm.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A photovoltaic inverter fault diagnosis method based on an improved DQN algorithm, comprising:
[0007] Step 1: Collect the status and environmental data of the PV inverter, and partition the data to obtain the training set, test set, and validation set;
[0008] Step 2: With the real historical fault information as the target, input the training set as the state s into the agent from the environment for parameter initialization to obtain the weight coefficients of the state-action function;
[0009] Step 3: Extract features from the model parameters based on the CNN model;
[0010] Step 4: Based on the extracted features and the action a in the action set in the agent, obtain the state-action function Q(s, a) containing the weight coefficients;
[0011] Step 5: Determine whether the current loop count has reached the maximum loop count; if not, based on the state-action function Q(s, a) corresponding to the action a selected by the greedy strategy and the real historical fault information, obtain the loss function; and obtain the reward based on the action a, and update the environment and enter the next state S', and repeat Steps 2 to 4; until the loop count exceeds the maximum loop count, or the loss function is less than the threshold; if so, terminate the loop and output the optimized model;
[0012] Step 6: Process the optimized model based on the validation set and test set to obtain the optimized DQN model;
[0013] Step 7: Based on the optimized DQN model, obtain the fault data of the inverter.
[0014] A further improvement of the present invention lies in:
[0015] Further, partitioning the data to obtain the training set, test set, and validation set is specifically: Based on the Hold-Out data partitioning method, the data is sequentially partitioned into the training set, validation set, and test set; the ratio of the training set, validation set, and test set is 6:2:2.
[0016] Further, the threshold is set manually and the threshold is a hyperparameter.
[0017] Further, the loss function is:
[0018] LossFuction regularize =β*Q(s t , a t )+α(Q(s t , a t )-r t +γmax a Q(s′,a′)) 2 (1)
[0019] Among them, α is the learning rate, which is one of the commonly used hyperparameters in reinforcement learning. Q(s t , a t ) is the Q-value in the current state, representing the Q-value corresponding to the environment in this state and action at the current time point t; r t represents the reward feedback by the environment at the current time point t; γ is the decay function, which is one of the hyperparameters in reinforcement learning; max a Q(s′, a′) is the state-action value function value with the highest Q-value feedback among all possible actions and corresponding states in the next moment based on the current moment. β * Q(s t , a t ) is a total activity weighting term to regularize the Q-value for each state in the learning process; β is a newly added hyperparameter in this optimized DQN algorithm and is set according to the actual situation.
[0020] Further, based on the validation set and the test set, the optimized model is processed to obtain the optimized DQN model. Specifically:
[0021] The optimized model is validated based on the validation set to obtain the best-performing mode; the best-performing mode is tested through the test set to evaluate the accuracy of the selected mode, and the optimized DQN model is obtained.
[0022] The photovoltaic inverter fault diagnosis system based on the improved DQN algorithm includes:
[0023] A partitioning module, which collects the state and environmental data of the photovoltaic inverter and partitions the data to obtain a training set, a test set, and a validation set;
[0024] An initialization module, which targets the real historical fault information and inputs the training set as the state s into the agent from the environment for parameter initialization to obtain the weight coefficient of the state-action function;
[0025] A feature extraction module, which extracts features from the model parameters based on the CNN model:
[0026] A first acquisition module, which obtains the state-action function Q(s, a) containing the weight coefficient based on the extracted features and the action a in the action set in the agent;
[0027] A judgment module, which judges whether the current number of loops reaches the maximum number of loops; if not, based on the state-action function Q(s, a) corresponding to the action a selected based on the greedy strategy and the true historical fault information, obtains a loss function; and obtains a reward based on the action a, and the environment is updated and enters the next state S'; until the number of loops exceeds the maximum number of loops or the loss function is less than the threshold; if so, terminates the loop and outputs the optimized model;
[0028] An optimization processing module, which processes the optimized model based on the validation set and the test set to obtain the most optimized DQN model;
[0029] A second acquisition module, which acquires the fault data of the inverter based on the most optimized DQN model.
[0030] A terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0031] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] Based on the collected data, the present invention obtains a training set, a test set, and a validation set. Based on the reinforcement learning process, it continuously optimizes the DQN algorithm model. On the premise of not affecting or even improving the accuracy, it proposes improvements in the convergence speed of the DQN algorithm, helps projects applying or intending to apply related algorithms to improve efficiency, save computing power, and save costs. At the same time, it proposes improvements in the generalization of the DQN algorithm, shortens the parameter tuning time of this algorithm in actual applications, saves the time of debuggers, and further helps application units reduce costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0035] Figure 1 It is a schematic diagram of an IGBT fault of a certain photovoltaic inverter;
[0036] Figure 2Flowchart of the photovoltaic inverter fault diagnosis method based on the improved DQN algorithm of the present invention;
[0037] Figure 3 It is the overall flowchart;
[0038] Figure 4 It is the schematic diagram of the neural network process;
[0039] Figure 5 It is the structural diagram of the photovoltaic inverter fault diagnosis system based on the improved DQN algorithm of the present invention. Specific implementation manners
[0040] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0041] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0042] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0043] In the description of the embodiments of the present invention, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings or the orientation or positional relationship when the product of this invention is usually placed. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation of the present invention. In addition, terms such as "first", "second", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.
[0044] In addition, if the term "horizontal" appears, it does not mean that the component is required to be absolutely horizontal, but it can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and does not mean that the structure must be completely horizontal, but it can be slightly inclined.
[0045] In the description of the embodiments of the present invention, it should also be noted that unless otherwise clearly specified and limited, if the terms "set", "install", "connected", "connected" are used, they should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0046] The following further describes the present invention in detail with reference to the drawings:
[0047] See Figure 2 , the present invention discloses a photovoltaic inverter fault diagnosis method based on an improved DQN algorithm, including:
[0048] S101: Collect the state and environmental data of the photovoltaic inverter, and divide the data to obtain a training set, a test set, and a validation set.
[0049] Based on the Hold-Out data division method, the data is sequentially divided into a training set, a validation set, and a test set; the ratio of the training set, the validation set, and the test set is 6:2:2.
[0050] S102: Taking the real historical fault information as the target, input the training set as the state s into the intelligent agent from the environment for parameter initialization to obtain the weight coefficient of the state-action function;
[0051] S103: Extract features from the model parameters based on the CNN model;
[0052] S104: Based on the extracted features and the action a in the action set in the intelligent agent, obtain the state-action function Q(s, a) including the weight coefficient;
[0053] S105: Determine whether the current number of loops has reached the maximum number of loops; if not, based on the state-action function Q(s, a) corresponding to the action a selected based on the greedy strategy and the real historical fault information, obtain the loss function; and obtain the reward based on the action a, and the environment is updated and enters the next state S', repeat S102 to S104; until the number of loops exceeds the maximum number of loops, or the loss function is less than the threshold; if so, terminate the loop and output the optimized model. The threshold is set manually and the threshold is a hyperparameter.
[0054] The loss function is:
[0055] LossFuction regularize = β * Q(s t , a t ) + α(Q(st , a t ) - r t +γmax a Q(S′, a′)) 2 (1)
[0056] Among them, α is the learning rate, which is one of the commonly used hyperparameters in reinforcement learning. Q(s t , a t ) is the Q value in the current state, representing the Q value corresponding to the environment in this state and action at the current time point t; r t represents the reward feedback by the environment at the current time point t; γ is the decay function, which is one of the hyperparameters in reinforcement learning; max a Q(s′, a′) is the state-action value function value with the highest Q value feedback among all possible actions and corresponding states in the next moment based on the current moment. β * Q(s t , a t ) is a total activity weighting term to regularize the Q value for each state in the learning process; β is the newly added hyperparameter in this optimized DQN algorithm and is set according to the actual situation.
[0057] S106: Based on the validation set and the test set, process the optimized model to obtain the optimal DQN model;
[0058] Validate the optimized model based on the validation set to obtain the best-performing model; test the best-performing model through the test set to evaluate the accuracy of the selected model and obtain the optimal DQN model.
[0059] S107: Based on the optimal DQN model, obtain the fault data of the inverter.
[0060] Example:
[0061] To solve the drawbacks of the existing machine learning methods for IGBT fault diagnosis in photovoltaic inverters, based on the DQN method, conduct fault learning and judgment. By combining the environment of reinforcement learning with the actual historical observation data of the power station and the data of each sensor, build unique long-term environmental parameters that conform to the photovoltaic power station, enabling the intelligent agent to fully consider the impact of external environment and internal environment such as power station equipment on the inverter and IGBT components when making decisions.
[0062] In the actual use stage, use an external device to directly read the sensor data of the IGBT inside the inverter, and it is applicable to the built power stations as a detachable external device. This device will have a high-speed data transmission function and wirelessly transmit the data through an encrypted channel to the remote reinforcement learning system in this application document for calculation. See Figure 3, this remote reinforcement learning system stores the built environment, environment parameters that match the station, and historical fault information of the station, and has completed the preliminary learning. At the same time, this system can also read other sensor data information transmitted to the centralized control room. After the system completes learning, it synchronizes the learning results with the centralized control side of the station and corrects the model according to the actual results.
[0063] See Figure 4 , in the preliminary learning stage, after the remote system included in this method obtains the historical real numerical data of the station, it will use it as the state (s) to input into the agent. The weight coefficients of the state-action function are initialized in the agent, and a three-layer CNN model is used for parameter feature extraction. According to the state-action function Q(s,a), the agent takes an action (a) based on the ε-greedy strategy to judge the IGBT fault type, and obtains the corresponding reward for this action by comparing the actual fault type in the current state in the dataset, and randomly updates the relevant parameters of the feature extraction network.
[0064] The setting of the reward in this method is based on the accuracy of the agent's selection result. For whether the selected fault type can correspond to the dataset label, it is divided into two types: the reward is incremented by one and the reward remains unchanged.
[0065] In addition to changing the photovoltaic scenario from the physical environment to the digital environment, defining the input, output content, and data types corresponding to the environment, state, action, and reward of fault diagnosis applications in the photovoltaic or power generation field scenarios under reinforcement learning, so that the reinforcement learning method can be applied, this patent also optimizes the loss function of the DQN algorithm used. The following is the loss function commonly used in DQN:
[0066] LossFuction DQN =α(Q(s t , a t ) - r t +γmax a Q(s′, a′)) 2 (1)
[0067] Among them, α is the learning rate, which is one of the common hyperparameters in reinforcement learning. Q(s t , a t ) is the state-action value function, representing the Q value corresponding to the environment in this state and action at the current time point t; r t represents the reward feedback by the environment at the current time point t; γ is the decay function, which is one of the hyperparameters in reinforcement learning; max a Q(s′, a′) is the state-action value function value with the highest Q value feedback among all possible actions and corresponding states in the next moment based on the current moment. The improvement made by this patent based on the above formula is shown in the following formula.
[0068] Loss Function regularize = β * Q(s t , a t ) + α(Q(s t , a t ) - r t + γ max a Q(s′, a′)) 2 (2)
[0069] Among them, β * Q(s t , a t ) is a weighted term that is always active to regularize the Q value at each state during the learning process. Q(s t , a t ) is the Q value in the current state. β is a hyperparameter newly added to this optimized DQN algorithm and can be set according to the actual situation. To ensure that the Q value is successfully regularized during the calculation, this hyperparameter should be less than 1, usually starting to adjust the parameter from below 0.5. This optimized algorithm modifies the original loss function of the DQN algorithm to make the Q value lower, thereby assisting in accelerating the convergence speed of the algorithm and increasing the robustness of the algorithm under the condition that the learning rate and step size remain unchanged.
[0070] Specific steps for the preliminary preparation work of the algorithm:
[0071] 1. The centralized control side sorts out auxiliary parameters such as weather and operation data at the same moment and historical fault information of the station according to time.
[0072] 2. Import parameters required for building the environment such as sunshine duration and air pressure into the remote system.
[0073] 3. Using the real historical fault information of the station as the target, the agent tries to learn.
[0074] 4. Give rewards to the agent according to the learning results, and then iterate the algorithm.
[0075] 5. Complete the learning and have a relatively stable algorithm in the current environment.
[0076] Specific steps for implementing the learning method:
[0077] 1. Existing sensors collect data on the equipment side of the station.
[0078] 2. The production data is locally transmitted to the network-connected external device beside the sensor.
[0079] 3. The external device transmits the data to the remote server through the Internet.
[0080] 4. The data enters the system equipped with the DQN algorithm in the server for deep reinforcement learning.
[0081] 5. After the learning is completed, the system sends the learning results to the large control screen of the station yard.
[0082] 6. The station yard staff verifies the learning results according to the actual results on the centralized control side.
[0083] 7. The station yard returns the verification results to the remote system.
[0084] 8. The system automatically receives the verification results.
[0085] 9. The system corrects the loss function in the algorithm according to the learning results.
[0086] See Figure 5 , this invention discloses a photovoltaic inverter fault diagnosis system based on an improved DQN algorithm, including:
[0087] A partitioning module, which collects the status and environmental data of the photovoltaic inverter, and partitions the data to obtain a training set, a test set, and a validation set;
[0088] An initialization module, which takes the real historical fault information as the target, inputs the training set as the state s into the agent from the environment for parameter initialization, and obtains the weight coefficients of the state-action function;
[0089] A feature extraction module, which extracts features from the model parameters based on the CNN model;
[0090] A first acquisition module, which obtains the state-action function Q(s, a) containing the weight coefficients based on the extracted features and the action a in the action set in the agent;
[0091] A judgment module, which judges whether the current number of loops has reached the maximum number of loops; if not, based on the state-action function Q(s, a) corresponding to the action a selected based on the greedy strategy and the real historical fault information, obtains the loss function; and obtains the reward based on the action a, and the environment is updated and enters the next state S'; until the number of loops exceeds the maximum number of loops, or the loss function is less than the threshold; if so, terminates the loop and outputs the optimized model;
[0092] An optimization processing module, which processes the optimized model based on the validation set and the test set to obtain the most optimized DQN model;
[0093] A second acquisition module, which obtains the fault data of the inverter based on the most optimized DQN model.
[0094] The terminal device provided by an embodiment of the present invention. The terminal device of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor executes the computer program, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0095] The computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention.
[0096] The terminal device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory.
[0097] The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0098] The memory can be used to store the computer program and / or module, and the processor realizes various functions of the terminal device by running or executing the computer program and / or module stored in the memory, and by calling the data stored in the memory.
[0099] If the modules / units integrated in the terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0100] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A photovoltaic inverter fault diagnosis method based on an improved DQN algorithm, characterized in that Including: Step 1: Collect the status and environmental data of the photovoltaic inverter, and divide the data to obtain a training set, a test set, and a validation set; Step 2: Taking the real historical fault information as the target, input the training set as the state s into the agent from the environment for parameter initialization to obtain the weight coefficient of the state-action function; Step 3: Extract features from the model parameters based on the CNN model; Step 4: Based on the extracted features and the action a in the action set of the agent, obtain the state-action function Q(s, a) containing the weight coefficient; Step 5: Determine whether the current loop count has reached the maximum loop count; if not, based on the state-action function Q(s, a) corresponding to the action a selected by the greedy strategy and the real historical fault information, obtain the loss function; and obtain the reward based on the action a, and update the environment, and enter the next state S', repeat steps 2 to 4; until the loop count exceeds the maximum loop count, or the loss function is less than the threshold; If so, terminate the loop and output the optimized model, and the loss function is: Among them, is the learning rate, which is one of the commonly used hyperparameters in reinforcement learning, is the Q value in the current state, representing the Q value corresponding to the environment in this state and action at the current time point t; represents the reward feedback by the environment at the current time point t; is the decay function, which is one of the hyperparameters in reinforcement learning; is the state-action value function value with the highest Q value feedback among all possible actions and corresponding states in the next moment based on the current moment, is a total activity weighting term to regularize the Q value for each state in the learning process; is a newly added hyperparameter in this optimized DQN algorithm and is set according to the actual situation; Step 6: Process the optimized model based on the validation set and the test set to obtain the optimized DQN model; Step 7: Based on the optimized DQN model, obtain the fault data of the inverter.
2. The photovoltaic inverter fault diagnosis method based on the improved DQN algorithm according to claim 1, wherein, The dividing the data to obtain a training set, a test set, and a validation set is specifically: based on the Hold-Out data division method, the data is sequentially divided into a training set, a validation set, and a test set; the ratio of the training set, the validation set, and the test set is 6:2:
2.
3. The photovoltaic inverter fault diagnosis method based on the improved DQN algorithm according to claim 2, characterized in that, The threshold is set manually and is a hyperparameter.
4. The photovoltaic inverter fault diagnosis method based on the improved DQN algorithm according to claim 1, wherein, The processing the optimized model based on the validation set and the test set to obtain the optimized DQN model is specifically: Validate the optimized model based on the validation set to obtain the mode with the best performance; test the mode with the best performance through the test set, evaluate the accuracy of the selected mode, and obtain the optimized DQN model.
5. A photovoltaic inverter fault diagnosis system based on an improved DQN algorithm, characterized in that, Including: A dividing module, which collects the status and environmental data of the photovoltaic inverter and divides the data to obtain a training set, a test set, and a validation set; An initialization module, which takes the real historical fault information as the target, inputs the training set as the state s into the agent from the environment for parameter initialization to obtain the weight coefficient of the state-action function; A feature extraction module, which extracts features from the model parameters based on the CNN model; A first obtaining module, which based on the extracted features and the action a in the action set of the agent, obtains the state-action function Q(s, a) containing the weight coefficient; A judgment module, which judges whether the current loop count has reached the maximum loop count; if not, based on the state-action function Q(s, a) corresponding to the action a selected by the greedy strategy and the real historical fault information, obtain the loss function; and obtain the reward based on the action a, and update the environment, and enter the next state S'; until the loop count exceeds the maximum loop count, or the loss function is less than the threshold; If so, terminate the loop and output the optimized model, and the loss function is: Among them, is the learning rate, which is one of the commonly used hyperparameters in reinforcement learning. is the Q-value in the current state, representing the Q-value corresponding to the environment in this state and action at the current time point t. represents the reward feedback by the environment at the current time point t. is the decay function, which is one of the hyperparameters in reinforcement learning. is the state-action value function value with the highest Q-value feedback among all possible actions and corresponding states in the next moment based on the current moment. is a total activity weighting term to regularize the Q-value for each state in the learning process. is a newly added hyperparameter in this optimized DQN algorithm and is set according to the actual situation. An optimization processing module, which processes the optimized model based on a validation set and a test set to obtain an optimized DQN model; A second acquisition module, which acquires fault data of an inverter based on the optimized DQN model.
6. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 1-4 are implemented.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1-4 are implemented.
Citation Information
Patent Citations
Deepfake detection method based on reinforcement learning DQN algorithm
CN113313046A
Improved DQN fault diagnosis method and system for gas turbine rotor system
CN115270867A