Wind turbine generator abnormal state detection method and system based on reinforcement learning

Through reinforcement learning methods, combined with sensor data and reward function to optimize the abnormal state detection of wind turbine units, the problem of insufficient flexibility in traditional methods to identify abnormalities in dynamic environments is solved, efficient and accurate wind turbine status monitoring is achieved, and the equipment adaptability and operational efficiency is improved.

CN120402304APending Publication Date: 2025-08-01GUANGDONG MINGYANG WIND POWER IND GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510326521.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

When facing complex and dynamic operating conditions, existing wind turbine abnormal state detection technology is difficult to quickly and accurately identify abnormalities. The traditional method is insufficient in flexibility, insufficient data accuracy and reliability, and limited model generalization capabilities.

Method used

The abnormal state detection method of wind turbine units based on reinforcement learning is adopted, and the operating state is constructed through sensor data, the reward function is designed for reinforcement learning, the Q value assignment is adjusted, and the parameters are optimized by machine learning to realize adaptive and efficient abnormal state detection.

Benefits of technology

It improves the accuracy and real-time identification of abnormal states of wind turbines, and can quickly adapt in a dynamic environment, optimize long-term benefits, reduce operation and maintenance costs, and improve equipment service life and operation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120402304A_ABST
    Figure CN120402304A_ABST
Patent Text Reader

Abstract

The invention discloses a reinforcement learning-based wind turbine generator abnormal state detection method and system. The method comprises the steps of obtaining preset sensor data; preprocessing the sensor data; constructing the operation state of the wind turbine generator according to sensor data, and performing Q value assignment of reinforcement learning on all control actions selected based on the operation state of the wind turbine generator; training a wind turbine generator abnormal state detection model based on reinforcement learning; evaluating the abnormal state detection accuracy and response speed of the wind turbine generator abnormal state detection model under different wind turbine generator operation states; according to the wind turbine generator abnormal state detection model and the obtained real-time sensor data, the state of the wind turbine generator is detected, and the abnormal state of the wind turbine generator is recognized; according to the method, the reinforcement learning method is utilized, various sensor data are fused, the abnormal state of the wind turbine generator is sensed, the operation and maintenance efficiency and safety of the wind turbine generator are effectively improved, and the feasible and intelligent method is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wind turbine state detection, and particularly relates to a method and system for detecting abnormal states of wind turbines based on reinforcement learning. Background Art

[0002] With the rapid development of renewable energy, wind power, as an important clean energy, has received increasing attention in terms of its power generation efficiency and equipment reliability. The identification and detection of abnormal states of wind turbines are one of the key technologies to ensure the safe and efficient operation of wind farms. Currently, the technical status in this field is mainly reflected in the following aspects: 1. Traditional monitoring technologies: Traditional methods for detecting abnormal states of wind turbines mainly rely on monitoring physical quantities such as vibration analysis, temperature monitoring, and oil contamination detection. These methods usually set thresholds based on experience and manually analyze data. When the parameters deviate from the normal range, the system will issue an alarm. The advantage of such methods is their relative simplicity and ease of implementation, but they also have the disadvantages of insufficient flexibility and slow response speed. Especially when facing complex and dynamic operating conditions, it is difficult to quickly and accurately identify abnormalities.

[0003] 2. Data-driven technologies: In recent years, with the progress of big data and sensor technologies, data-driven abnormal detection methods have gradually become the mainstream. These methods usually use machine learning technologies to analyze historical operation data and establish models to monitor abnormal states. Common algorithms include support vector machine (SVM), decision tree, and random forest, etc. These algorithms can learn from a large amount of data and identify potential abnormal patterns. Compared with traditional methods, the detection rate and flexibility have been significantly improved.

[0004] 3. Application of deep learning: Deep learning also occupies an important position in the detection of abnormal states of wind turbines. By using convolutional neural network (CNN) and recurrent neural network (RNN), researchers can better process complex time-series data and multi-dimensional features. Such methods can automatically extract features and significantly improve the detection accuracy and efficiency. For example, deep learning can identify subtle abnormal changes in signals such as wind speed, vibration, and noise, effectively preventing faults from occurring.

[0005] 4. Rise of reinforcement learning: In recent years, reinforcement learning, as a new machine learning method, has begun to be applied to the state identification of wind turbines. Through interaction with the environment, reinforcement learning can self-learn and optimize under dynamically changing working conditions, providing a new idea for the detection of abnormal states. This method can optimize long-term benefits and improve the operation and maintenance efficiency of wind turbines.

[0006] Overall, current wind turbine abnormality identification and detection technology has evolved from traditional empirical methods to data-driven and intelligent approaches. Despite continuous technological advancements, practical applications still face various challenges, such as data accuracy and reliability, and model generalization. Therefore, further improving the accuracy and real-time performance of detection technology remains a critical issue in this field. Summary of the Invention

[0007] The purpose of the present invention is to overcome the shortcomings of the existing technology and propose a method and system for detecting abnormal conditions of wind turbines based on reinforcement learning. By obtaining feedback through interaction with the environment, the system can quickly adapt to new data and conditions, thereby improving the accuracy and real-time performance of identifying abnormal conditions of wind turbines.

[0008] The purpose of the present invention is achieved through the following technical solution: A method for detecting abnormal state of a wind turbine generator system based on reinforcement learning, comprising the following steps: S1. Obtaining preset sensor data; S2, preprocessing the sensor data to normalize it; S3. Construct the wind turbine operating state based on the sensor data, design a reinforcement learning reward function, and assign Q values to all control actions selected based on the wind turbine operating state through reinforcement learning. S4. Enable the preset intelligent agent to select the corresponding control action according to the current operating state of the wind turbine and update the corresponding Q value assignment. Repeat this step until the preset number of training steps N is reached, completing the training of the wind turbine abnormal state detection model based on reinforcement learning; S5. Evaluate the abnormal state detection accuracy and response speed of the wind turbine abnormal state detection model under different wind turbine operating conditions, and adjust the reward function, learning rate, and discount factor of the reinforcement learning based on the evaluation results to obtain a constructed wind turbine abnormal state detection model; S6. Detect the status of the wind turbine generator set based on the abnormal status detection model of the wind turbine generator set and the acquired real-time sensor data to identify its abnormal status.

[0009] Furthermore, the step S1 includes: The sensors include a vibration sensor, a temperature sensor, a wind speed sensor, a current and voltage sensor, and a noise reduction sensor.

[0010] Furthermore, the step S1 includes: The sensor communication is connected to a sensor anomaly detection system to identify sensor faults and monitor sensor data anomalies. If an anomaly is found in the sensor data, the sensor is replaced.

[0011] Further, the step S3 includes: Construct the operating state of the wind turbine according to the sensor data. The representation of the operating state includes sensor data, the timestamp of the state, and a snapshot of historical data. The sensor data includes vibration, temperature, wind speed, current, and voltage.

[0012] Furthermore, step S3 includes: Design the reward function of reinforcement learning. The reward function of reinforcement learning includes positive rewards and negative rewards, and assign Q-values for reinforcement learning to all control actions selected based on the operating state of the wind turbine. The control actions that result in positive rewards include normal operation, successfully identifying an anomaly and taking correct measures. Assign a Q-value of +1 to the control actions for normal operation, and assign a Q-value of +10 to the control actions for successfully identifying an anomaly and taking correct measures. The control actions that result in negative rewards include abnormal vibration, abnormal temperature, and overload anomaly. Assign a Q-value of -10 to the control actions for abnormal vibration, assign a Q-value of -15 to the control actions for abnormal temperature, and assign a Q-value of -20 to the control actions for overload anomaly.

[0013] Furthermore, step S5 includes: S501. Define evaluation metrics: First, clarify the key metrics for evaluating the performance of the evaluation model, including: Accuracy: The ratio of the number of detected abnormal states to the total number of abnormal states; Response speed: The time required from the occurrence of an abnormal state to the detection of the anomaly; S502. Data collection and analysis: Collect the accuracy and response speed data of the model detecting abnormal states under different operating states of the wind turbine; analyze the data to identify the performance of the model; S503. According to the evaluation results, formulate a parameter adjustment strategy, including: a. Adjust the reward function: Increase the reward value: In the states where the model performs poorly, increase the value of the reward function to encourage the model to pay more attention to these states where the model performs poorly; Dynamic reward adjustment: Dynamically adjust the reward function according to the performance of the model in different states to enable the model to develop evenly in each state; b. Adjust the learning rate: Decrease the learning rate: If the model performs unstably in a certain state, appropriately decrease the learning rate to enable the model to adjust the strategy more carefully; State-based learning rate: Set different learning rates for different states to enable the model to learn more carefully in key states; c. Adjust the discount factor, and the value range of the discount factor is from 0 to 1: Increase the discount factor: In the state of slow reaction speed, increase the discount factor to make the model pay more attention to future rewards, so as to adjust the strategy faster; Dynamic discount factor: Dynamically adjust the discount factor according to the real-time performance of the model to optimize the emphasis on future rewards; S504. The design of the automatic adjustment mechanism includes: Feedback loop: After the model detects an abnormal state, automatically adjust the parameters according to the feedback result; Machine learning optimization: Use machine learning algorithms to optimize the reward function, learning rate and discount factor; S505. The experimental verification includes: Comparison experiment: Conduct comparison experiments under different parameter settings to evaluate the performance improvement of the model; Real-time monitoring: During actual operation, monitor the accuracy and reaction speed of the model in real time to ensure the effectiveness of parameter adjustment; S506. Continuous optimization: According to the experimental results, continuously optimize the model parameters to ensure that it can maintain good detection performance under different operating states.

[0014] A wind turbine abnormal state detection system based on reinforcement learning, which is used to implement the above-mentioned wind turbine abnormal state detection method based on reinforcement learning, includes: Environment definition module, which is used to define the operating state of the wind turbine, control actions, and the Q-value assignment of the reward function; Agent design module, which is used to design an agent for operating state perception, control action selection, and learning optimization strategy; State representation module, which constructs the operating state of the wind turbine according to the sensor data; Action selection module, which selects control actions based on the operating state of the wind turbine; Reward function design module, which assigns Q-values to each control action according to the actual control action; Reinforcement learning algorithm training module, which starts from the initial state to make the agent select control actions, observes the environmental feedback, updates the Q-value or strategy, and loops until the preset termination condition is reached; Evaluation module, which is used to evaluate the performance of the agent in different environments and record the accuracy and reaction speed of abnormal state detection.

[0015] A non-transitory computer-readable medium storing instructions, when the instructions are executed by a processor, execute the steps of the above-mentioned wind turbine abnormal state detection method based on reinforcement learning.

[0016] A computing device includes a processor and a memory for storing programs executable by the processor. When the processor executes the programs stored in the memory, the above-mentioned method for detecting abnormal states of wind turbine generators based on reinforcement learning is implemented.

[0017] Compared with the prior art, the present invention has the following advantages and beneficial effects: 1. Strong adaptability: During the operation of wind turbines, there are many uncertainties and changes, including environmental factors, equipment aging, and sudden failures. Traditional monitoring and detection methods often rely on preset rules or models and are difficult to cope with these dynamic changes. However, reinforcement learning algorithms can obtain feedback through interaction with the environment, quickly adapt to new data and states, and thus improve the accuracy of abnormal state recognition. The present invention can, when faced with unusual vibration patterns of wind turbine generators, continuously learn and adjust its strategy to identify potential faults.

[0018] 2. Optimization of long-term benefits: The present invention encourages the optimization of long-term returns rather than only focusing on short-term results. In the monitoring of wind turbine generators, it not only involves identifying the current abnormal states but also the anticipation and handling of potential future faults. This strategy can effectively reduce maintenance costs and extend the service life of the equipment. By early identifying specific abnormal patterns and taking corresponding preventive measures, the overall operating efficiency of the wind turbine generators is improved, thereby providing long-term economic benefits for the operation of the wind turbine generators.

[0019] 3. Handling high-dimensional state spaces: The operating data of wind turbine generators usually comes from high-dimensional information generated by multiple sensors, such as temperature, vibration, acoustic signals, etc. Traditional monitoring methods often face challenges when dealing with multi-dimensional data. However, the present invention has good feature learning capabilities, can efficiently handle high-dimensional state spaces, extract important features, and ignore redundant information. Multi-sensor fusion enables the accurate identification and analysis of the complex operating states of wind turbine generators, thereby improving the efficiency of fault detection.

[0020] 4. Flexibility of reinforcement learning strategies: The flexibility of reinforcement learning enables it to adapt to different monitoring requirements and application scenarios. For example, the Q-Learning algorithm or policy-based algorithms in model-free reinforcement learning can be used for customized adjustment according to the specific type of wind turbine generator and the working environment. This flexibility enables the wind turbine generator state detection system to continuously optimize and update according to real-time data, ensuring higher accuracy and real-time performance.

[0021] 5. Ensemble learning and multi-task processing: The present invention can be combined with supervised learning. First, use historical data to train a preliminary model, and then continuously optimize and adjust the model through reinforcement learning. This method can significantly improve the reliability and efficiency of abnormal state detection and adapt to wind turbine systems of different scales and types.

[0022] In the identification and detection of abnormal states of wind turbines, by utilizing the adaptability, long-term optimization ability, and flexibility provided by reinforcement learning, not only can the success rate of fault detection be improved, but also the operation and maintenance costs can be reduced, achieving efficient management of equipment. These advantages make the present invention have broad application prospects in the wind power field. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a training flow chart of an abnormal state detection model for wind turbines based on reinforcement learning. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] Embodiment 1 Refer to Figure 1 As shown, the method for detecting abnormal states of wind turbines based on reinforcement learning provided in this embodiment includes the following steps: S1. Obtain various preset sensor data; the sensors include vibration sensors, temperature sensors, wind speed sensors, current and voltage sensors, and noise reduction sensors; the sensors are communicatively connected to a sensor abnormal detection system to identify faults of the sensors, monitor abnormalities in the sensor data, and replace the sensors if abnormalities in the sensor data are found.

[0025] S2. Perform preprocessing of normalizing the sensor data. S3. Construct the operating state of the wind turbine according to the sensor data, and design a reward function for reinforcement learning to assign Q values for reinforcement learning to all control actions selected based on the operating state of the wind turbine. Construct the operating state of the wind turbine according to the sensor data. The representation of the operating state includes sensor data, the time stamp of the state, and a historical data snapshot. The sensor data includes vibration, temperature, wind speed, current, and voltage.

[0026] Design a reward function for reinforcement learning. The reward function for reinforcement learning includes positive rewards and negative rewards, and assign Q values for reinforcement learning to all control actions selected based on the operating state of the wind turbine.

[0027] The control actions that result in positive rewards include normal operation, successfully identifying abnormalities and taking correct measures. Assign a Q value of +1 to the control actions for normal operation, and assign a Q value of +10 to the control actions for successfully identifying abnormalities and taking correct measures.

[0028] The control actions for negative rewards include abnormal vibration, abnormal temperature, and overload abnormality. A Q-value assignment of -10 is given to the control action for abnormal vibration, a Q-value assignment of -15 is given to the control action for abnormal temperature, and a Q-value assignment of -20 is given to the control action for overload abnormality.

[0029] S4. Enable the preset agent to select the corresponding control action according to the current operating state of the wind turbine generator set, and update the corresponding Q-value assignment. Repeat this step until the preset training step N is reached to complete the training of the abnormal state detection model for wind turbine generator sets based on reinforcement learning; S5. Evaluate the abnormal state detection accuracy and response speed of the abnormal state detection model for wind turbine generator sets under different operating states of the wind turbine generator set. Adjust the reward function, learning rate, and discount factor of the reinforcement learning according to the evaluation results to obtain the constructed abnormal state detection model for wind turbine generator sets, including: S501. Define the evaluation metrics: First, clarify the key metrics for evaluating the model performance, including: Accuracy: The ratio of the number of detected abnormal states to the total number of abnormal states; Response Time: The time required from the occurrence of an abnormal state to the detection of the abnormality.

[0030] S502. Data collection and analysis: Under different operating states of the wind turbine generator set, collect the accuracy and response speed data of the model for detecting abnormal states; analyze the data to identify the states in which the model performs well and the states in which it performs poorly.

[0031] S503. According to the evaluation results, formulate a parameter adjustment strategy: a. Adjust the reward function: Increase the reward value: In states where the model performs poorly, such as high wind speed, strong wind, etc., increase the value of the reward function to encourage the model to pay more attention to these states.

[0032] Dynamic reward adjustment: Dynamically adjust the reward function according to the performance of the model in different states to enable the model to develop evenly in each state.

[0033] b. Adjust the learning rate: Decrease the learning rate: If the model performs unstably in certain states, appropriately decrease the learning rate to enable the model to adjust the strategy more carefully and avoid performance degradation caused by radical updates.

[0034] State-specific learning rate: Set different learning rates for different states to enable the model to learn more carefully in key states.

[0035] c. Adjust the discount factor, and the value range of the discount factor is from 0 to 1; Increase the discount factor: In the state with a slower reaction speed, increase the discount factor to make the model pay more attention to future rewards, thereby adjusting the strategy faster.

[0036] Dynamic discount factor: Dynamically adjust the discount factor according to the real-time performance of the model to optimize the emphasis on future rewards.

[0037] S504. Automatic adjustment mechanism: To improve the efficiency and accuracy of adjustment, design an automatic adjustment mechanism: Feedback loop: After the model detects an abnormal state, automatically adjust the parameters according to the feedback results.

[0038] Machine learning optimization: Use machine learning algorithms, such as genetic algorithms, Bayesian optimization, etc., to optimize the reward function, learning rate, and discount factor.

[0039] S505. Experimental verification: After adjusting the parameters, it is necessary to verify the adjustment effect through experiments: Comparative experiment: Conduct comparative experiments under different parameter settings to evaluate the performance improvement of the model.

[0040] Real-time monitoring: During actual operation, monitor the accuracy and reaction speed of the model in real time to ensure the effectiveness of parameter adjustment.

[0041] S506. Continuous optimization: According to the experimental results, continuously optimize the model parameters to ensure that it can maintain good detection performance under different operating states.

[0042] S6. According to the wind turbine abnormal state detection model and the acquired real-time sensor data, detect the state of the wind turbine and identify its abnormal state.

[0043] Through the above solution, using the reinforcement learning method and integrating various types of sensor data, the abnormal state perception of the wind turbine is realized, effectively improving the operation and maintenance efficiency and safety of the wind turbine. It provides a feasible and intelligent solution for solving common abnormal states and promotes the intelligent development of the wind power industry.

[0044] Embodiment 2 The wind turbine abnormal state detection system based on reinforcement learning provided in this embodiment is used to implement the wind turbine abnormal state detection method based on reinforcement learning described in Embodiment 1, and includes: An environment definition module, used to define the operating state of the wind turbine, control actions, and the Q-value assignment of the reward function, including: State Space: The sensors include: Vibration sensor: Monitor the vibration level of the generator set; Temperature sensor: Monitor the temperature of important components such as generators and pitch systems; Wind speed sensor: Obtain wind speed and direction; Current and voltage sensors: Monitor the output performance of the generator; Noise reduction sensor: Monitor abnormal sounds.

[0045] State representation includes: Real-time data read by each sensor; Labor load status.

[0046] Action Space: Drive instruction: Change the blade angle to control the power output of the generator.

[0047] Stop or restart the generator.

[0048] Maintenance instruction: Send an alarm to the operation and maintenance team and recommend maintenance.

[0049] Reward Function: Positive rewards include: During normal operation, give a small positive reward (+1); When an abnormal state is successfully identified and addressed, such as timely adjusting the blade angle, give a larger positive reward (+10).

[0050] Negative rewards include: When there is abnormal vibration, give a larger negative reward (-10); When the temperature is too high, give a negative reward (-15); When the labor load exceeds the safety limit, give a larger negative reward (-20).

[0051] Agent design module, used to design an agent for operating state perception, control action selection, and learning and optimizing strategies, including: State perception: Read data from multiple sensors to construct the current operating state of the wind turbine.

[0052] Action selection: Select corresponding control actions based on the current operating state of the wind turbine.

[0053] Learning mechanism: Continuously optimize the strategy through interaction with the environment to improve operating safety and efficiency.

[0054] State representation module, construct the operating state of the wind turbine according to sensor data; Action selection module, select control actions based on the operating state of the wind turbine; Reward function design module, assign Q values to each control action according to the actual control action; A reinforcement learning algorithm training module sets the state and action space, starts from the initial state to enable the agent to select control actions, observes the environmental feedback, updates the Q value or policy, and loops until a preset termination condition is reached. An evaluation module is used to evaluate the performance of the agent in different environments and record the accuracy and response speed of abnormal state detection.

[0055] Embodiment 3 This embodiment discloses a non-transitory computer-readable medium storing instructions, which, when executed by a processor, perform the steps of the method for detecting abnormal states of a wind turbine based on reinforcement learning according to Embodiment 1.

[0056] The non-transitory computer-readable medium in this embodiment can be a medium such as a disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), a USB flash drive, or a mobile hard disk.

[0057] Embodiment 4 This embodiment discloses a computing device including a processor and a memory for storing programs executable by the processor. When the processor executes the programs stored in the memory, it implements the method for detecting abnormal states of a wind turbine based on reinforcement learning according to Embodiment 1.

[0058] The computing device described in this embodiment can be a desktop computer, a laptop computer, a smart phone, a PDA handheld terminal, a tablet computer, a programmable logic controller (PLC), or other terminal devices with processor functions.

[0059] The above-described embodiments are only the preferred embodiments of the present invention and do not limit the scope of implementation of the present invention. Therefore, any changes made according to the shape and principle of the present invention should be covered by the protection scope of the present invention.

Claims

1. An abnormal state detection method for wind turbines based on reinforcement learning, characterized in that, It includes the following steps: S1. Obtain various preset sensor data; S2. Perform preprocessing of normalizing the sensor data; S3. Construct the operating state of the wind turbine based on the sensor data, design a reward function for reinforcement learning, and assign Q-values for reinforcement learning to all control actions selected based on the operating state of the wind turbine; S4. Enable a preset intelligent agent to select corresponding control actions according to the current operating state of the wind turbine, and update the corresponding Q-value assignment. Repeat this step until the preset training step N is reached to complete the training of the wind turbine abnormal state detection model based on reinforcement learning; S5. Evaluate the abnormal state detection accuracy and response speed of the wind turbine abnormal state detection model under different wind turbine operating states, adjust the reward function, learning rate, and discount factor of reinforcement learning according to the evaluation results to obtain the constructed wind turbine abnormal state detection model; S6. Detect the state of the wind turbine according to the wind turbine abnormal state detection model and the obtained real-time sensor data, and identify its abnormal state.

2. The abnormal state detection method of a wind turbine based on reinforcement learning according to claim 1, characterized in that The step S1 includes: The sensors include vibration sensors, temperature sensors, wind speed sensors, current and voltage sensors, and noise reduction sensors.

3. The method for detecting abnormal states of a wind turbine based on reinforcement learning according to claim 2, wherein, The step S1 includes: The sensors are communicatively connected to a sensor abnormal detection system to identify faults of the sensors, monitor the abnormality of the sensor data, and replace the sensors if the abnormality of the sensor data is found.

4. The abnormal state detection method for a wind turbine based on reinforcement learning according to claim 1, characterized in that, The step S3 includes: Construct the operating state of the wind turbine according to the sensor data. The representation of the operating state includes sensor data, the time stamp of the state, and a historical data snapshot. The sensor data includes vibration, temperature, wind speed, current, and voltage.

5. A method for detecting abnormal states of a wind turbine based on reinforcement learning according to claim 1, characterized in that, The step S3 includes: Design a reward function for reinforcement learning. The reward function for reinforcement learning includes positive rewards and negative rewards, and assign Q-values for reinforcement learning to all control actions selected based on the operating state of the wind turbine; The control actions for which positive rewards are given include normal operation, successfully identifying abnormalities and taking correct measures. Assign a Q-value of +1 to the control actions for normal operation, and assign a Q-value of +10 to the control actions for successfully identifying abnormalities and taking correct measures; The control actions for which negative rewards are given include vibration abnormalities, temperature abnormalities, and overload abnormalities. Assign a Q-value of -10 to the control actions for vibration abnormalities, assign a Q-value of -15 to the control actions for temperature abnormalities, and assign a Q-value of -20 to the control actions for overload abnormalities.

6. The abnormal state detection method of a wind turbine based on reinforcement learning according to claim 1, characterized in that The step S5 includes: S501. Define evaluation metrics: First, clarify the key metrics for evaluating the model performance, including: Accuracy: The ratio of the number of detected abnormal states to the total number of abnormal states; Response speed: The time required from the occurrence of an abnormal state to the detection of the abnormality; S502. Data collection and analysis: Under different wind turbine operating states, collect the accuracy and response speed data of the model detecting abnormal states; analyze the data to identify the performance of the model; S503. According to the evaluation results, formulate a parameter adjustment strategy, including: a. Adjust the reward function: Increase the reward value: When the model performs poorly, increase the value of the reward function to encourage the model to pay more attention to these states where the model performs poorly; Dynamic reward adjustment: Dynamically adjust the reward function according to the model's performance in different states to enable the model to develop evenly in each state; b. Adjust the learning rate: Decrease the learning rate: If the model's performance is unstable in a certain state, appropriately decrease the learning rate to enable the model to adjust the strategy more meticulously; State-specific learning rate: Set different learning rates for different states to enable the model to learn more meticulously in critical states; c. Adjust the discount factor, where the value range of the discount factor is from 0 to 1: Increase the discount factor: In the state with a slower reaction speed, increase the discount factor to make the model pay more attention to future rewards, thereby adjusting the strategy faster; Dynamic discount factor: Dynamically adjust the discount factor according to the model's real-time performance to optimize the emphasis on future rewards; S504. The designed automatic adjustment mechanism includes: Feedback loop: Enable the model to automatically adjust parameters according to the feedback results after detecting abnormal states; Machine learning optimization: Use machine learning algorithms to optimize the reward function, learning rate, and discount factor; S505. Experimental verification includes: Comparative experiment: Conduct comparative experiments under different parameter settings to evaluate the performance improvement of the model; Real-time monitoring: During actual operation, monitor the accuracy and reaction speed of the model in real time to ensure the effectiveness of parameter adjustment; S506. Continuous optimization: Continuously optimize the model parameters according to the experimental results to ensure that it can maintain good detection performance in different operating states.

7. An abnormal state detection system for a wind turbine based on reinforcement learning, characterized in that, A device for implementing the reinforcement learning-based wind turbine abnormal state detection method according to any one of claims 1-6, includes: An environment definition module for defining the operating state of the wind turbine, control actions, and the Q-value assignment of the reward function; An agent design module for designing an agent to perform operating state perception, control action selection, and learning and optimizing strategies; A state representation module for constructing the operating state of the wind turbine based on sensor data; An action selection module for selecting control actions based on the operating state of the wind turbine; A reward function design module for performing Q-value assignment to each control action according to the actual control action; A reinforcement learning algorithm training module that enables the agent to select control actions starting from the initial state, observe the environmental feedback, update the Q-value or strategy, and loop until a preset termination condition is reached; An evaluation module for evaluating the performance of the agent in different environments and recording the accuracy and reaction speed of abnormal state detection.

8. A non-transitory computer-readable medium storing instructions, characterized in that, When the instruction is executed by a processor, perform the steps of the reinforcement learning-based wind turbine abnormal state detection method according to any one of claims 1-6.

9. A computing device, comprising a processor and a memory for storing processor-executable programs, characterized in that, When the processor executes the program stored in the memory, implement the reinforcement learning-based wind turbine abnormal state detection method according to any one of claims 1-6.