Fan SCADA data analysis method based on large language model

By introducing a large language model in the SCADA data analysis of fans, processing historical data and analyzing wind farm data in real time, the problems of low data processing efficiency and insufficient intelligence in traditional methods are solved, and efficient fault detection and optimized operation and maintenance management are achieved.

CN120145274APending Publication Date: 2025-06-13CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510319082.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The SCADA data processing method of traditional fans is inefficient, making it difficult to identify and predict fan failures in real time, and the system is not intelligent enough, making it difficult to cope with complex and changing wind farm environments.

Method used

The SCADA data analysis method of fan based on large language models is adopted, and the fan failure is identified by obtaining historical data for preprocessing, selecting machine learning models, and using real-time data for training, identifying and predicting fan failures.

Benefits of technology

It significantly improves data processing efficiency and fault detection accuracy, enhances the automation and intelligence level of the system, can automatically analyze and interpret complex data, identify and predict potential faults in real time, and optimizes the operation and maintenance management efficiency and safety of wind farms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145274A_ABST
    Figure CN120145274A_ABST
Patent Text Reader

Abstract

The invention discloses a fan SCADA data analysis method based on a large language model. The fan SCADA data analysis method comprises the following steps: 1) acquiring historical fan SCADA data; 2) preprocessing the historical fan SCADA data to obtain fan SCADA data to be analyzed; 3) selecting a machine learning model by the large language model based on to-be-analyzed fan SCADA data; 4) the big language model trains the selected machine learning model by using the to-be-analyzed fan SCADA data to obtain a fan SCADA data analysis model; and 5) acquiring real-time fan SCADA data, inputting the real-time fan SCADA data into the fan SCADA data analysis model, and identifying and / or predicting fan faults. According to the method, a large amount of complex data can be automatically analyzed and explained, potential faults and abnormal states can be identified and predicted in real time, and the operation and maintenance management efficiency and safety of the wind power plant are greatly optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis, and specifically to a method for analyzing wind turbine SCADA data based on large language models. Background Art

[0002] The workflow of data science is inherently complex, involving interrelated tasks such as data processing, feature engineering, and model training. As data and requirements continue to evolve, solving these tasks requires iterative refinement and real-time adjustment. Recent research has integrated large language models into data science tasks, leveraging their extensive knowledge and coding capabilities. These methods mainly focus on isolated tasks such as feature engineering, model selection, and hyperparameter optimization, usually operating within a fixed workflow. However, they lack an overall assessment of the end-to-end workflow, making it difficult to evaluate the complete data science process. In addition, these methods usually have difficulty handling real-time changes in intermediate data and dynamically adapting to changing task dependencies.

[0003] In the field of wind energy, wind turbine SCADA (Supervisory Control and Data Acquisition) systems have been widely used in wind farms, mainly for monitoring and collecting the operating data of turbines. Wind turbine SCADA systems generate a large amount of operating data in real time, which is usually collected in real time at high frequencies, forming a huge time series dataset that includes not only the status information of the turbines but also their operating conditions. Traditional data processing methods face challenges of low efficiency and dependence on manual intervention, and are difficult to meet the requirements of wind turbine fault prediction and maintenance optimization. Specifically, the following problems and disadvantages mainly exist:

[0004] 1. Low data processing efficiency: Traditional methods rely on manual intervention for data preprocessing and analysis, and are less efficient when dealing with large-scale data, making it difficult to achieve rapid response.

[0005] 2. Poor interpretability of on-site data systems: The data system presents the status of the wind turbine through real-time SCADA data, but it is difficult for inexperienced on-site staff to directly identify potential problems from these data.

[0006] 3. Insufficient intelligence: The existing systems have limited intelligence levels, mostly relying on preset rules and simple algorithms, lacking the ability of deep learning and adaptive adjustment, and are difficult to cope with the complex and changeable environment of wind farm operation. Summary of the Invention

[0007] The object of the present invention is to provide a method for analyzing wind turbine SCADA data based on large language models, including the following steps:

[0008] 1) Obtain historical wind turbine SCADA data;

[0009] 2) Preprocess the historical fan SCADA data to obtain the fan SCADA data to be analyzed;

[0010] 3) The large language model selects a machine learning model based on the fan SCADA data to be analyzed;

[0011] 4) The large language model uses the fan SCADA data to be analyzed to train the selected machine learning model to obtain a fan SCADA data analysis model;

[0012] 5) Obtain real-time fan SCADA data and input it into the fan SCADA data analysis model to identify and / or predict fan faults.

[0013] Furthermore, the fan SCADA data includes wind speed, power, and temperature data.

[0014] Furthermore, the fan SCADA data is obtained through sensor monitoring.

[0015] Furthermore, the steps for preprocessing the fan SCADA data include: missing value processing, outlier detection and exclusion, and time series normalization.

[0016] Furthermore, the methods for missing value processing include: interpolation filling and deleting incomplete data.

[0017] Furthermore, the outliers are extreme values and extreme outliers.

[0018] Furthermore, in step 3), if the target variable type of the fan SCADA data to be analyzed is binary, the selected machine learning model is a classification model;

[0019] If the size of the fan SCADA data to be analyzed is, or the characteristics are, the optimization algorithm adopted by the machine learning model is.

[0020] Furthermore, in step 3), if the target variable type of the fan SCADA data to be analyzed is binary, the selected classification model is used as the machine learning model; if the number of records of the fan SCADA data to be analyzed is less than 10,000 and the number of features is less than 100, the optimization algorithms selected are gradient descent, Newton's method, or L-BFGS; if the number of records of the fan SCADA data to be analyzed is greater than 10,000 or the number of features is greater than 100, the optimization algorithms selected are stochastic gradient descent (SGD), Adam, Adagrad, or RMSprop.

[0021] Further, in step 4), the large language model trains the selected machine learning model through the train_and_visualize method and evaluates the performance of the machine learning model; if the accuracy of the machine learning model after training does not reach the threshold, or the model has overfitting, underfitting, data characteristic mismatch, or excessive computational resource consumption, the hyperparameters are adjusted or the machine learning model is replaced, and training is carried out again. Further, after identifying and / or predicting the fan fault, the fan fault and the real-time fan SCADA data are also visualized.

[0022] The technical effect of the present invention is beyond doubt. The present invention proposes a method for analyzing fan SCADA data based on a large language model, which significantly improves the data processing efficiency and fault detection accuracy by integrating advanced machine learning technologies and natural language processing capabilities, while enhancing the automation and intelligence level of the system. The present invention can automatically analyze and interpret a large amount of complex data, identify and predict potential faults and abnormal states in real time, and greatly optimize the operation and maintenance management efficiency and safety of the wind farm.

[0023] Specifically, the beneficial effects of the present invention are as follows:

[0024] (1) Efficiency improvement: Compared with the traditional SCADA system, the present invention significantly reduces the need for manual intervention through the automated data preprocessing and model training processes, and greatly improves the data processing speed.

[0025] (2) Intelligence and self-adaptability: The system of the present invention can automatically optimize the model parameters according to real-time feedback, while the traditional system usually lacks such self-adaptive ability, so the present invention performs more stably and reliably in a dynamic environment.

[0026] (3) User interaction: The interactive interface provided by the present invention enables non-technical users to easily understand and operate the system, improving the user experience and operation convenience. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a framework diagram of a method for analyzing fan SCADA data according to the present invention.

[0028] Figure 2 It is a diagram of the data preprocessing result of a large language model according to the present invention.

[0029] Figure 3 It is a schematic diagram of a simple operation interface for fan SCADA data analysis.

[0030] Figure 4 It is a schematic diagram of a complex interaction interface for fan SCADA data analysis. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The present invention will be further described below in conjunction with embodiments, but it should not be understood that the above-mentioned subject matter scope of the present invention is limited to the following embodiments. Without departing from the above technical idea of the present invention, various substitutions and changes made according to common general technical knowledge and customary means in the art should be included within the protection scope of the present invention.

[0032] Embodiment 1:

[0033] See Figures 1 to 4 , a method for analyzing wind turbine SCADA data based on a large language model, comprising the following steps:

[0034] 1) Obtain historical wind turbine SCADA data;

[0035] 2) Preprocess the historical wind turbine SCADA data to obtain the wind turbine SCADA data to be analyzed;

[0036] 3) The large language model selects a machine learning model based on the wind turbine SCADA data to be analyzed;

[0037] 4) The large language model uses the wind turbine SCADA data to be analyzed to train the selected machine learning model to obtain a wind turbine SCADA data analysis model;

[0038] 5) Obtain real-time wind turbine SCADA data and input it into the wind turbine SCADA data analysis model to identify and / or predict wind turbine failures.

[0039] The wind turbine SCADA data includes wind speed, power, and temperature data.

[0040] The wind turbine SCADA data is obtained through sensor monitoring.

[0041] The steps for preprocessing the wind turbine SCADA data include: missing value processing, outlier detection and exclusion, and time series normalization.

[0042] The methods for missing value processing include: interpolation filling and deleting incomplete data.

[0043] Outliers are outliers and extreme values.

[0044] In step 3), if the target variable type of the wind turbine SCADA data to be analyzed is binary, the selected machine learning model is a classification model;

[0045] If the size of the wind turbine SCADA data to be analyzed is, or the characteristic is, the optimization algorithm adopted by the machine learning model is.

[0046] In step 3), the large language model selects a machine learning model through the LLMHelper module. If the target variable type of the wind turbine SCADA data to be analyzed is binary, a classification model is selected as the machine learning model; if the amount of wind turbine SCADA data to be analyzed is less than 10,000 records and the number of features is less than 100, gradient descent, Newton's method, or L-BFGS is selected as the optimization algorithm adopted by the machine learning model; if the amount of wind turbine SCADA data to be analyzed is greater than 10,000 records or the number of features is greater than 100, stochastic gradient descent (SGD), Adam, Adagrad, or RMSprop is selected as the optimization algorithm adopted by the machine learning model.

[0047] In step 4), the large language model trains the selected machine learning model through the train_and_visualize method and evaluates the performance of the machine learning model; if the accuracy of the machine learning model after training does not reach the threshold, or there are problems such as overfitting, underfitting, data characteristic mismatch, excessive computational resource consumption, etc. in the model, the hyperparameters are adjusted or a more suitable model is replaced according to the specific problems (for example, random forest is suitable for high-dimensional data and non-linear relationships, and gradient boosting tree is suitable for complex non-linear relationships and higher accuracy requirements), and training is carried out again.

[0048] After identifying and / or predicting a wind turbine fault, the wind turbine fault and real-time wind turbine SCADA data are also visualized.

[0049] Embodiment 2:

[0050] A method for analyzing wind turbine SCADA data based on a large language model includes the following steps:

[0051] 1) Obtain historical wind turbine SCADA data;

[0052] 2) Preprocess the historical wind turbine SCADA data to obtain the wind turbine SCADA data to be analyzed;

[0053] 3) The large language model selects a machine learning model based on the wind turbine SCADA data to be analyzed;

[0054] 4) The large language model uses the wind turbine SCADA data to be analyzed to train the selected machine learning model to obtain a wind turbine SCADA data analysis model;

[0055] 5) Obtain real-time wind turbine SCADA data and input it into the wind turbine SCADA data analysis model to identify and / or predict wind turbine faults.

[0056] Embodiment 3:

[0057] A method for analyzing wind turbine SCADA data based on a large language model, the technical content is the same as that of Example 2. Further, the wind turbine SCADA data includes wind speed, power, and temperature data.

[0058] Example 4:

[0059] A method for analyzing wind turbine SCADA data based on a large language model, the technical content is the same as any one of Examples 2-3. Further, the wind turbine SCADA data is obtained by sensor monitoring.

[0060] Example 5:

[0061] A method for analyzing wind turbine SCADA data based on a large language model, the technical content is the same as any one of Examples 2-4. Further, the steps for preprocessing the wind turbine SCADA data include: missing value processing, outlier detection and exclusion, and time series standardization.

[0062] Example 6:

[0063] A method for analyzing wind turbine SCADA data based on a large language model, the technical content is the same as any one of Examples 2-5. Further, the methods for missing value processing include: interpolation filling and deleting incomplete data.

[0064] Example 7:

[0065] A method for analyzing wind turbine SCADA data based on a large language model, the technical content is the same as any one of Examples 2-6. Further, the outliers are outliers and extreme values.

[0066] Example 8:

[0067] A method for analyzing wind turbine SCADA data based on a large language model, the technical content is the same as any one of Examples 2-7. Further, in step 3), if the target variable type of the wind turbine SCADA data to be analyzed is binary, then a classification model is selected as the machine learning model; if the amount of wind turbine SCADA data to be analyzed is less than 10,000 records and the number of features is less than 100, then gradient descent, Newton's method, or L-BFGS is selected as the optimization algorithm adopted by the machine learning model; if the amount of wind turbine SCADA data to be analyzed is greater than 10,000 records or the number of features is greater than 100, then stochastic gradient descent (SGD), Adam, Adagrad, or RMSprop is selected as the optimization algorithm adopted by the machine learning model.

[0068] Example 9:

[0069] A method for analyzing wind turbine SCADA data based on a large language model, the technical content is the same as any one of Examples 2-8. Further, in step 3), the large language model selects the machine learning model through the LLMHelper module.

[0070] Example 10:

[0071] A method for analyzing wind turbine SCADA data based on a large language model, the technical content is the same as any one of Examples 2-9. Further, in step 4), the large language model trains the selected machine learning model through the train_and_visualize method and evaluates the performance of the machine learning model; if the accuracy of the machine learning model after training does not reach the threshold, or there are problems such as overfitting, underfitting, data characteristic mismatch, and excessive computational resource consumption in the model, the hyperparameters are adjusted according to the specific problems or a more suitable model is replaced (for example, random forest is suitable for high-dimensional data and non-linear relationships, and gradient boosting tree is suitable for complex non-linear relationships and higher accuracy requirements), and training is carried out again.

[0072] Example 11:

[0073] A method for analyzing wind turbine SCADA data based on a large language model, the technical content is the same as any one of Examples 2-10. Further, after identifying and / or predicting a wind turbine fault, the wind turbine fault and real-time wind turbine SCADA data are also visualized.

[0074] Example 12:

[0075] A method for analyzing wind turbine SCADA data based on a large language model is as follows:

[0076] The specific results of the present invention can be referred to in the appendix Figure 1 , Figure 1 which is a framework diagram of a method for analyzing wind turbine SCADA data of the present invention. First is the data preprocessing module. This module first loads the SCADA data of the wind turbine, including key parameters such as wind speed, power, and temperature. Automatic technology is used to clean the data, including missing value filling and outlier processing, using interpolation and an outlier detection strategy based on the IQR method. Secondly is the model selection and optimization module: The system recommends suitable machine learning models (such as random forest, support vector machine, or neural network) through the LLM and automatically adjusts the model parameters according to the data characteristics to optimize the model training process. Finally is the free interactive analysis module: This module uses the trained model to analyze real-time data, identify and predict potential faults and anomalies. Through natural language generation technology, the complex data analysis results are converted into easy-to-understand text and graphical outputs. 2. Summary of the Invention

[0078] The design of the wind turbine SCADA data interpreter aims to achieve efficient processing and real-time interpretation of turbine status data through intelligent and automated means. Relying on large language models (LLMs) and related intelligent tools, the interpreter constructs a complete automated system from data preprocessing to model selection and optimization, and then to operating status interpretation and anomaly analysis. Figure 1 Shows the overall framework of the SCADA data interpreter.

[0079] 2.1 Data Preprocessing and Cleaning

[0080] The data preprocessing module designs a multi-step cleaning process, including data loading, missing value handling, anomaly detection, and normalization of time series. In the data loading phase, the interpreter dynamically selects the loading method according to the data volume. For large-scale data, parallel processing is adopted to improve the loading efficiency (such as using Dask), while smaller data sets are directly loaded using Pandas to ensure optimal performance and resource allocation. For missing value handling, the interpreter automatically adopts interpolation or deletes incomplete data according to the continuity requirements of the data set to ensure the integrity of the time series. In addition, the interpreter uses statistical methods to detect and mark anomalies in the SCADA data, identifying outliers and extreme values to exclude noise from the analysis set. Finally, the processed time series data is normalized to minimize the impact of different scales on the analysis results and lay the foundation for subsequent modeling and analysis in terms of data quality.

[0081] 2.2 Model Selection and Optimization

[0082] After data preprocessing and cleaning, the LLMHelper module of the LLM interacts with the user to automatically recommend a suitable model. In this process, the ModelTrainingAgent class first analyzes the type of target variable (binary, multi-class, or continuous) in the training data to determine whether to use a classification model or a regression model. For example, according to the data size and characteristics, it may recommend algorithms such as DBSCAN, Random-ForestClassifier, or XGBoost to optimize the prediction of the turbine operating status.

[0083] The choose_model method of this class dynamically selects the most suitable machine learning model according to the model suggestions provided by LLMHelper. Subsequently, the train_and_visualize method trains the selected model and evaluates its performance. If the accuracy of the model does not reach the predefined 85% threshold, the LLM will provide optimization suggestions, such as adjusting hyperparameters or replacing the model, and then retrain the model to improve the accuracy.

[0084] In addition, the ModelTrainingAgent is responsible for conducting in-depth analysis of the turbine's operating status, including state clustering using the DBSCAN algorithm and generating explanatory text for the operating status through the LLMHelper based on the clustering results. During this process, data visualization charts such as power curve charts are also generated and saved. Finally, the analysis results, status explanations, and charts are compiled into a report to provide detailed operation and maintenance suggestions for maintenance personnel.

[0085] 2.3 Interactive Analysis of SCADA Data

[0086] Supported by the data preprocessing, model selection and optimization module, the SCADA data analysis method provides an integrated user interface for data analysis tasks such as data status interpretation and anomaly detection. Designed with Tkinter, the interface presents a small window where users can instantly view the turbine's operating status, anomaly detection results, and data analysis charts, while providing text explanations and operation suggestions.

[0087] The interface provides natural language explanations generated by the LLM to make the analysis results more understandable. Users can clearly see the current key parameters such as wind speed and power and quickly judge whether the equipment is operating normally. For abnormal data or equipment failures, the interface provides cause analysis and suggestions to help maintenance personnel take prompt action. This interactive interface significantly improves the usability of SCADA data, enabling non-technical users to intuitively understand the data, thus promoting efficient monitoring and maintenance decision-making in wind farm operations.

Claims

1. A method for analyzing wind turbine SCADA data based on a large language model, characterized in that: The following steps are involved: 1) Obtain historical wind turbine SCADA data. 2) Preprocess the historical wind turbine SCADA data to obtain the wind turbine SCADA data to be analyzed; 3) The large language model selects the machine learning model based on the wind turbine SCADA data to be analyzed; 4) The large language model uses the wind turbine SCADA data to be analyzed to train the selected machine learning model to obtain the wind turbine SCADA data analysis model; 5) Obtain real-time wind turbine SCADA data and input it into the wind turbine SCADA data analysis model to identify and / or predict wind turbine failures.

2. A wind turbine SCADA data analysis method based on a large language model according to claim 1, characterized in that: The wind turbine SCADA data includes wind speed, power and temperature data.

3. A wind turbine SCADA data analysis method based on a large language model according to claim 2, characterized in that: Wind turbine SCADA data is obtained through sensor monitoring.

4. A wind turbine SCADA data analysis method based on a large language model according to claim 1, characterized in that: The steps of preprocessing wind turbine SCADA data include: missing value processing, outlier detection and exclusion, and time series standardization.

5. A wind turbine SCADA data analysis method based on a large language model according to claim 4, characterized in that: Methods for handling missing values ​​include: interpolation filling and deletion of incomplete data.

6. A wind turbine SCADA data analysis method based on a large language model according to claim 4, characterized in that: Outliers are outliers and extreme values.

7. A wind turbine SCADA data analysis method based on a large language model according to claim 1, characterized in that: In step 3), if the target variable type of the wind turbine SCADA data to be analyzed is binary, the classification model is selected as the machine learning model; if the amount of wind turbine SCADA data to be analyzed is less than 10,000 records and the number of features is less than 100, gradient descent, Newton's method or L-BFGS is selected as the optimization algorithm used by the machine learning model; if the amount of wind turbine SCADA data to be analyzed is greater than 10,000 records or the number of features is greater than 100, stochastic gradient descent, Adam, Adagrad or RMSprop is selected as the optimization algorithm used by the machine learning model.

8. A wind turbine SCADA data analysis method based on a large language model according to claim 1, characterized in that: In step 3), the large language model selects the machine learning model through the LLMHelper module.

9. A wind turbine SCADA data analysis method based on a large language model according to claim 1, characterized in that: In step 4), the large language model trains the selected machine learning model through the train_and_visualize method and evaluates the performance of the machine learning model; if the accuracy of the machine learning model after training does not reach the threshold, or the model is overfitting, underfitting, data feature mismatch, or computing resource consumption is too large, the hyperparameters are adjusted or the machine learning model is replaced and retrained.

10. A wind turbine SCADA data analysis method based on a large language model according to claim 1, characterized in that: After identifying and / or predicting wind turbine failures, real-time wind turbine SCADA data is also visualized.