Fault diagnosis and optimization method based on rough set and random forest and related equipment
By fusing multi-source heterogeneous data using rough set and random forest methods, a high-precision diagnostic model is constructed and a real-time optimization strategy is generated, which solves the problem of intelligent diagnosis and operation and maintenance of new energy power systems and improves the system's intelligence level and operational efficiency.
Patent Information
- Application Number
- CN202610104653.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-26
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies struggle to effectively integrate data from multiple systems and parameters in new energy power systems, resulting in delayed fault warnings, high false alarm rates, a lack of intelligent operation and maintenance strategies, frequent unplanned outages, and high operation and maintenance costs.
The rough set algorithm is used to reduce the features of multi-source heterogeneous data and integrate them across systems. A random forest diagnostic model is constructed, and multi-objective optimization techniques are combined to generate operation and maintenance strategies, thus realizing a closed loop of intelligent diagnosis and dynamic optimization.
It has improved the intelligence level and operational efficiency of the operation and maintenance of new energy power systems, and reduced operation and maintenance costs and unplanned downtime frequency through high-precision fault diagnosis and real-time optimization strategies.
Smart Images

Figure CN122065173A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automatic optimization technology for new energy sources, specifically to a fault diagnosis and optimization method and related equipment based on rough sets and random forests. Background Technology
[0002] With the large-scale development of new energy power systems, the operating data of equipment such as wind turbines and photovoltaic power plants are characterized by massive volume, heterogeneity, high dimensionality, and strong real-time performance.
[0003] Traditional equipment monitoring and fault diagnosis methods typically rely on a single data source and threshold alarms, making it difficult to effectively extract early fault characteristics from complex data involving multiple systems and parameters. This results in problems such as delayed warnings, high false alarm rates, and limited diagnostic dimensions. Furthermore, existing operation and maintenance strategies are mostly based on fixed rules or manual experience, lacking intelligent optimization capabilities that coordinate with real-time equipment status, operating environment, and power grid dispatch requirements. This leads to limited operational efficiency, frequent unplanned downtime, and high operation and maintenance costs.
[0004] Currently, although some studies have applied machine learning algorithms such as random forests to equipment fault classification or rough set theory to feature reduction, most are limited to single algorithms or isolated subsystems. A closed-loop technical system capable of integrating multi-source heterogeneous data governance, intelligent diagnosis, and dynamic operation optimization has not yet been formed. Especially when dealing with the collaborative operation and maintenance of multiple types of equipment such as wind turbines and photovoltaic systems in new energy power plants, how to achieve cross-system data fusion, construct interpretable and iterative diagnostic models, and generate safe and executable optimization strategies in real time based on diagnostic results remains a critical technical bottleneck that the industry urgently needs to overcome. Summary of the Invention
[0005] The purpose of this invention is to provide a fault diagnosis and optimization method and related equipment based on rough set and random forest to overcome the problems existing in the prior art. This invention realizes feature reduction and cross-system fusion of multi-source heterogeneous data through rough set algorithm, constructs a high-precision and interpretable random forest diagnostic model, and generates safe, economical and executable operation and maintenance strategies and operation instructions in real time based on the early warning information output by the model using multi-objective optimization technology. It effectively breaks through the closed-loop technical bottleneck from data fusion, intelligent diagnosis to dynamic optimization, and significantly improves the intelligent operation and maintenance level and overall operation efficiency of new energy power systems.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a fault diagnosis and optimization method based on rough sets and random forests, comprising the following steps: Step 1: Collect multi-source heterogeneous operation data through the new energy power system; Step 2: Input the multi-source heterogeneous operating data into the constructed fault diagnosis model and output early warning information, which includes fault categories; wherein, the training method of the fault diagnosis model includes using a rough set algorithm to perform knowledge reduction on the multidimensional dataset, and inputting the knowledge-reduced equipment operating sample set into a random forest model; Step 3: Based on the early warning information and real-time operational data, conduct multi-objective optimization analysis through a big data platform to generate decision support scheme data.
[0007] In some embodiments, the steps for constructing the fault diagnosis model include: Historical multi-source heterogeneous operation data are collected from new energy power systems; Historical multi-source heterogeneous operation data is cleaned, denoised, and normalized. A heterogeneous storage strategy is adopted to uniformly store and manage the normalized historical multi-source heterogeneous operation data, resulting in a multidimensional dataset covering equipment health, operating efficiency, and fault characteristics. Rough set algorithm is used to perform knowledge reduction on multidimensional dataset to extract key fault feature attributes of multidimensional dataset and obtain a knowledge-reduced set of equipment operation samples. The reduced set of equipment operation samples is input into the random forest model for training, generating a fault diagnosis decision forest composed of multiple decision trees. The fault category with the most votes in the fault diagnosis decision forest is selected as the final comprehensive diagnosis result of the random forest through a voting mechanism, thus completing the construction of the fault diagnosis model.
[0008] In some embodiments, the multi-source heterogeneous operating data includes wind turbine tower vibration data, main shaft temperature data, gearbox grease status data, yaw system parameter data, pitch system status data, blade icing image data, wind measurement system deviation data, inverter electrical parameter data, and photovoltaic module infrared images. The historical multi-source heterogeneous operating data includes historical wind turbine tower vibration data, historical main shaft temperature data, historical gearbox grease status data, historical yaw system parameter data, historical pitch system status data, historical blade icing image data, historical wind measurement system deviation data, historical inverter electrical parameter data, and historical photovoltaic module infrared images.
[0009] In some embodiments, the rough set algorithm includes defining indistinguishable relations, calculating upper and lower approximation precision, and roughness.
[0010] In some embodiments, the step of inputting the knowledge-reduced device operation sample set into a random forest model for training, generating a fault diagnosis decision forest composed of multiple decision trees, and selecting the fault category with the most votes in the fault diagnosis decision forest as the final comprehensive diagnostic result output by the random forest through a voting mechanism, specifically includes: The Bootstrap method is used to perform random sampling with replacement on the knowledge-reduced set of equipment operation samples to generate several independent and overlapping training subsets. For each training subset, an independent decision tree is constructed using the decision tree algorithm. Each decision tree outputs an independent fault category judgment, resulting in several fault category judgments. The fault category with the most votes is selected as the final comprehensive diagnostic result of the random forest through a voting mechanism.
[0011] In some embodiments, the decision support scheme data includes wind turbine start-up and shutdown strategies, yaw adjustment suggestions, icing treatment schemes, maintenance plans, and operating parameter optimization strategies.
[0012] In some embodiments, the step of generating decision support scheme data by performing multi-objective optimization analysis through a big data platform based on early warning information and real-time operational data specifically includes: The big data platform receives and integrates early warning information and real-time operational data to obtain a comprehensive decision input matrix. The system calls a pre-defined multi-objective optimization model to analyze and calculate the comprehensive decision input matrix, obtaining one or a set of Pareto optimal solutions, and then converts the optimal solution set into actionable decision support scheme data.
[0013] Secondly, this invention provides a fault diagnosis and optimization system based on rough sets and random forests, comprising: The data acquisition module is used to collect multi-source heterogeneous operation data from the new energy power system; The early warning information output module is used to input multi-source heterogeneous operating data into the constructed fault diagnosis model and output early warning information, which includes fault categories; wherein, the training method of the fault diagnosis model includes using a rough set algorithm to perform knowledge reduction on the multidimensional dataset, and inputting the knowledge-reduced equipment operating sample set into a random forest model; The optimization module is used to perform multi-objective optimization analysis based on early warning information and real-time operational data through a big data platform, and generate decision support solution data.
[0014] Thirdly, the present invention provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.
[0015] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.
[0016] The above technical solution has the following advantages or beneficial effects: Firstly, this invention provides a fault diagnosis and optimization method based on rough set and random forest. By using the rough set algorithm to achieve feature reduction and cross-system fusion of multi-source heterogeneous data, a high-precision and interpretable random forest diagnostic model is constructed. Based on the early warning information output by the model, a safe, economical and executable operation and maintenance strategy and operation instructions are generated in real time using multi-objective optimization technology. This effectively breaks through the closed-loop technical bottleneck from data fusion, intelligent diagnosis to dynamic optimization, and significantly improves the intelligent operation and maintenance level and overall operation efficiency of new energy power systems.
[0017] In some embodiments, a heterogeneous storage strategy is used to achieve unified governance of multi-source historical data, forming a high-quality multidimensional dataset; rough set algorithm is used for knowledge reduction, effectively extracting key fault features and improving model interpretability; finally, a highly robust diagnostic decision forest is constructed through random forest ensemble learning, realizing accurate and stable identification and classification of complex faults, providing a reliable and efficient intelligent diagnostic foundation for subsequent real-time optimization, and significantly improving the accuracy and anticipation of fault warnings.
[0018] In some embodiments, by explicitly covering a full-dimensional data system encompassing core equipment such as wind turbines and photovoltaics, from mechanical vibration and thermal state to electrical performance and image features, a panoramic perception and fusion of multi-equipment and multi-state information of new energy power plants is achieved. This comprehensive data foundation not only provides the rough set-random forest model with complete training samples covering multiple physical fields such as mechanical, electrical, and environmental aspects, significantly improving the model's ability to identify complex faults and latent defects, but also ensures that subsequent optimization strategies can be based on the real and comprehensive operating status of the equipment, thereby generating more accurate and on-site safety control and operation and maintenance decisions.
[0019] In some embodiments, by introducing indistinguishable relationships, upper and lower approximation precision, and roughness calculations from rough set algorithms, information redundancy and uncertainty can be quantitatively evaluated and eliminated from massive, high-dimensional heterogeneous data, thereby accurately extracting the key feature subset most relevant to the fault, effectively improving the training efficiency and generalization ability of the subsequent random forest model, and enhancing the interpretability of the fault diagnosis process.
[0020] In some embodiments, diverse training subsets are constructed through Bootstrap sampling, and multiple decision trees with differences are generated based on random feature selection, which effectively improves the model's generalization ability and anti-overfitting performance. Finally, a majority voting mechanism is used to integrate the judgments of all trees, ensuring the stability and high reliability of the diagnostic results, enabling it to output robust and accurate comprehensive fault diagnosis conclusions when facing complex and ever-changing new energy equipment operation scenarios.
[0021] In some embodiments, by concretizing the abstract optimization results into a series of directly executable operation and maintenance instructions such as wind turbine start-up and shutdown, yaw adjustment, icing treatment, maintenance scheduling and parameter optimization, the precise implementation from intelligent diagnosis to on-site operation is realized. This ensures that the optimization strategy is not only at the theoretical level, but is transformed into a schedulable, traceable and evaluable closed-loop action plan, which greatly improves the timeliness and accuracy of operation and maintenance response and the overall system's executable intelligence level.
[0022] In some embodiments, by integrating early warning information with real-time operating conditions into a unified decision input, and using a multi-objective optimization model to solve for the Pareto optimal solution set, the coordinated optimization of objectives such as safety, economy, and efficiency under multiple constraints is achieved; finally, the mathematically optimal solution is transformed into directly executable operation and maintenance instructions, ensuring that the generated decision scheme is not only theoretically optimal, but also feasible on site.
[0023] Secondly, this invention provides a fault diagnosis and optimization system based on rough set and random forest. Through modular design, it seamlessly integrates three major functions: data acquisition, intelligent diagnosis, and dynamic optimization. The early warning information output module is based on an advanced rough set-random forest fusion model to achieve real-time and accurate fault diagnosis of multi-source heterogeneous data. The optimization module generates operation and maintenance decision schemes that take into account safety, economy, and efficiency based on the diagnosis results and using multi-objective optimization technology. The entire system is supported by a big data platform, forming a complete automated closed loop from "state perception" to "optimization execution", which significantly improves the intelligent operation and maintenance level and overall operating efficiency of new energy power plants.
[0024] Thirdly, the present invention provides a computer device that, through a processor executing a specific computer program, can efficiently implement the steps of the method of the present invention. When performing data processing tasks, the computer device can accurately perform numerical calculations and logical judgments, avoiding errors caused by human factors. At the same time, since the computer program has high stability and reliability, it can ensure the accuracy and consistency of the data processing results.
[0025] Fourthly, the present invention provides a computer-readable storage medium. By programming the steps of the method of the present invention into a computer program and storing it on the computer-readable storage medium, users can easily load these programs onto any compatible computer device and execute them without rewriting or converting the code, which greatly improves the convenience and flexibility of program execution. Attached Figure Description
[0026] Figure 1 This is a flowchart illustrating a fault diagnosis and optimization method based on rough set and random forest, as shown in some embodiments of this specification. Figure 2This is a schematic diagram of a single decision tree according to some embodiments of this specification; Figure 3 This is a schematic diagram of the structure of a computer device according to some embodiments of this specification. Detailed Implementation
[0027] The present invention will be further described in detail below with reference to specific embodiments. These descriptions are for explanation purposes only and are not intended to limit the scope of the invention. To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention. It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0028] The purpose of this invention is to provide a fault diagnosis and optimization method and related equipment based on rough set and random forest to overcome the problems existing in the prior art. This invention realizes feature reduction and cross-system fusion of multi-source heterogeneous data through rough set algorithm, constructs a high-precision and interpretable random forest diagnostic model, and generates safe, economical and executable operation and maintenance strategies and operation instructions in real time based on the early warning information output by the model using multi-objective optimization technology. It effectively breaks through the closed-loop technical bottleneck from data fusion, intelligent diagnosis to dynamic optimization, and significantly improves the intelligent operation and maintenance level and overall operation efficiency of new energy power systems.
[0029] Example: This embodiment provides a fault diagnosis and optimization method based on rough sets and random forests. See [link to relevant documentation]. Figure 1 This includes the following steps: Step 1: Collect multi-source heterogeneous operation data through the new energy power system.
[0030] In some embodiments, the multi-source heterogeneous operating data includes wind turbine tower vibration data, main shaft temperature data, gearbox grease status data, yaw system parameter data, pitch system status data, blade icing image data, wind measurement system deviation data, inverter electrical parameter data, and photovoltaic module infrared images.
[0031] In some embodiments, the new energy industry consists of industrial enterprises operating 24 / 7, resulting in an extremely large amount of data generated daily, with storage units typically measured in TB or even PB. For example, a conventional SCADA system, based on a sampling interval of 3-4 seconds per measurement point, would generate 12B / frame × 0.3 frames / second × 86400 seconds / day × 365 days × 10000 points = 1.03TB annually.
[0032] The main sources of new energy power data are extracted from various business systems in the current power production process. Existing systems such as Huaneng Group's "Smart Operation and Maintenance System," "Huaneng Jiangxi Clean Energy Daily Operation and Accounting System," "New Energy Lean Management Platform," Xinrui Zone 1 Monitoring System, and Smart Access Control System can serve as major components of the data sources. In the future, meteorological data, management system data, and other extended business data will be introduced. These data types not only include a large amount of structured data but also some unstructured data such as documents and images.
[0033] Step 2: Input the multi-source heterogeneous operating data into the constructed fault diagnosis model and output early warning information, which includes fault categories; wherein, the training method of the fault diagnosis model includes using a rough set algorithm to perform knowledge reduction on the multidimensional dataset, and inputting the knowledge-reduced equipment operating sample set into a random forest model.
[0034] In some embodiments, the steps for constructing the fault diagnosis model include: Step 2.1: Collect historical multi-source heterogeneous operation data through the new energy power system.
[0035] In some embodiments, the historical multi-source heterogeneous operating data includes historical wind turbine tower vibration data, historical main shaft temperature data, historical gearbox grease status data, historical yaw system parameter data, historical pitch system status data, historical blade icing image data, historical wind measurement system deviation data, historical inverter electrical parameter data, and historical photovoltaic module infrared images.
[0036] Step 2.2 involves cleaning, denoising, and normalizing the historical multi-source heterogeneous operating data. A heterogeneous storage strategy is then used to uniformly store and manage the normalized historical multi-source heterogeneous operating data, resulting in a multidimensional dataset covering equipment health, operating efficiency, and fault characteristics.
[0037] In some embodiments, historical multi-source heterogeneous operational data that has undergone cleaning and normalization are stored in domains using differentiated storage engines based on their data type, access frequency, and business needs; time-series data (such as vibration and temperature curves) are stored in time-series databases for efficient compression and fast retrieval; structured operational logs are stored in relational databases to ensure transaction consistency; and unstructured images and documents (such as infrared thermal images and maintenance reports) are stored in distributed file systems. Secondly, by constructing a unified data connector and metadata management layer, a virtual "logical data lake" is established on top of the above-mentioned storage systems. This logical layer performs unified cataloging and indexing of all data and provides standardized access interfaces to the outside world.
[0038] In some embodiments, a heterogeneous data storage strategy is adopted to establish underlying data communication channels. Different storage engines are used for different data types. By constructing distributed file systems, distributed data warehouses, non-relational databases (such as columnar databases, document databases, graph databases, etc.), and relational databases, centralized storage and unified management of various types of data are achieved, meeting the low-cost, high-performance storage needs of diverse and large-scale data. Simultaneously, connectors are built between the various storage systems to enable rapid data fusion and form a data mart.
[0039] Step 2.3: Use the rough set algorithm to perform knowledge reduction on the multidimensional dataset to extract key fault feature attributes from the multidimensional dataset and obtain the equipment operation sample set after knowledge reduction.
[0040] In some embodiments, the rough set algorithm includes defining indistinguishable relations, calculating upper and lower approximation precision, and roughness.
[0041] In some embodiments, a rough set algorithm is used to perform knowledge reduction on a multidimensional dataset, specifically including: Define the domain of device operating parameters U, the set of conditional attributes C, and the set of decision attributes D; construct the indistinguishable relation IND(P), calculate the attribute reduction set G based on IND(P) such that IND(G) = IND(P) and G is independent, and obtain the reduced attribute set; reconstruct the sample set based on the reduced attribute set for subsequent random forest model training.
[0042] In some embodiments, the formula for the rough set algorithm is: ; In the formula, Indicates roughness; R Representing knowledge; C This indicates the number of elements in the set, satisfying the interval [0, 1]. Representing approximate knowledge; To represent approximate knowledge.
[0043] Step 2.4: Input the knowledge-reduced set of equipment operation samples into the random forest model for training, generate a fault diagnosis decision forest composed of multiple decision trees, and select the fault category with the most votes in the fault diagnosis decision forest as the final comprehensive diagnosis result output by the random forest through a voting mechanism, thus completing the construction of the fault diagnosis model.
[0044] In some embodiments, the step of inputting the knowledge-reduced device operation sample set into a random forest model for training, generating a fault diagnosis decision forest composed of multiple decision trees, and selecting the fault category with the most votes in the fault diagnosis decision forest as the final comprehensive diagnostic result output by the random forest through a voting mechanism, specifically includes: Step 2.4.1 uses the Bootstrap method to perform random sampling with replacement on the knowledge-reduced equipment operation sample set to generate several independent and overlapping training subsets. Each subset retains the main features of the original data while introducing a certain degree of randomness and variability. The output of this step lays the foundation for the subsequent construction of diverse base classifiers.
[0045] Step 2.4.2: For each training subset, an independent decision tree is constructed using the decision tree algorithm. Each decision tree outputs an independent fault category judgment, resulting in several fault category judgments. When splitting each non-leaf node of each tree, the algorithm does not select the optimal split from all features, but rather randomly selects a subset of features from all features, and then selects the best split point within that subset using indicators such as information gain or Gini coefficient. This process ensures that the growth path and structure of each tree are random, effectively enhancing the model's generalization ability and suppressing overfitting.
[0046] In some embodiments, see Figure 2Before training, the random forest algorithm needs to perform multiple sampling on a sample set composed of operating parameters collected from power equipment to generate multiple sets for training, namely the training set. The basic steps for constructing a random forest are as follows: Initialization process, sample collection, random sampling using the Bootstrap method to obtain a set of W training samples X = {x1, x2, ..., xW}, where x represents a single training sample; using the training sample set X, a decision tree F = {f1, f2, ..., fW} is generated using a decision tree generation algorithm, where f represents a single decision tree; when selecting parameter attributes at each non-leaf node, several attributes are randomly selected from all parameter attributes as the splitting attributes of the current node, and child nodes are generated according to the information gain value; pruning is performed on each decision tree to remove bad nodes, which is used to simplify the entire forest, prevent overfitting, and improve algorithm efficiency; by testing each decision tree with samples, the corresponding fault categories C1(x), C2(x), ..., CW(x) are obtained; the category with the most output fault categories among all decision trees in the decision forest is used as the output result of the entire random forest using a voting method.
[0047] Step 2.4.3: The fault category with the most votes is output as the final comprehensive diagnosis result of the random forest through a voting mechanism. The ensemble strategy cleverly summarizes the judgment results of multiple weak classifiers (single trees) that may not be completely accurate but have differences, thereby obtaining a more robust and reliable strong classification conclusion.
[0048] Step 3: Based on the early warning information and real-time operational data, conduct multi-objective optimization analysis through a big data platform to generate decision support scheme data.
[0049] In some embodiments, the decision support scheme data includes wind turbine start-up and shutdown strategies, yaw adjustment suggestions, icing treatment schemes, maintenance plans, and operating parameter optimization strategies.
[0050] Step 3, which involves generating decision support scheme data by performing multi-objective optimization analysis based on early warning information and real-time operational data through a big data platform, specifically includes: Step 3.1: The big data platform receives and integrates early warning information and real-time operating data to obtain a comprehensive decision input matrix. The platform aligns and correlates the structured diagnostic results (such as fault type, probability, and location) with real-time operating data (such as wind speed, irradiance, and power grid dispatch instructions) from SCADA and environmental monitoring systems in time and space to form a comprehensive decision input matrix that integrates equipment health status and system operating environment.
[0051] Step 3.2: Call the preset multi-objective optimization model to analyze and calculate the comprehensive decision input matrix to obtain one or a set of Pareto optimal solutions, and convert the optimal solution set into actionable decision support scheme data.
[0052] In some embodiments, the big data platform fully utilizes the most advanced big data-related thinking, methods, and tools, and based on the source and characteristics of new energy power data, adopts targeted strategies and patterns to meet the professional needs of new energy information flow in processing and functional applications, thereby forming a full-process technical framework that integrates collection, storage, computing, analysis, and application.
[0053] In some embodiments, the service platform serves as a crucial display and interface platform for the big data system platform. Through the service platform, direct interaction with users is possible, generating visual interfaces for analysis results and creating customized reports. It also generates optimal auxiliary decision-making solutions for power enterprise information systems, providing reference for operators and maintenance personnel. The big data service platform is developed based on the functional interfaces of existing intelligent systems, directly connecting various systems and capturing required data for secondary analysis and processing.
[0054] In one embodiment of the present invention, a fault diagnosis and optimization system based on rough set and random forest is provided, comprising: The data acquisition module is used to collect multi-source heterogeneous operation data from the new energy power system; The early warning information output module is used to input multi-source heterogeneous operating data into the constructed fault diagnosis model and output early warning information, which includes fault categories; wherein, the training method of the fault diagnosis model includes using a rough set algorithm to perform knowledge reduction on the multidimensional dataset, and inputting the knowledge-reduced equipment operating sample set into a random forest model; The optimization module is used to perform multi-objective optimization analysis based on early warning information and real-time operational data through a big data platform, and generate decision support solution data.
[0055] See Figure 3In one embodiment of the present invention, a computer device is provided, comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, and is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the computer storage medium to realize a corresponding method flow or corresponding function. The processor described in this embodiment of the present invention can be used for the operation of fault diagnosis and optimization methods based on rough sets and random forests.
[0056] In one embodiment of the present invention, a computer-readable storage medium is provided, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and extended storage media supported by the computer device. The computer-readable storage medium provides storage space that stores the terminal's operating system; and the storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk storage device. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the fault diagnosis and optimization method based on rough sets and random forests in the embodiment.
[0057] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0058] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0059] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0060] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0061] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A fault diagnosis and optimization method based on rough sets and random forests, characterized in that, Includes the following steps: Collect multi-source heterogeneous operation data through new energy power systems; Multi-source heterogeneous operating data is input into the constructed fault diagnosis model, and early warning information is output, including fault categories; wherein, the training method of the fault diagnosis model includes using a rough set algorithm to perform knowledge reduction on the multidimensional dataset, and inputting the knowledge-reduced equipment operating sample set into a random forest model; Based on early warning information and real-time operational data, multi-objective optimization analysis is conducted through a big data platform to generate decision support solution data.
2. The fault diagnosis and optimization method based on rough set and random forest according to claim 1, characterized in that, The steps for constructing the fault diagnosis model include: Historical multi-source heterogeneous operation data are collected from new energy power systems; Historical multi-source heterogeneous operation data is cleaned, denoised, and normalized. A heterogeneous storage strategy is adopted to uniformly store and manage the normalized historical multi-source heterogeneous operation data, resulting in a multidimensional dataset covering equipment health, operating efficiency, and fault characteristics. Rough set algorithm is used to perform knowledge reduction on multidimensional dataset to extract key fault feature attributes of multidimensional dataset and obtain a knowledge-reduced set of equipment operation samples. The reduced set of equipment operation samples is input into the random forest model for training, generating a fault diagnosis decision forest composed of multiple decision trees. The fault category with the most votes in the fault diagnosis decision forest is selected as the final comprehensive diagnosis result of the random forest through a voting mechanism, thus completing the construction of the fault diagnosis model.
3. The fault diagnosis and optimization method based on rough set and random forest according to claim 2, characterized in that, The multi-source heterogeneous operating data includes wind turbine tower vibration data, main shaft temperature data, gearbox grease status data, yaw system parameter data, pitch system status data, blade icing image data, wind measurement system deviation data, inverter electrical parameter data, and photovoltaic module infrared images. The historical multi-source heterogeneous operating data includes historical wind turbine tower vibration data, historical main shaft temperature data, historical gearbox grease status data, historical yaw system parameter data, historical pitch system status data, historical blade icing image data, historical wind measurement system deviation data, historical inverter electrical parameter data, and historical photovoltaic module infrared images.
4. The fault diagnosis and optimization method based on rough set and random forest according to claim 2, characterized in that, The rough set algorithm includes defining indistinguishable relations, calculating upper and lower approximation accuracy, and roughness.
5. The fault diagnosis and optimization method based on rough set and random forest according to claim 2, characterized in that, The process involves inputting the knowledge-reduced set of equipment operation samples into a random forest model for training, generating a fault diagnosis decision forest composed of multiple decision trees. A voting mechanism is then used to select the fault category with the most votes in the fault diagnosis decision forest as the final comprehensive diagnostic result output by the random forest. Specifically, this includes: The Bootstrap method is used to perform random sampling with replacement on the knowledge-reduced set of equipment operation samples to generate several independent and overlapping training subsets. For each training subset, an independent decision tree is constructed using the decision tree algorithm. Each decision tree outputs an independent fault category judgment, resulting in several fault category judgments. The fault category with the most votes is selected as the final comprehensive diagnostic result of the random forest through a voting mechanism.
6. The fault diagnosis and optimization method based on rough set and random forest according to claim 1, characterized in that, The decision support solution data includes wind turbine start-up and shutdown strategies, yaw adjustment suggestions, icing treatment plans, maintenance plans, and operating parameter optimization strategies.
7. The fault diagnosis and optimization method based on rough set and random forest according to claim 1, characterized in that, The process of generating decision support scheme data based on early warning information and real-time operational data through a big data platform for multi-objective optimization analysis includes: The big data platform receives and integrates early warning information and real-time operational data to obtain a comprehensive decision input matrix. The system calls a pre-defined multi-objective optimization model to analyze and calculate the comprehensive decision input matrix, obtaining one or a set of Pareto optimal solutions, and then converts the optimal solution set into actionable decision support scheme data.
8. A fault diagnosis and optimization system based on rough sets and random forests, characterized in that, include: The data acquisition module is used to collect multi-source heterogeneous operation data from the new energy power system; The early warning information output module is used to input multi-source heterogeneous operating data into the constructed fault diagnosis model and output early warning information, which includes fault categories; wherein, the training method of the fault diagnosis model includes using a rough set algorithm to perform knowledge reduction on the multidimensional dataset, and inputting the knowledge-reduced equipment operating sample set into a random forest model; The optimization module is used to perform multi-objective optimization analysis based on early warning information and real-time operational data through a big data platform, and generate decision support solution data.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1-7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-7.