A method, apparatus, electronic device, and storage medium for predicting equipment health index.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本申请提供一种设备健康指数预测方法、装置、电子设备及存储介质,用于解决现有技术中风力发电设备的监测数据偏差导致风力发电设备的健康指数的预测准确性低的技术问题
本技术方案通过对整个目标取值区间的遍历,生成大量候选预测方案,而非仅给出单一点预测。当根据某一取值估计的监测数据与预设条件存在较大偏差时,择优机制会优先将其排除,仍有其他候选取值可供选择;当根据某一取值估计的监测数据与预设条件高度一致时,择优机制将其保留。原始采集的监测数据在本技术方案中用于为预设条件中历史分布、历史统计特征等参照标准的构建提供来源,即使原始监测数据存在偏差或波动,这些偏差或波动既不会直接影响择优判断的参照基准,也不会直接进入预测结果,从而降低了由于输入的监测数据的偏差导致风力发电设备健康指数的预测失准风险。
Smart Images

Figure CN122359250B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of condition monitoring of wind power generation equipment, and in particular to a method, device, electronic equipment and storage medium for predicting equipment health index. Background Technology
[0002] The Health Index (HI) of wind power equipment is often used to describe and / or measure the health status of core transmission components such as gearboxes and main bearings. Predicting the health index can help maintenance personnel understand the degradation trend of the equipment and thus assist in making corresponding maintenance decisions.
[0003] Existing health index prediction methods are based on monitoring data such as vibration data, data collected by Supervisory Control and Data Acquisition (SCADA) systems, and oil wear particle analysis data to build prediction models. However, prediction models assume that the input features are accurate and reliable during training and application. The monitoring data deviates from the actual state due to noise caused by large differences in sampling frequency, the accuracy of acquisition equipment, and data transmission errors, which further reduces the accuracy of the prediction results. Summary of the Invention
[0004] This application provides a method, apparatus, electronic device, and storage medium for predicting equipment health index, which addresses the technical problem of low prediction accuracy of the health index of wind power generation equipment due to deviations in monitoring data.
[0005] Firstly, this application provides a method for predicting equipment health index, including: Obtain the target health index range of the target wind power generation equipment within the target time period; Based on the target value range and the monitoring data estimation model, multiple monitoring data related to the target health index at multiple values are obtained. Each monitoring data corresponds to data of at least one data type in the preset data types, including mechanical vibration type, operating condition type and lubrication wear type. The monitoring data estimation model is obtained by training the initial estimation model based on the sample health index of the sample wind power generation equipment and the sample monitoring data corresponding to each sample health index. Target monitoring data that meets preset conditions are obtained from the plurality of monitoring data. The preset conditions are determined based on evaluation parameters, which include at least one of probability P-value and divergence. The target value corresponding to the target monitoring data is used as the predicted value of the target health index for the target time period.
[0006] Secondly, this application provides a device for predicting equipment health index, comprising: The acquisition module is used to obtain the target health index range of the target wind power generation equipment within the target time period; The estimation module is used to obtain multiple monitoring data related to the target health index at multiple values based on the target value range and the monitoring data estimation model. Each monitoring data corresponds to data of at least one data type in the preset data types. The preset data types include mechanical vibration type, operating condition type and lubrication wear type. The monitoring data estimation model is obtained by training an initial estimation model based on the sample health index of the sample wind power generation equipment and the sample monitoring data corresponding to each sample health index. The selection module is used to obtain target monitoring data that meets preset conditions from the plurality of monitoring data. The preset conditions are determined based on evaluation parameters, which include at least one of probability P-value and divergence. The output module is used to take the target value corresponding to the target monitoring data as the predicted value of the target health index in the target time period.
[0007] Thirdly, this application provides a computer device including a memory and a processor, the memory having a computer program executable on the processor, the processor executing the program to implement the steps of the aforementioned data prediction method.
[0008] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the aforementioned data prediction method.
[0009] Compared with the prior art, the technical solution of the present invention has the following beneficial effects: This technical solution generates a large number of candidate prediction schemes by traversing the entire target value range, rather than providing a single-point prediction. When the monitoring data estimated based on a certain value deviates significantly from the preset conditions, the selection mechanism will prioritize excluding it, while other candidate values remain available. When the monitoring data estimated based on a certain value is highly consistent with the preset conditions, the selection mechanism will retain it. The original collected monitoring data in this technical solution is used to provide a source for constructing reference standards such as historical distribution and historical statistical characteristics in the preset conditions. Even if the original monitoring data has deviations or fluctuations, these deviations or fluctuations will not directly affect the reference benchmark for selection judgment, nor will they directly enter the prediction results, thereby reducing the risk of inaccurate prediction of the wind power equipment health index due to deviations in the input monitoring data. Attached Figure Description
[0010] The accompanying drawings used are briefly described below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0011] Figure 1 A flowchart illustrating a method for predicting equipment health index provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a monitoring data estimation model provided in an embodiment of this application; Figure 3 A flowchart illustrating a monitoring data estimation method provided in an embodiment of this application; Figure 4 A flowchart illustrating an initial monitoring data amplification method provided in this application embodiment; Figure 5 A flowchart illustrating another initial monitoring data amplification method provided in this application embodiment; Figure 6 A flowchart illustrating a monitoring data selection method provided in an embodiment of this application; Figure 7 A flowchart illustrating another monitoring data selection method provided in this application embodiment; Figure 8 This is a schematic diagram of the structure of a device for predicting the health index of an equipment, provided in an embodiment of this application. Detailed Implementation
[0012] The technical solutions in this application will now be described with reference to the accompanying drawings.
[0013] To facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and effect. For example, "first instruction" and "second instruction" are used to distinguish different user instructions and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.
[0014] It should be noted that, in this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.
[0015] Furthermore, "at least one" refers to one or more, while "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can mean: a, or b, or c, or a and b, or a and c, or b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0016] Furthermore, the terms "comprising" and "having," and any variations thereof, in the embodiments and drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0017] According to the first aspect of this application, with reference to the appendix Figure 1 As shown, this application claims protection for a method for predicting equipment health index, comprising the following steps: S101: Obtain the target health index range of the target wind power generation equipment within the target time period.
[0018] In this embodiment, the target wind power generation equipment is the wind power generation equipment whose health index is to be predicted. Wind power generation equipment refers to a device used to convert wind energy into electrical energy, typically a wind turbine generator set. The health index is used to quantify the comprehensive health status of the wind power generation equipment. The health index value can be preset to [0, 1], or other values. Specifically, when the health index of a wind power generation equipment is 0, it indicates that the equipment is completely ineffective or faulty; when the health index is 1, it indicates that the equipment is brand new or in a completely healthy state. The health index of the target wind power generation equipment is used as the target health index.
[0019] In this embodiment, the target time period refers to the future time interval to be predicted. The target value range refers to the possible value range of the target health index within the target time period. The target value range can be the same as the preset value range of the health index, or it can be obtained based on prior knowledge or preliminary analysis of current observation data. The target value range is based on the complete preset range, excluding health index candidate values that are physically impossible, contradict historical degradation trends, or significantly conflict with current multi-source data, in order to narrow down the candidate set.
[0020] For example, if the health index of the target wind power generation equipment was 0.85 at the previous moment, and under no special circumstances, the health status of the target wind power generation equipment is unlikely to deteriorate drastically in a short period of time, then values less than 0.4 or greater than 0.9 can be excluded.
[0021] For example, if the current vibration amplitude of the target wind power generation equipment is extremely low, values less than 0.4 can be manually excluded.
[0022] For example, if the target wind power generation equipment has just been replaced with a new product, values less than 0.5 can be manually excluded in order to adapt to wind power generation equipment with different production quality.
[0023] The above are merely examples and should not be construed as limiting the embodiments of this application.
[0024] S102: Based on the target value range and the monitoring data estimation model, obtain multiple monitoring data related to the target health index at multiple values. Each monitoring data corresponds to data of at least one data type in the preset data types, including mechanical vibration type, operating condition type and lubrication wear type.
[0025] It should be noted that the monitoring data refers to data that reflects the operating status of wind power generation equipment. The mechanical vibration data refers to the vibration data of key components of the wind power generation equipment, with a sampling frequency typically between 2kHz and 20kHz. This mechanical vibration data can be obtained from vibration data of components such as gearboxes and bearings collected by a Condition Monitoring System (CMS). When collecting this mechanical vibration data, the CMS system directly senses the mechanical vibration of the equipment during operation through acceleration sensors installed on the bearing housings or shells of key components such as gearboxes and main bearings. The mechanical vibration signals are then amplified, filtered, and converted from analog to digital to obtain the mechanical vibration data. The operating condition data refers to the operating condition data of the wind power generation equipment. This operating condition data can be collected by SCADA, with a sampling period of 1 second or 1 minute, etc. The operating condition data includes at least one of the following: wind speed, generator speed, output power / torque, gearbox oil temperature, and bearing temperature. The lubrication and wear data is used to characterize the degree of wear of the components of the wind power generation equipment. The data on the lubrication wear type can be obtained through online sensors or periodic offline laboratory analysis, with a frequency period of 1 day, 1 week, or 1 month.
[0026] In this embodiment, after training the monitoring data estimation model, the target value range is discretized to obtain multiple candidate values; each candidate value is input into the monitoring data estimation model, and the corresponding monitoring data is output. The monitoring data estimation model is obtained by training an initial estimation model based on the sample health index of the sample wind power generation equipment and the sample monitoring data corresponding to each sample health index.
[0027] It should be noted that the discretization of the target value range can be based on an equal step size, that is, the target value range is evenly divided into multiple candidate values according to a preset fixed step size. Alternatively, it can be based on a non-equal step size; for example, the target value range can be further divided into several sub-intervals. Based on the historical data distribution density of the target health index in each sub-interval, a short step size is used in sub-intervals with higher historical data frequency to improve candidate accuracy, while a long step size is used in sub-intervals with lower historical data frequency to reduce redundancy. Alternatively, some key candidate values can be manually selected, and other candidate values can be obtained by combining them using other discretization methods. The above are merely examples and should not be construed as limiting the embodiments of this application.
[0028] It should be noted that the input to the monitoring data estimation model can be the target health index, or it can include other covariate data related to the monitoring data, in order to improve the model's fit and estimation accuracy.
[0029] It should be noted that when the monitoring data includes data of multiple preset data types, the monitoring data estimation model can be a multi-output monitoring data estimation model, simultaneously outputting data of multiple preset data types, or multiple estimation sub-models can be constructed, each of which is used to estimate the corresponding preset data type based on the target health index. The above are merely examples and should not be construed as limiting the embodiments of this application.
[0030] In this embodiment, the first quarter of 2026 is taken as the target time period, and the quarterly abrasive particle concentration is used as the monitoring data for detailed explanation. The monitoring data estimation model can be constructed and trained based on a linear regression model, as follows: The initial estimation model is constructed. The inputs to the initial estimation model are the quarterly health index of the wind turbine and the ambient temperature at the wind turbine's location. The output is the abrasive particle concentration of the wind turbine in the corresponding quarter. The initial estimation model consists of an input layer, a linear combination layer, and an output layer. The input layer receives the values of the wind turbine's health index and the ambient temperature at the wind turbine's location to obtain an input vector, denoted as X. The linear combination layer is a weight coefficient matrix used to perform linear regression calculations based on the weight vector W and the bias term b. The output layer outputs the abrasive particle concentration of the corresponding wind turbine in the corresponding quarter, denoted as... The connection relationship of the initial estimation model can be expressed as follows: .
[0031] Dataset Preparation. Obtain the sample health index, average ambient temperature at the location of each sample wind turbine in each quarter, and abrasive particle concentration for each sample wind turbine in each quarter for the first 15 years of 2026. Pair the sample data according to time points; each sample data point includes the sample health index, ambient temperature, and abrasive particle concentration for one sample wind turbine. Obtain the training dataset based on all sample data. Divide the training dataset into a training set and a test set in an 8:2 ratio. The training set is used to train the initial estimation model, and the test set is used to test the prediction accuracy and other prediction performance of the initial estimation model.
[0032] It should be noted that the target sample wind power generation equipment may or may not be one of the sample wind power generation equipment. The sample health index can be regarded as the sample value of the health index of the sample power generation equipment.
[0033] Model Training. The linear combination layer of the initial estimation model is initialized. The sample health index and ambient temperature of the wind power generation equipment corresponding to each sample data point are input into the initial estimation model to output the estimated abrasive particle concentration, i.e., the estimated monitoring data. The optimal weight vector W and bias term b are solved using the least squares method, with the training objective being to minimize the mean square error between the estimated monitoring data and the sample monitoring data.
[0034] S103: Obtain target monitoring data that meets preset conditions from the plurality of monitoring data. The preset conditions are determined based on evaluation parameters, which are used to evaluate the rationality of the monitoring data. The higher the rationality of the monitoring data, the higher the probability that the candidate value used to estimate the monitoring data is the target value of the target health index; the lower the rationality of the monitoring data, the lower the probability that the candidate value used to estimate the monitoring data is the target value of the target health index.
[0035] In this embodiment, the evaluation parameters include at least one of probability P-value and divergence, and may also include other evaluation parameters.
[0036] It should be noted that the probability P-value is based on the idea of statistical hypothesis testing, which assesses whether the distribution of monitoring data estimated from a certain value of a health indicator differs significantly from historical normal fluctuation patterns. The larger the P-value, the better the estimated monitoring data matches historical patterns, and the more reasonable the corresponding health index value; the smaller the P-value, the lower the degree of matching between the estimated monitoring data and historical patterns, and the more unreasonable the corresponding health index value.
[0037] It should be noted that the divergence is used to measure the difference between the probability distribution of monitoring data estimated from a certain value of the health index and the probability distribution of historically observed monitoring data. Commonly used divergences include KL divergence (Kullback-Leibler Divergence) or JS divergence (Jensen-Shannon Divergence), but they are not limited to these. The larger the divergence, the greater the difference between the two probability distributions, and the lower the reasonable value of the corresponding health index; the smaller the divergence, the smaller the difference between the two probability distributions, and the higher the reasonable value of the corresponding health index.
[0038] It should be noted that the preset conditions can be obtained based on the evaluation parameters. When the evaluation parameters include a P-value, the preset condition can be that the P-value is less than a P-value threshold, or it can be that the monitoring data with the largest P-value is selected. When the evaluation parameters include divergence, the preset condition can be that the divergence is less than a divergence threshold, or it can be that the monitoring data with the smallest divergence is selected. When the preset conditions include both P-value and divergence, the preset condition can be that both the P-value and divergence are less than a divergence threshold simultaneously.
[0039] It should be noted that if only one monitoring data point meets the preset condition, that monitoring data point will be used as the target monitoring data point. If multiple monitoring data points meet the preset condition, one monitoring data point can be randomly selected from all the monitoring data points that meet the preset condition as the target monitoring data point. If no monitoring data point meets the preset condition, an early warning message can be generated to provide feedback to the user.
[0040] The above-described rationality assessment method is merely an example and should not be construed as limiting the embodiments of this application.
[0041] S104: The target value corresponding to the target monitoring data is used as the predicted value of the target health index in the target time period.
[0042] Understandably, this method generates a large number of candidate prediction schemes by traversing the entire target value range, rather than providing a single-point prediction. When the monitoring data estimated based on a certain value deviates significantly from the preset conditions, the optimization mechanism will prioritize eliminating it, while other candidate values are still available. When the monitoring data estimated based on a certain value is highly consistent with the preset conditions, the optimization mechanism will retain it. The optimization mechanism ensures that the final predicted value is the solution that best fits the preset conditions and data structure. Deviations in the input data are intercepted by the optimization mechanism, thereby significantly reducing the risk of prediction inaccuracies caused by local anomalies in the original data.
[0043] In some feasible implementations, the initial estimation model can also be constructed based on eXtremeGradient Boosting (XGBoost), which is obtained by integrating multiple sequentially trained Classification and Regression Trees (CART) as base learners. The initial estimation model includes a feature input layer, a multi-round iterative tree construction layer, and a result accumulation layer. The feature input layer receives the health index of the wind power generation equipment, the quarterly average temperature of the wind power generation equipment, and the quarterly average wind speed to obtain a feature vector. The multi-round iterative tree construction layer sequentially trains K regression decision trees, with each new tree learning the prediction residuals of the already trained regression decision trees. The result accumulation layer accumulates the prediction results of all leaf nodes of the regression decision trees tree by tree to obtain the average generator speed of the corresponding wind power generation equipment in the corresponding quarter, i.e., the estimated data. The initial estimation model can be expressed as: ; in, This refers to the monitoring and estimation data. Let X represent the prediction function of the k-th regression decision tree, where K is the total number of regression decision trees, and X represents the feature vector.
[0044] It should be noted that the method for constructing the training dataset for the initial estimation model can be similar to the methods described above, and will not be described in detail here. First, a baseline XGBoost model is established using default hyperparameters, and time series cross-validation is used to evaluate its performance. In time series cross-validation, data from earlier time periods are retained for training, and data from later time periods are used for validation, to simulate the data structure of real-world scenarios and avoid overfitting evaluation bias caused by the scrambling of time sequences in traditional random cross-validation. A hyperparameter search grid is defined, covering at least the following parameters: maximum tree depth, learning rate, subsampling rate, and regularization strength. A random search method is used to sample several sets of hyperparameter combinations from the parameter grid, and the root mean square error of the validation set in time series cross-validation is used as the evaluation metric to select the optimal hyperparameter combination.
[0045] In this embodiment, the maximum depth of the tree can be preset to 3-5 layers. The learning rate can be preset to [0.01, 0.1]. The subsampling rate can be preset to [0.8, 0.9]. The regularization strength can be preset to 1-2.
[0046] In this embodiment, the initial estimation model is trained using the optimal hyperparameter combination and the training set. An early stopping mechanism is employed during training to prevent the model from overlearning on the training data. If the evaluation metric on the test set no longer decreases within a specified number of consecutive epochs (e.g., 10 epochs), the addition of new trees is stopped, and the current number of trees, K, is used as the final model size.
[0047] In some feasible implementations, refer to the appendix. Figure 2 As shown, the initial estimation model can also be constructed based on a Long Short-Term Memory (LSTM) recurrent neural network. The initial estimation model includes an input layer 210, an LSTM layer 220, and a fully connected layer 230.
[0048] In this embodiment, the input layer 210 is used to receive input features, which can be represented as a three-dimensional tensor of (batch size, time step, feature dimension). The input features include the health index of the wind power generation equipment and the ambient wind speed at the location of the wind power generation equipment. Data from four quarters is used as a sample, meaning each sample data includes the health index, ambient temperature, and ambient wind speed of the same wind power generation equipment for four consecutive quarters.
[0049] In this embodiment, the LSTM layer 220 comprises two units, referred to as the first LSTM layer 221 and the second LSTM layer 222. Each unit utilizes three gating mechanisms—a forget gate, an input gate, and an output gate—to selectively remember and forget long-range dependencies in the input sequence, in conjunction with the transmission of cell states. The first LSTM layer 221 receives a three-dimensional tensor input (batch size, time step, feature dimension) and outputs a three-dimensional tensor (batch size, time step, hidden dimension). The hidden dimension of the output is 64. The second LSTM layer 222 receives a three-dimensional tensor input (batch size, time step, hidden dimension) and outputs a three-dimensional tensor input (batch size, time step, hidden dimension). The hidden dimension of the output is 64. A Dropout layer 223 is introduced between or after the two units to randomly discard a certain proportion of neuron activation values during training, preventing the model from becoming overly dependent on certain features and enhancing the model's robustness in noisy environments.
[0050] In this embodiment, the fully connected layer 230 is used to gradually compress the high-dimensional temporal features and regress them to the final data estimate output after the LSTM layer 220 completes the mapping from the original input features to the high-dimensional temporal features. The fully connected layer can be configured as a three-layer structure: the first fully connected layer 231 maps the hidden state of the LSTM layer 220 at the last time step to an intermediate dimension, that is, compresses the hidden dimension to 32 dimensions; the second fully connected layer 232 further compresses the hidden dimension to 16; and the third fully connected layer 233 outputs the monitoring estimate data. Meanwhile, except for the third fully connected layer 233 which directly outputs continuous values, the other fully connected layers are subsequently processed by batch normalization and ReLU activation function 234.
[0051] It should be noted that the training dataset for the initial estimation model can be similar to that described above, and will not be elaborated further here. The hyperparameters of the LSTM network are determined based on the feature dimensions and sample size of the dataset. The hidden dimension of the LSTM layer output can be set according to a hierarchical rule based on the number of features: if the feature dimension is less than 10, the hidden dimension is set to 32; if the feature dimension is between 10 and 30, the hidden dimension is set to 64; if the feature dimension is between 30 and 50, the hidden dimension is set to 128. The number of LSTM layers can be set to 2 to provide sufficient temporal modeling depth while keeping the number of parameters controllable. The dropout rate can be set to 0.2 to effectively prevent overfitting when the sample size is limited. The number of fully connected layers and intermediate dimensions can be adjusted according to the actual feature dimensions and task complexity.
[0052] In this implementation, the mean squared error (MSE) is used as the loss function for training, which is the mean of the squares of the differences between the model estimates and the actual observed values. The optimizer can be the Adam (Adaptive Moment Estimation) optimizer, which has adaptive learning rate adjustment capabilities, is insensitive to initial hyperparameter values, and converges quickly, making it suitable for training deep networks such as LSTM. The initial learning rate can be set to 0.001. The LSTM model is trained iteratively using the training set data. In each iteration, the training set is input into the model in batches, the loss function value is calculated, the gradients of the parameters of each layer are calculated using the backpropagation algorithm, and the parameters are updated by the Adam optimizer. The loss curves on the training and validation sets are monitored during training. An early stopping mechanism is employed: if the validation set loss no longer decreases within a certain number of consecutive iterations (e.g., 20 iterations), training is terminated and the parameters are restored to the checkpoint parameters where the validation set loss is minimized, avoiding overfitting to the training data.
[0053] It should be noted that the above-described methods for constructing and training the initial estimation model are merely examples and should not be construed as limiting the embodiments of this application.
[0054] Since the update frequency of the health index of wind power equipment may differ from the update frequency of the monitoring data (e.g., the update frequency of the health index is monthly or quarterly, while the update cycle of the monitoring data is 1 second or 1 minute), the monitoring data estimation model first maps the target health index to initial monitoring data of the same frequency and data type as the monitoring data, and then performs expansion processing based on the initial monitoring data. For example, when estimating the daily average rotational speed of a wind power equipment in the first quarter of 2026 based on its health index, the monitoring data estimation model first estimates the average rotational speed of the wind power equipment in the quarter based on its health index in the first quarter of 2026, and then performs expansion processing on the average rotational speed of the wind power equipment in the quarter to obtain the daily average rotational speed of the wind power equipment in the quarter.
[0055] See attached document Figure 3 As shown, in some feasible implementations, step S102 further includes the following sub-steps: S301: Based on a preset step size, the target value range is divided into multiple discrete values. The preset step size can be obtained by pre-setting, for example, 0.1 can be used as the preset step size.
[0056] S302: Input each discrete value into the monitoring data estimation model to obtain initial monitoring data corresponding to each discrete value. The initial monitoring data includes statistical values of data corresponding to each preset data type within the target time period. The update frequency of the initial monitoring data is the same as the update frequency of the health index of the target wind power generation equipment, and the preset data type of the initial monitoring data is the same as the preset data type of the monitoring data. The construction and training methods of the monitoring data estimation model are similar to those described above and will not be repeated here.
[0057] It should be noted that the initial detection data can be determined based on the update frequency of the target health index and the specific type of the monitoring data, such as total value, average value, median, month-on-month growth rate, etc.
[0058] S303: The initial monitoring data corresponding to each discrete value is amplified to obtain a corresponding monitoring data, until the plurality of monitoring data are obtained, wherein each monitoring data includes the value of each preset data type within a unit time period during the target time period.
[0059] It should be noted that the unit duration is shorter than the duration corresponding to the target time period. The amplification process is used to convert the low-frequency initial monitoring data into high-frequency monitoring data.
[0060] The target health index is typically updated quarterly or annually, while the monitoring data can be updated at intervals of 1 second, 1 minute, 1 hour, daily, monthly, and quarterly, leading to a mismatch in frequency. By expanding the sample size, the initial monitoring data with the same frequency as the target health index is decomposed into estimated results matching the frequency of the monitoring data. This expands the comparison from a single total quantity to a comparison of monitoring data with richer time-series characteristics, improving the discriminative power of the selection process and the reliability of the evaluation results.
[0061] See attached document Figure 4 As shown, in some feasible implementations, after step S302, the following sub-steps are further included: S401: Obtain historical data for each of the preset data types prior to the target time period. The historical data refers to the actual observation data of the same preset data type during historical periods prior to the target time period. The historical data can be selected with the same time span as the target time period, or it can be selected from the most recent several periods of data immediately adjacent to the target time period.
[0062] S402: Obtain the historical distribution of the historical data for each of the preset data types. The historical distribution refers to the characteristics extracted from the historical data that describe the relative proportion or distribution pattern of the data among various sub-cycles.
[0063] S403: Based on the historical distribution of each preset data type, the total value of the corresponding preset data type in the initial monitoring data within the target time period is expanded to obtain a corresponding monitoring data.
[0064] In this embodiment, the historical distribution can be represented by the proportion of each sub-cycle to the total amount of the entire cycle, the standardized value sequence of each sub-cycle, or other data structures that can characterize the distribution pattern.
[0065] For example, taking the cumulative effective value of gearbox vibration of a wind power generation device in the first quarter of 2026 as an example to expand the sample to the daily effective value of gearbox vibration of the same wind power generation device in the first quarter of 2026. The daily allocation ratio can be calculated based on the daily effective value of gearbox vibration of the wind power generation device in the first quarter of 2025. This daily allocation ratio represents the proportion of the effective value of gearbox vibration on that day to the cumulative effective value of gearbox vibration in the first quarter of 2025. Assuming that the daily allocation pattern in 2026 is the same as in 2025, the daily effective value of vibration of the wind power generation device in the first quarter of 2025 is used to expand the sample to the cumulative effective value of gearbox vibration of the wind power generation device in the first quarter of 2026 according to the said daily allocation ratio. The specific method is as follows: ; ; in, This represents the daily allocation ratio for day i. This represents the effective value of gearbox vibration on day i in the first quarter of 2025 for the wind power generation equipment. This indicates the cumulative effective value of gearbox vibration of the wind power generation equipment in the first quarter of 2025. This represents the estimated effective value of gearbox vibration on day i in the first quarter of 2026 for the wind power generation equipment. This indicates the cumulative effective value of gearbox vibration of the wind power generation equipment in the first quarter of 2026.
[0066] In some feasible implementations, when the monitored data is rotational speed, the sample can be expanded based on the average value. Compared to a total-based expansion method, the ratio between the daily rotational speed and the quarterly average rotational speed can be used to calculate the expansion ratio. The corresponding expansion result is obtained by multiplying the daily average rotational speed of the wind power equipment in the first quarter of 2026 with the daily expansion ratio.
[0067] To improve accuracy, the average of all expansion ratios corresponding to the same date in multiple historical years can be calculated and used as the expansion ratio.
[0068] It should be noted that the amplification results of each of the preset data types collectively constitute one monitoring data point corresponding to the same value of the target health index. If the monitoring data includes multiple preset data types, the amplification operation is performed independently for the historical data and historical distribution of each preset data type to obtain their respective monitoring data.
[0069] The expansion operation based on historical distribution only requires obtaining historical periodic data, calculating the daily distribution ratio, and performing a multiplicative distribution to complete the expansion. The calculation process is intuitive and transparent, making it easy for decision-makers to understand and verify.
[0070] See attached document Figure 5 In some feasible implementations, after step S401, the following sub-steps are further included: S501: Based on the historical data and data prediction model of each preset data type, obtain the data distribution of each preset data type in the target time period, denoted as the predicted distribution.
[0071] In this embodiment, the monitoring data prediction model is obtained by training an initial prediction model based on sample monitoring data of at least one preset data type. Each preset data type includes data for that preset data type within a first time period and its distribution within a second time period. The first time period precedes the second time period. The first time period serves as the historical observation period for input to the monitoring data prediction model, containing the actual observation sequence of the preset data type at each unit duration. The second time period serves as the time period for training the data prediction model as a supervision signal. The label or target learned by the model is the data distribution within the second time period, i.e., the relative distribution pattern of data values in each sub-cycle within the second time period. The length of each sub-cycle can be equal to the unit duration.
[0072] It should be noted that the method for constructing the monitoring data prediction model can be selected according to the actual situation. Some feasible implementation examples are given below.
[0073] In some feasible implementations, the actual observation sequence of each sub-cycle in the previous period immediately preceding the target time period can be obtained as the baseline sequence for expansion. The historical average month-on-month growth rate of the observations in each sub-cycle during the same period in the historical data is calculated. Starting from the last sub-cycle value of the baseline sequence, the values are successively multiplied by the historical average month-on-month growth rate of the corresponding sub-cycle to progressively estimate the preliminary values of each sub-cycle within the target time period, forming a preliminary estimation sequence. The sum of the values of each sub-cycle in the preliminary estimation sequence is calculated to obtain the preliminary estimated total. The ratio between the preliminary estimated total and the initial monitoring data output by the monitoring data estimation model is calculated to obtain the scaling factor. Each sub-cycle value in the preliminary estimation sequence is multiplied by the scaling factor to obtain the expanded monitoring data.
[0074] In some feasible implementations, monitoring data for the target time period is predicted based on historical data, and then adjusted based on the initial monitoring data. Methods such as the Seasonal Autoregressive Integrated Moving Average (SARIMA) model, classical seasonal decomposition and combination, and machine learning can be used.
[0075] For example, the initial prediction model is constructed using an LSTM model. The structure of the initial prediction model can be similar to the initial estimation model constructed based on the aforementioned LSTM model, and will not be described in detail here. The input of the initial prediction model is the monitoring data for the first time period, and the output is the data distribution for the second time period.
[0076] In terms of constructing the training dataset, the sample data collection method described in step S102 above can be referred to, and the training dataset can be constructed according to the pairing relationship of "first time period - second time period".
[0077] For example, the specific construction method is as follows: The length of the first time period is set to L, and the length of the second time period is set to H. The length of the second time period can be the same as the length of the target time period. Starting from the beginning of the historical sequence, the observations of L sub-cycles from day t to day t+L-1 are taken as the input sequence, and the data distributions of H sub-cycles from day t+L to day t+L+H-1 are taken as the supervision signal to form a training sample. Subsequent training samples are constructed sequentially according to a sliding window with a set step size (e.g., 1 day or 1 quarter). All training samples are divided into a training set, a validation set, and a test set in an 8:2 ratio.
[0078] In terms of model training and evaluation, the loss function can be the mean squared error. The loss function value that minimizes the difference between the predicted distribution and the true distribution on the training set is used to ensure that the model output approximates the true mapping relationship between the "previous sequence" and the "later distribution" in history as closely as possible. The training process and model optimization methods are similar to those for the monitoring data estimation model described above, and will not be repeated here.
[0079] S502: Based on the predicted distribution of each preset data type, the total value of the corresponding preset data type in the initial monitoring data within the target time period is expanded to obtain a corresponding monitoring data. The expansion operation can refer to step S403, and will not be repeated here.
[0080] In this technical solution, the data prediction model outputs the predicted distribution, that is, the relative relationship of each sub-cycle value. The expansion step applies the total value output by the monitoring data estimation model as a total constraint to the predicted distribution through proportional scaling or total alignment. This ensures that the expanded high-frequency monitoring data sequence strictly corresponds to the value of the currently assumed health index at the total level, and conforms to the temporal evolution law of the monitoring data itself at the distribution level.
[0081] In some feasible implementations, refer to the appendix. Figure 6 As shown, when the evaluation parameter is the P value, the following steps are included before step 104: S601: For any one of the multiple monitoring data, calculate the statistical value corresponding to each of the preset data types in the arbitrary monitoring data.
[0082] In this embodiment, the statistical value refers to a general quantitative indicator extracted from the high-frequency monitoring data sequence obtained by sample expansion, which can describe certain characteristics of the data sequence, such as the average value, variance, year-on-year growth rate, and month-on-month growth rate.
[0083] In this embodiment, for a monitoring data obtained by extrapolating a candidate value of the health index and expanding the sample, the mean and variance are calculated for each preset data type.
[0084] S602: Based on the statistical values corresponding to the data of each of the preset data types, obtain the probability P-value of any one of the monitoring data. Under the statistical hypothesis testing framework, the hypothesis is set as "the statistical characteristics of the monitoring data are consistent with the historical normal pattern" or "the monitoring data comes from the historical normal distribution". The probability P-value (or simply P-value) refers to the probability of observing the current statistical value or a more extreme result than the current statistical value under the condition that the hypothesis is true. The larger the P-value, the higher the consistency between the current monitoring data and historical patterns, and the more reasonable the corresponding candidate value of the health index is; the smaller the P-value, the greater the possibility that there is a significant difference between the current monitoring data and historical patterns, and the more likely the corresponding candidate value of the health index is to deviate from the normal range.
[0085] In this embodiment, hypothesis testing is performed on each statistical value of the monitoring data for each preset data type, and a P-value is calculated. That is, the P-value of the monitoring data includes the P-value corresponding to the mean and the P-value corresponding to the variance. After calculating all P-values for each preset data type, they are combined into the P-value corresponding to the monitoring data. Specific merging methods can include weighted average, minimum value method, etc.
[0086] The following is a detailed explanation of the steps for calculating the P-value corresponding to the average value for any one of the preset data types in the monitoring data. The method for calculating the P-value corresponding to the variance is similar and will not be repeated here.
[0087] Suppose that the preset data type has M historical contemporaries that can be referenced before the target time period. Each historical contemporaneous period and the target time period contain N sub-periods. For the m-th historical contemporaneous period (m=1, 2, ..., M), the sequence of actual observations for each sub-period within this historical contemporaneous period is denoted as […]. The average for the same period in history is denoted as The historical mean is a sample set consisting of the averages of M historical data from the same period, denoted as . Calculate the mean of the sample set of historical means (denoted as ). ) and standard deviation (denoted as ).
[0088] The estimated monitoring data is denoted as the data of the preset data type. Calculate their average value, and denote it as . .
[0089] Setting the null hypothesis The estimated average value of the monitoring data obtained in this preset data type. The sample set corresponding to the historical mean comes from the same population. An alternative hypothesis is defined. The estimated average value of the monitoring data obtained in this preset data type. There is a significant difference from historical normal levels. Under the condition that H0 holds, the t-statistic approximately follows a sequence with M degrees of freedom. The t-distribution of 1.
[0090] Constructing the t-statistic based on a two-tailed test: ; Calculate the P-value corresponding to the average value using the following method: ; Where |t| is the absolute value of the t-statistic; Let M be the degree of freedom The cumulative distribution function of the t-distribution of 1, i.e., in a distribution with M degrees of freedom In a t-distribution of 1, the probability that a random variable takes a value less than or equal to a given value |t| is expressed as the sum of the probabilities of the two tails of a random variable whose absolute value exceeds |t| in the t-distribution.
[0091] It should be noted that the above are merely examples and should not be construed as limiting the embodiments of this application.
[0092] S603: From the plurality of monitoring data, determine the monitoring data whose P-value is less than the P-value threshold as the target monitoring data. The P-value threshold refers to a pre-set critical value standard used to determine whether the P-value meets the significance requirement. The P-value threshold can be pre-set.
[0093] In this embodiment, after obtaining the P-value of the monitoring data corresponding to each candidate value of the target health index, each P-value is compared with the P-value threshold. Monitoring data with P-values less than the P-value threshold are selected as the range of target monitoring data. If multiple monitoring data meet the preset conditions, the monitoring data corresponding to the minimum P-value can be further selected or the target monitoring data can be obtained through random selection. If no monitoring data meets the preset conditions, an early warning message can be generated, or the monitoring data corresponding to the minimum P-value can be directly taken as the target monitoring data.
[0094] The probability P-value is used as a quantitative indicator for judging reasonableness. The smaller the P-value, the lower the probability of observing the monitoring data under the assumption of historical patterns, and the lower the reasonableness of the target health index for that candidate value. This mechanism eliminates the subjectivity and inconsistency of human judgment and improves the traceability and interpretability of the entire selection process.
[0095] Each data point is evaluated from multiple dimensions, including mean and variance, and the degree of conformity between the monitored data and historical patterns is characterized by both central tendency and dispersion. Cross-evaluation ensures that random biases in a single dimension do not lead to overall evaluation failure, thus enhancing the robustness of the evaluation results.
[0096] In some feasible implementations, the evaluation parameter is the divergence, a quantitative indicator used to measure the degree of difference between two probability distributions. The smaller the divergence value, the closer the two distributions are, the more consistent the calculated data is with historical patterns, and the more reasonable the corresponding candidate health index value. Commonly used divergence indicators include Kullback-Leibler divergence (KL divergence) and Jensen-Shannon divergence (JS divergence). One or both can be calculated simultaneously, depending on the sensitivity to distribution differences.
[0097] See attached document Figure 7 As shown, the following steps are included before step 104: S701: Obtain the historical data of each of the preset data types prior to the target time period.
[0098] S702: For any one of the multiple monitoring data, calculate the divergence value of each of the preset data types of the monitoring data based on the similarity between the historical data of each preset data type and the data corresponding to the preset data type in the arbitrary data.
[0099] In this embodiment, historical data and all estimated monitoring data are binned to construct corresponding frequency distribution histograms. The bin boundaries for historical data and estimated monitoring data are identical to ensure that the bins in their histograms are completely consistent. Specifically, the union of the data values of the preset data type in the historical data and all estimated monitoring data is taken. The minimum value in this union is used as the lower bound of the preset data type, and the maximum value is used as the upper bound. The interval formed by the lower bound of the preset data type and the upper bound is divided into B bins of equal width. The number of bins B can be preset.
[0100] Based on the boundaries of each bucket, the frequency of each monitored data falling into each bucket is statistically analyzed using historical data and estimated data, resulting in a frequency vector for each bucket. The sum of all elements in each frequency vector is 1.
[0101] Furthermore, to avoid the influence of zero frequencies on subsequent divergence calculations, the frequencies of each monitored data point falling into each bucket can be statistically analyzed using both historical data and estimated data, resulting in their respective frequency vectors. Laplace smoothing is then applied to the frequencies of each bucket, i.e., a minimum positive value is added to the frequency of each bucket, such as 1 or 0.1. Finally, normalization is performed on each processed frequency vector to obtain the final frequency vector.
[0102] In this embodiment, for each preset data type, the frequency vector corresponding to the historical data is denoted as... The frequency vector of the j-th monitoring data obtained by estimation is denoted as... The frequency vector of the j-th monitoring data is calculated and denoted as . corresponding divergence .
[0103] Taking the KL divergence value as an example, the calculation method of the KL divergence value includes: ; in, Let b be the probability value of the frequency vector corresponding to the historical data in the b-th bucket. The probability value of the frequency vector of the j-th monitoring data obtained is the probability value of the b-th bucket.
[0104] S703: Based on the divergence value corresponding to each of the preset data types, obtain the comprehensive divergence value of any one of the monitoring data.
[0105] In this embodiment, after calculating the divergence values of each preset data type, all divergence values of the target health index corresponding to the same candidate value are merged into the comprehensive divergence value. Optional merging methods include weighted average method, maximum value method, etc.
[0106] S704: From the multiple monitoring data, determine the data whose comprehensive divergence value is less than the divergence threshold as the target monitoring data. The divergence threshold can be preset.
[0107] In this embodiment, the comprehensive divergence value corresponding to each monitoring data is compared with the divergence threshold, and data with a comprehensive divergence value less than the divergence threshold are selected as target monitoring data. If multiple candidate monitoring data all meet the preset conditions, the monitoring data corresponding to the minimum comprehensive divergence value can be further selected or randomly selected to obtain the target monitoring data. If no monitoring data meets the preset conditions, an early warning message can be generated, or the monitoring data corresponding to the minimum comprehensive divergence value can be directly selected as the target monitoring data.
[0108] Divergence measures the relative differences at the probability distribution level and naturally possesses dimensionlessness and normalization properties. Their distribution vectors, after binning and normalization, are compared within a unified probability space. This allows for the comprehensive consideration of the rationality of heterogeneous data such as mechanical vibration, operating conditions, and lubrication wear on the same scale, reducing the risk of evaluation distortion caused by dimensional differences.
[0109] Divergence measures the differences between the entire probability distribution and can capture more comprehensive information about the distribution pattern. For example, even if the mean and variance of the estimated monitoring data are consistent with the historical data, skewed changes in the distribution shape can be captured by divergence, thus improving the reasonableness of the predicted values of the wind power generation equipment resistance index.
[0110] In one feasible implementation, when the evaluation parameters include both the P-value and the divergence, the method further includes determining a comprehensive score for any one of the monitoring data based on the probability P-value and the comprehensive divergence value. The preset condition includes a comprehensive score greater than a scoring threshold. Step S104 includes determining the monitoring data with a comprehensive score greater than the scoring threshold from the plurality of monitoring data as the target monitoring data.
[0111] In this embodiment, the method for calculating the probability P value of any monitoring data and the method for calculating the comprehensive divergence value are similar to the methods described above, and will not be repeated here.
[0112] The comprehensive divergence value can first be positively transformed by taking a negative value, taking the reciprocal, or taking exponential decay, so that the larger the value, the better the rationality it represents. Then, the rate P value and the processed comprehensive divergence value are weighted and summed according to preset weights to obtain the corresponding comprehensive score.
[0113] It should be noted that the above methods are merely examples and should not be considered as limitations on the embodiments of this application.
[0114] According to the second aspect of this application, with reference to the appendix Figure 8 As shown, this application provides a device for predicting equipment health index, comprising: The acquisition module 810 is used to acquire the target health index range of the target wind power generation equipment within the target time period. The estimation module 820 is used to obtain multiple monitoring data related to the target health index at multiple values based on the target value range and the monitoring data estimation model. Each monitoring data corresponds to data of at least one data type in the preset data types. The preset data types include mechanical vibration type, operating condition type and lubrication wear type. The monitoring data estimation model is obtained by training an initial estimation model based on the sample health index of the sample wind power generation equipment and the sample monitoring data corresponding to each sample health index. Selection module 830 is used to obtain target monitoring data that meets preset conditions from the plurality of monitoring data. The preset conditions are determined according to evaluation parameters, which include at least one of probability P-value and divergence. The output module 840 is used to use the target value corresponding to the target monitoring data as the predicted value of the target health index in the target time period.
[0115] In some feasible implementations, the estimation module includes a first submodule 821, a second submodule 822, and a third submodule 823: The first submodule 821 is used to divide the target value range into multiple discrete values according to a preset step size; The second submodule 822 is used to input each of the discrete values into the monitoring data estimation model to obtain initial monitoring data corresponding to each of the discrete values. The initial monitoring data includes statistical values of data corresponding to each of the preset data types within the target time period. The third submodule 823 is used to perform a sampling process on the initial monitoring data corresponding to each discrete value to obtain a corresponding monitoring data, until the multiple monitoring data are obtained. Each monitoring data includes the value of each preset data type within a unit time period during the target time period.
[0116] In some feasible implementations, the third submodule 823 further includes: Obtain historical data for each of the preset data types prior to the target time period; Obtain the historical distribution of the historical data for each of the preset data types; Based on the historical distribution of each preset data type, the total value of the corresponding preset data type in the initial monitoring data within the target time period is expanded to obtain a corresponding monitoring data.
[0117] In some feasible implementations, the third submodule 823 further includes: Obtain historical data for each of the preset data types prior to the target time period; Based on the historical data and data prediction model of each preset data type, the predicted distribution of each preset data type in the target time period is obtained. The monitoring data prediction model is obtained by training an initial prediction model based on sample monitoring data of at least one preset data type. The sample monitoring data of each preset data type includes the monitoring data of the corresponding preset data type in the first time period and the data distribution in the second time period. The first time period is before the second time period. Based on the predicted distribution of each preset data type, the total value of the corresponding preset data type in the initial monitoring data within the target time period is expanded to obtain a corresponding monitoring data.
[0118] In some feasible implementations, the evaluation parameter is the P value, the preset condition includes the P value being less than a P value threshold, and the system further includes an evaluation module 850, which is used for: For any one of the multiple monitoring data, calculate the statistical value corresponding to each of the preset data types in the monitoring data, where the statistical value includes the mean and variance; Based on the statistical values corresponding to the data of each of the preset data types, the probability P value of any one monitoring data is obtained; The selection module 830 is further configured to determine, from the plurality of monitoring data, the monitoring data whose P value is less than the P value threshold as the target monitoring data.
[0119] In some feasible implementations, the evaluation parameter is the divergence, the preset condition includes a comprehensive divergence value less than a divergence threshold, and the system further includes an evaluation module 850, which is used for: Obtain historical data for each of the preset data types prior to the target time period; For any one of the multiple monitoring data, the divergence value of each of the preset data types is calculated based on the similarity between the historical data of each preset data type and the data of the corresponding preset data type in the any one monitoring data. Based on the divergence value corresponding to each of the preset data types, the comprehensive divergence value of any monitoring data is obtained; The selection module 830 is further configured to determine, from the plurality of monitoring data, the monitoring data whose comprehensive divergence value is less than the divergence threshold as the target monitoring data.
[0120] In some feasible implementations, the evaluation parameters are the P-value and the divergence, the preset conditions include a comprehensive score greater than a scoring threshold, and the system further includes an evaluation module 850, which is used for: For any one of the multiple monitoring data, calculate the statistical value corresponding to each of the preset data types in that one monitoring data, where the statistical value includes the mean and variance; and obtain the probability P value of the one monitoring data based on the statistical value corresponding to each of the preset data types. Obtain historical data for each of the preset data types prior to the target time period; Based on the similarity between the historical data of each preset data type and the data corresponding to the preset data type in any monitoring data, calculate the divergence value of each preset data type in any monitoring data; based on the divergence value corresponding to each preset data type, obtain the comprehensive divergence value of any monitoring data. The comprehensive score of any one of the monitoring data is determined based on the probability P-value and the comprehensive divergence value. The selection module 830 is further configured to determine, from the plurality of monitoring data, the monitoring data whose comprehensive score is greater than the scoring threshold as the target monitoring data.
[0121] According to a third aspect of this application, this application provides a computer device including a memory and a processor, the memory having a computer program executable on the processor, the processor executing the program to implement the steps of the aforementioned method.
[0122] According to a fourth aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the aforementioned method.
[0123] In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software. The steps of the method disclosed in the embodiments of this application can be directly manifested as execution by a hardware processor, or as a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor executes the instructions in the memory, combining them with its hardware to complete the steps of the above method. To avoid repetition, detailed descriptions are omitted here.
[0124] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0125] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0126] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0127] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0128] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0129] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0130] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for predicting equipment health index, characterized in that, include: Obtain the target health index range of the target wind power generation equipment within the target time period; Based on the target value range and the monitoring data estimation model, multiple monitoring data related to the target health index at multiple values are obtained. Each monitoring data corresponds to data of at least one data type in the preset data types, including mechanical vibration type, operating condition type and lubrication wear type. The monitoring data estimation model is obtained by training the initial estimation model based on the sample health index of the sample wind power generation equipment and the sample monitoring data corresponding to each sample health index. Target monitoring data that meets preset conditions are obtained from the plurality of monitoring data. The preset conditions are determined based on evaluation parameters, which include at least one of probability P-value and divergence. The target value corresponding to the target monitoring data is used as the predicted value of the target health index in the target time period; The step of obtaining multiple monitoring data related to the target health index across multiple values based on the target value range and the monitoring data estimation model includes: According to the preset step size, the target value range is divided into multiple discrete values; Each discrete value is input into the monitoring data estimation model to obtain initial monitoring data corresponding to each discrete value. The initial monitoring data is the total value of data corresponding to each preset data type within the target time period. The initial monitoring data corresponding to each discrete value is amplified to obtain a corresponding monitoring data, until the plurality of monitoring data are obtained. Each monitoring data includes the value of each preset data type within a unit time period during the target time period.
2. The method according to claim 1, characterized in that, The step of amplifying the initial monitoring data corresponding to each discrete value to obtain a corresponding monitoring data includes: Obtain historical data for each of the preset data types prior to the target time period; Obtain the historical distribution of the historical data for each of the preset data types; Based on the historical distribution of each preset data type, the total value of the corresponding preset data type in the initial monitoring data within the target time period is expanded to obtain a corresponding monitoring data.
3. The method according to claim 1, characterized in that, The step of amplifying the initial monitoring data corresponding to each discrete value to obtain a corresponding monitoring data includes: Obtain historical data for each of the preset data types prior to the target time period; Based on the historical data and data prediction model of each preset data type, the predicted distribution of each preset data type in the target time period is obtained. The data prediction model is obtained by training an initial prediction model based on sample monitoring data of at least one preset data type. The sample monitoring data of each preset data type includes the monitoring data of the corresponding preset data type in the first time period and the data distribution in the second time period. The first time period is before the second time period. Based on the predicted distribution of each preset data type, the total value of the corresponding preset data type in the initial monitoring data within the target time period is expanded to obtain a corresponding monitoring data.
4. The method according to any one of claims 1-3, characterized in that, The evaluation parameter is the P value. Before obtaining the target monitoring data that meets the preset conditions from the plurality of monitoring data, the method further includes: For any one of the multiple monitoring data, calculate the statistical value corresponding to each of the preset data types in the any one monitoring data, the statistical value including the mean and variance; Based on the statistical values corresponding to the data of each of the preset data types, the probability P value of any one monitoring data is obtained; The preset conditions include a P-value less than a P-value threshold. Target monitoring data satisfying the preset conditions is obtained from the plurality of monitoring data, including: From the multiple monitoring data, the monitoring data with a P value less than the P value threshold are determined as the target monitoring data.
5. The method according to any one of claims 1-3, characterized in that, The evaluation parameter is the divergence. Before obtaining the target monitoring data that meets the preset conditions from the plurality of monitoring data, the method further includes: Obtain historical data for each of the preset data types prior to the target time period; For any one of the multiple monitoring data, the divergence value of each of the preset data types is calculated based on the similarity between the historical data of each preset data type and the data of the corresponding preset data type in the any one monitoring data. Based on the divergence value corresponding to each of the preset data types, the comprehensive divergence value of any monitoring data is obtained; The preset conditions include a comprehensive divergence value less than a divergence threshold. Target monitoring data satisfying the preset conditions is obtained from the plurality of monitoring data, including: From the multiple monitoring data, the monitoring data with a comprehensive divergence value less than the divergence threshold is determined as the target monitoring data.
6. The method according to any one of claims 1-3, characterized in that, The evaluation parameters are the P-value and the divergence. Before obtaining the target monitoring data that meets the preset conditions from the plurality of monitoring data, the method further includes: For any one of the multiple monitoring data, calculate the statistical value corresponding to each of the preset data types in that one monitoring data, where the statistical value includes the mean and variance; and obtain the probability P value of the one monitoring data based on the statistical value corresponding to each of the preset data types. Obtain historical data for each of the preset data types prior to the target time period; Based on the similarity between the historical data of each preset data type and the data corresponding to the preset data type in any monitoring data, calculate the divergence value of each preset data type in any monitoring data; based on the divergence value corresponding to each preset data type, obtain the comprehensive divergence value of any monitoring data. The comprehensive score of any one of the monitoring data is determined based on the probability P-value and the comprehensive divergence value. The preset conditions include a comprehensive score greater than a score threshold, and target monitoring data that meets the preset conditions is obtained from the multiple monitoring data, including: From the multiple monitoring data, the monitoring data with a comprehensive score greater than the score threshold is determined as the target monitoring data.
7. A device for predicting equipment health index, characterized in that, include: The acquisition module is used to obtain the target health index range of the target wind power generation equipment within the target time period; The estimation module is used to obtain multiple monitoring data related to the target health index at multiple values based on the target value range and the monitoring data estimation model. Each monitoring data corresponds to data of at least one data type in the preset data types. The preset data types include mechanical vibration type, operating condition type and lubrication wear type. The monitoring data estimation model is obtained by training an initial estimation model based on the sample health index of the sample wind power generation equipment and the sample monitoring data corresponding to each sample health index. The selection module is used to obtain target monitoring data that meets preset conditions from the plurality of monitoring data. The preset conditions are determined based on evaluation parameters, which include at least one of probability P-value and divergence. The output module is used to take the target value corresponding to the target monitoring data as the predicted value of the target health index in the target time period; The estimation module includes a first submodule, a second submodule, and a third submodule: The first submodule is used to divide the target value range into multiple discrete values according to a preset step size; The second submodule is used to input each of the discrete values into the monitoring data estimation model to obtain initial monitoring data corresponding to each of the discrete values. The initial monitoring data includes statistical values of data corresponding to each of the preset data types within the target time period. The third submodule is used to amplify the initial monitoring data corresponding to each discrete value to obtain a corresponding monitoring data, until the multiple monitoring data are obtained. Each monitoring data includes the value of each preset data type within a unit time period during the target time period.
8. A computer device comprising a memory and a processor, the memory having a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Equipment health monitoring method and device, equipment and storage medium
CN121765342A
A method, device and medium for health risk assessment of an environmental pollution event
CN122155119A