Equipment health monitoring method and device, equipment and storage medium
By collecting equipment operating parameters and dynamically adjusting the reference benchmark using a baseline generation model and feature extraction network, the accuracy problem of equipment health monitoring under dynamic operating conditions is solved, achieving higher monitoring accuracy and reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-03-31
AI Technical Summary
Existing equipment health monitoring methods use fixed benchmarks, which cannot adapt to dynamic changes in operating conditions, resulting in insufficient monitoring accuracy.
By collecting the operating parameters of the target equipment, calculating the reference benchmark using a pre-trained baseline generation model, comparing the monitoring data to obtain the difference data, and combining feature extraction and distribution analysis, the health status judgment is dynamically adjusted.
It improves the accuracy and reliability of equipment health monitoring, adapts to different operating conditions, and reduces misjudgments and omissions.
Smart Images

Figure CN121765342A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, device, and storage medium for monitoring equipment health. Background Technology
[0002] With the increasing intelligence level of the manufacturing industry, equipment health monitoring has become a key technology for ensuring production continuity and reducing maintenance costs. Traditional equipment monitoring methods mainly rely on fixed thresholds or historical statistical data to establish monitoring benchmarks, and judge the health status of equipment by comparing the deviation of real-time operating parameters with preset benchmarks.
[0003] However, in actual production environments, equipment operating conditions are constantly and dynamically changing. Different combinations of factors such as load, speed, and ambient temperature can lead to significant differences in normal operating parameters. Existing fixed-benchmark monitoring methods cannot adapt to these dynamic changes in operating conditions. They are prone to misinterpreting normal parameter fluctuations as equipment malfunctions during operating condition transitions, or failing to detect actual equipment degradation in a timely manner because parameters are still within fixed threshold ranges, resulting in insufficient accuracy and reliability of monitoring. Summary of the Invention
[0004] The main objective of this invention is to solve the technical problem that existing equipment health monitoring methods use fixed benchmarks, which cannot adapt to dynamic changes in operating conditions, resulting in insufficient monitoring accuracy. This invention provides a method for monitoring equipment health, the method comprising: The operating parameters of the target equipment during operation are collected to obtain monitoring data corresponding to the operating status of the target equipment; A reference benchmark is calculated based on the monitoring data using a pre-trained baseline generation model, and the monitoring data is compared with the reference benchmark to obtain difference data characterizing the health status of the target device. Feature extraction and distribution analysis are performed on the differential data to obtain a health feature representation reflecting the degree of equipment degradation. The health status of the target device is determined based on the health characteristics, and the health monitoring results of the target device are obtained.
[0005] The present invention also provides a device for monitoring equipment health, the device comprising: The data acquisition module is used to collect the operating parameters of the target equipment during operation and obtain monitoring data corresponding to the operating status of the target equipment. The baseline comparison module is used to calculate a reference benchmark based on the monitoring data using a pre-trained baseline generation model, and compare the monitoring data with the reference benchmark to obtain difference data characterizing the health status of the target device. The feature analysis module is used to extract features and perform distribution analysis on the differential data to obtain a health feature representation that reflects the degree of equipment degradation. The status determination module is used to determine the health status of the target device based on the health characteristics, and obtain the health monitoring results of the target device.
[0006] The present invention also provides a device health monitoring apparatus, comprising: a memory and at least one processor, wherein the memory stores instructions, and the memory and the at least one processor are interconnected via a line; the at least one processor invokes the instructions in the memory to cause the device health monitoring apparatus to perform the steps of the device health monitoring method described above.
[0007] The present invention also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the steps of the above-described device health monitoring method.
[0008] The aforementioned equipment health monitoring method, device, equipment, and storage medium obtain monitoring data by collecting operating parameters of the target equipment; calculate a reference benchmark based on the monitoring data using a pre-trained baseline generation model, and compare the monitoring data with the reference benchmark to obtain difference data; extract features and perform distribution analysis on the difference data to obtain health feature representations; and determine the health status based on the health feature representations to obtain the monitoring results. This method dynamically calculates the reference benchmark based on real-time operating conditions using a baseline generation model, replacing the traditional fixed threshold method. It employs a deep feature extraction network to capture equipment degradation features and uses probability distribution divergence values to quantify the health status, solving the problem of poor adaptability of existing methods under dynamic operating conditions and improving the accuracy and reliability of equipment health monitoring.
[0009] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0010] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0011] Figure 1This is a schematic diagram of the first embodiment of the device health monitoring method in this invention; Figure 2 This is a schematic diagram of a second embodiment of the device health monitoring method in this invention; Figure 3 This is a schematic diagram of one embodiment of the device health monitoring device in this invention; Figure 4 This is a schematic diagram of one embodiment of the device health monitoring device in this invention. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] The terms "comprising" and "having," and any variations thereof, used in the embodiments of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0014] To facilitate understanding of this embodiment, a device health monitoring method disclosed in this invention will first be described in detail. For example... Figure 1 As shown, this method includes the following steps: 101. Collect the operating parameters of the target equipment during operation to obtain monitoring data corresponding to the operating status of the target equipment; In this embodiment, the step of collecting operating parameters of the target device during operation to obtain monitoring data corresponding to the device's operating status includes: determining target monitoring parameters for the target device, where the target monitoring parameters are key operating parameters characterizing the device's health status; collecting multiple operating parameters during the operation of the target device, calculating the correlation between each operating parameter and the target monitoring parameter to obtain a correlation degree value for each operating parameter; sorting the operating parameters according to the correlation degree values, selecting operating parameters with correlation degree values greater than a preset threshold to obtain an input parameter set; synchronously collecting the input parameter set and the target monitoring parameters within a preset time window to obtain time series data containing timestamps and parameter values, and normalizing the time series data to obtain the monitoring data.
[0015] Specifically, target monitoring parameters refer to operating parameters that directly reflect changes in the health status of equipment. Taking the engine cooling system as an example, the target monitoring parameter could be the engine outlet coolant temperature. This parameter comprehensively reflects the heat dissipation performance of the cooling system and the engine's thermal load status. When the cooling system deteriorates, such as when the radiator is clogged or the cooling fan efficiency decreases, the coolant temperature will rise abnormally. Therefore, this parameter is suitable as a characterization parameter for health status.
[0016] After determining the target monitoring parameter, it is necessary to collect multiple operating condition parameters that may affect it. These operating condition parameters include, but are not limited to, engine speed, fuel supply, intake pressure, exhaust temperature, ambient air temperature, and vehicle speed. These operating condition parameters affect the value of the target monitoring parameter from different dimensions. Among them, engine speed and fuel supply mainly affect the heat generation of the engine, intake pressure and exhaust temperature reflect the engine's load status, and ambient air temperature and vehicle speed affect the heat dissipation efficiency of the cooling system.
[0017] For correlation calculation, this embodiment uses the Pearson correlation coefficient as a quantitative indicator of correlation. The Pearson correlation coefficient measures the degree of linear correlation between two variables, with a value ranging from -1 to 1. The closer the absolute value is to 1, the stronger the correlation. During calculation, the sampled value of each operating condition parameter needs to be paired with the sampled value of the target monitoring parameter at the corresponding time. The ratio of the covariance to the standard deviation of the two sets of data is calculated using statistical methods to obtain the correlation coefficient.
[0018] In one embodiment, historical data of the device operating in a healthy state for a period of time can be collected. This historical data includes time series of all candidate operating condition parameters and target monitoring parameters. For example, collecting data for one hour with a sampling period of one second can yield 3600 data samples. For each operating condition parameter, the Pearson correlation coefficient between it and the target monitoring parameter is calculated to obtain the correlation degree value of each operating condition parameter.
[0019] Specifically, in a monitoring application of an engine cooling system, the correlation values of various operating parameters were calculated as follows: engine speed had a correlation of 0.86, fuel supply was 0.83, intake pressure was 0.79, exhaust temperature was 0.76, ambient air temperature was 0.76, and vehicle speed was 0.76. It can be seen that engine speed has the highest correlation with coolant temperature. This is because engine speed directly determines the engine's thermal load and the coolant pump's flow rate; the higher the engine speed, the more heat the engine generates, and the faster the coolant circulation speed.
[0020] After obtaining the correlation coefficient values, a selection is made based on a preset threshold. This preset threshold can be set to 0.7, meaning that parameters with an absolute correlation coefficient greater than 0.7 are selected as modeling inputs. This threshold is chosen based on the statistical definition of strong correlation; generally, an absolute correlation coefficient greater than 0.7 indicates a strong correlation between two variables and has significant predictive value. In the example above, the correlation coefficients of all six parameters are greater than 0.7; therefore, these six parameters constitute the input parameter set and will serve as the input variables for the subsequent baseline generation model.
[0021] It should be noted that the correlation between operating parameters and target monitoring parameters may vary for different types of equipment. For example, for CNC machine tools, spindle temperature may be highly correlated with parameters such as spindle speed, cutting load, and coolant flow rate; for compressors, exhaust temperature may be correlated with parameters such as compression ratio, intake air temperature, and operating frequency. In some embodiments, other correlation measurement methods such as mutual information and Spearman's rank correlation coefficient can also be used to accommodate nonlinear correlations and ensure that operating parameters that have a significant impact on the target parameters can be accurately identified.
[0022] After determining the set of input parameters, synchronous timing data acquisition is required. The preset time window can be set to match the working cycle of the equipment. For example, for a reciprocating engine, it can be set to the complete working cycle time; for continuously operating equipment, it can be set to the duration of a specific operating condition. In this embodiment, the time window can be set to 10 minutes to 1 hour to ensure that sufficient data samples are collected to reflect the dynamic characteristics of the equipment under different operating conditions. Synchronous acquisition refers to measuring all operating condition parameters and target monitoring parameters in the input parameter set at the same time to ensure that there is a temporal correspondence between the values of each parameter.
[0023] The collected time-series data includes two parts: timestamps and parameter values. The timestamp records the acquisition time of each data point, typically in the format of year, month, day, hour, minute, and second, accurate to the second or millisecond level. The parameter values record the actual measured values of each operating condition parameter at that moment. For example, a data record can be represented as: timestamp 2025-01-15 10:30:05, engine speed 1500 rpm, fuel supply 45 mg / cycle, coolant temperature 85℃, intake pressure 1.2 bar, exhaust temperature 420℃, ambient temperature 25℃, vehicle speed 60 km / h.
[0024] Because the dimensions and numerical ranges of parameters vary significantly under different operating conditions—for example, engine speed may range from 800 to 2200 rpm, ambient temperature from -20°C to 40°C, and fuel supply from 20 to 80 mg / cycle—normalization is necessary to bring all parameters to the same numerical scale. This embodiment employs the min-max normalization method, which maps the value of each parameter to a range of 0 to 1.
[0025] The specific normalization process is as follows: First, determine the minimum and maximum values of the parameter across all collected data. Then, for each data point, subtract the minimum value from its original value, and divide by the difference between the maximum and minimum values to obtain the normalized value. For example, if the engine speed in a certain data collection is 1500 rpm, the minimum value of this parameter in the dataset is 800 rpm, and the maximum value is 2200 rpm, then the normalized value is (1500-800) / (2200-800) = 0.5. Through this normalization process, the values of all parameters are mapped to the range of 0 to 1, eliminating the differences in dimensions and numerical scales between different parameters.
[0026] The necessity of normalization lies in the fact that if raw numerical values are used directly for modeling, parameters with larger numerical ranges will dominate during model training, while the influence of parameters with smaller numerical ranges may be ignored. Normalization ensures that each parameter contributes more reasonably to the model, enabling the model to comprehensively consider the influence of all input parameters and improving the accuracy of baseline generation and health assessment.
[0027] 102. A reference benchmark is calculated based on the monitoring data using a pre-trained baseline generation model, and the monitoring data is compared with the reference benchmark to obtain difference data characterizing the health status of the target device; In this embodiment, the step of calculating a reference benchmark based on the monitoring data using a pre-trained baseline generation model and comparing the monitoring data with the reference benchmark to obtain difference data characterizing the health status of the target device includes: inputting the operating parameters in the monitoring data into the baseline generation model, obtaining predicted values for each layer through cascaded prediction of a multi-layer decision tree; performing a weighted summation of the predicted values for each layer to obtain a reference benchmark corresponding to the current operating condition; and calculating the difference between the actual measured values in the monitoring data and the reference benchmark at corresponding time points to obtain the difference data.
[0028] Specifically, the baseline generation model is constructed using the gradient boosting decision tree algorithm. This algorithm builds multiple decision trees sequentially, with each new tree specifically designed to fit the prediction residuals of all previous trees, thereby progressively improving the model's prediction accuracy. Unlike parallel ensemble algorithms such as random forests, the core advantage of gradient boosting decision trees lies in their ability to iteratively optimize and approximate the true value, making them particularly suitable for handling complex nonlinear relationships between operating parameters and equipment parameters.
[0029] After the monitoring data is input into the baseline generation model, it first enters the first-layer decision tree for prediction. This decision tree makes decisions based on the input operating parameters, such as engine speed of 1500 rpm and fuel supply of 45 mg / cycle, through a series of decision nodes. Each decision node sets a splitting condition for a specific operating parameter, such as determining whether the engine speed is greater than 1200 rpm. Based on the decision result, the data is guided to different child nodes, ultimately reaching the leaf node to output the predicted value for that layer. Assume the predicted value output by the first-layer decision tree is 83.5℃.
[0030] It should be noted that the first-layer decision tree was trained on health status data, and its predicted values represent a rough estimate of the target parameters that the equipment should have when it is in a healthy state under the current operating conditions. However, the prediction accuracy of a single decision tree is limited and may contain some errors, so subsequent layers of decision trees are needed for correction.
[0031] After obtaining the first-level predictions, the data continues to be fed into the second-level decision tree. During training, the second-level decision tree specifically learns the prediction residuals from the first level—the difference between the first-level predictions and the actual health status labels. When real-time data is input, the second-level decision tree outputs a correction value, such as 1.2℃, to correct for the prediction bias of the first level. This correction value can be positive or negative, indicating whether the first-level prediction was too high or too low.
[0032] The same process is repeated in subsequent decision tree layers. The third layer of the decision tree learns the residuals between the cumulative predictions and actual values from the first and second layers, and outputs a second correction value, such as 0.5℃. The fourth and fifth layers follow the same pattern, with each layer correcting the cumulative error of all preceding layers. In this embodiment, the baseline generation model can contain 300 layers of decision trees. Through this progressive error correction, the model can achieve high prediction accuracy.
[0033] After obtaining the predicted values from each layer of the decision tree, a weighted summation is required to obtain the final reference baseline. In this weighted summation process, the predicted value from the first layer of the decision tree serves as the base value, and the correction values from subsequent layers are weighted and accumulated according to a preset learning rate. The learning rate is a coefficient between 0 and 1, used to control the contribution of each layer's correction value to the final prediction. In this embodiment, the learning rate can be set to 0.08.
[0034] The specific weighted summation process is as follows: First, the first-layer predicted value of 83.5℃ is taken as the initial baseline. Then, the second-layer correction value of 1.2℃ is multiplied by the learning rate of 0.08 to obtain 0.096℃, which is added to the initial baseline to obtain 83.596℃. Next, the third-layer correction value of 0.5℃ is multiplied by 0.08 to obtain 0.04℃, which is also added to obtain 83.636℃. This process continues until the 300th layer, with the correction values of each layer being weighted according to the learning rate and then accumulated to finally obtain the reference baseline, such as 85.2℃.
[0035] It's important to note that the purpose of introducing a learning rate is to prevent overfitting. If the correction values from each layer are directly summed, the model may overfit the training data, leading to decreased generalization ability on new data. By setting a smaller learning rate, each decision tree layer contributes only a small portion of the correction, requiring more trees to work together to complete the prediction task. This enhances the model's stability and generalization ability. The specific value of the learning rate can be adjusted based on the actual application results, generally between 0.01 and 0.3.
[0036] After obtaining the reference baseline, it needs to be compared with the actual measured value in the monitoring data. The actual measured value refers to the true value of the target monitoring parameter collected at the same time, such as the actual coolant temperature measured by the equipment at the current time being 87.5℃. The difference between the actual measured value of 87.5℃ and the reference baseline of 85.2℃ is calculated, resulting in a difference of 2.3℃.
[0037] The physical meaning of the discrepancy data lies in quantifying the degree of deviation between the current state and the healthy state of the equipment. If the equipment is in a healthy state, its actual measured value should be close to the reference baseline, and the discrepancy data will be small, usually fluctuating within the measurement error range. However, when the equipment degrades, such as when the radiator of the cooling system becomes partially blocked, leading to a decrease in heat dissipation efficiency, the coolant temperature will rise under the same operating conditions, and the actual measured value will be significantly higher than the reference baseline, resulting in a significant increase in the discrepancy data.
[0038] In one embodiment, for time-series data, the aforementioned reference baseline calculation and difference comparison need to be performed for each time point. Assuming data is collected at 3600 time points over one hour, the operating parameters at each time point need to be input into the baseline generation model to obtain 3600 reference baseline values. These are then compared with the actual measured values at the corresponding time points to calculate the differences, resulting in 3600 difference data points, constituting a time series of difference data. This time series of difference data comprehensively records the dynamic process of the equipment deviating from its healthy state throughout the entire monitoring period, providing foundational data for subsequent feature extraction and health assessment.
[0039] In another embodiment, the multi-layered decision tree structure of the baseline generation model enables it to adapt to different operating conditions. When operating parameters change, such as when a vehicle transitions from driving on a flat road to climbing an incline, engine speed and load increase. The baseline generation model then recalculates the reference baseline based on the new combination of operating parameters and outputs the target parameter values that the equipment should have under the new operating conditions. This dynamic adaptability solves the problem that traditional fixed threshold methods cannot cope with changes in operating conditions and avoids misjudging normal operating condition transitions as equipment malfunctions.
[0040] Furthermore, the baseline generation model is trained through the following steps: acquiring historical operating data of the target device in a healthy state, the historical operating data including operating parameters and corresponding true values of the operating parameters; using the operating parameters in the historical operating data as input and the corresponding true values of the operating parameters as labels to train an initial decision tree model to obtain an initial prediction model; inputting the operating parameters into the initial prediction model to obtain initial predicted values, calculating a residual sequence based on the initial predicted values and the true values of the operating parameters to obtain first-round residual data; using the operating parameters as input and the first-round residual data... The first residual decision tree is trained as the fitting target to obtain the first residual prediction model. The initial prediction model and the first residual prediction model are then weighted and combined according to a preset learning rate to obtain the first round ensemble model. The operating condition parameters from the historical operating data are input into the first round ensemble model to obtain the current predicted value. The current residual is calculated based on the current predicted value and the actual value of the operating parameters. It is determined whether the current residual is less than a preset convergence threshold. If not, the current residual is used as the new fitting target to repeat the residual decision tree training and model ensemble steps. If yes, the current ensemble model is used as the baseline generation model.
[0041] Specifically, the training process of the baseline generation model employs an iterative optimization mechanism using gradient boosting decision trees. This training process enables the model to automatically learn the complex mapping relationship between the equipment's operating parameters and runtime parameters under healthy conditions, without the need for manually setting fixed thresholds or empirical formulas.
[0042] First, acquire historical operating data of the target equipment in a healthy state. The healthy state refers to the period during which the equipment is in normal working condition without any degradation or malfunction. The collection period for historical operating data can be set from one week to one month to cover the equipment's operation under different working conditions. For example, for a vehicle engine cooling system, operating data can be collected under various typical working conditions such as urban roads, highways, and hill climbing. The historical operating data includes two parts: first, operating parameters, such as engine speed, fuel supply, and intake pressure, which constitute the model input; second, the actual values of operating parameters, such as the actual measured value of coolant temperature, which constitute the labels for model training, representing the true performance of the equipment's healthy state under the corresponding working conditions.
[0043] When training the initial decision tree model, the CART regression tree from the decision tree algorithm is used as the base learner. The CART tree constructs its tree structure using a recursive binary search method, selecting a working parameter and a split point at each node to minimize the sum of the variances of the data within the two child nodes after the split. For example, the root node might choose engine speed as the split variable, with a split point of 1200 rpm. Samples with speeds less than 1200 rpm are assigned to the left child node, and samples with speeds greater than or equal to 1200 rpm are assigned to the right child node. This process is repeated recursively until a stopping condition is met, such as the tree depth reaching 3 levels or the number of samples in a leaf node decreasing to less than 10.
[0044] The trained initial decision tree model can predict operating parameters, but its prediction accuracy is limited. Suppose there is a set of training samples whose actual coolant temperature corresponds to 85.0℃, while the initial decision tree predicts 83.5℃, resulting in a prediction error of 1.5℃. This error is the residual. The operating parameters of all training samples are input into the initial prediction model, and the difference between the predicted value and the actual value for each sample is calculated, yielding a residual sequence. For example, a training set containing 10,000 samples will produce 10,000 residual values, which constitute the first round of residual data.
[0045] The key technological innovation lies in the subsequent residual fitting process. Unlike traditional methods that attempt to directly predict the target value, gradient boosting decision trees employ a progressive error correction strategy. Specifically, the residual data from the first round is used as a new fitting target to retrain a decision tree, called the first residual decision tree. This tree no longer predicts the coolant temperature itself, but instead learns the prediction error patterns of the initial model. For example, the first residual tree might find that when the engine speed is between 1500-1800 rpm and the ambient temperature is above 30℃, the initial model's predictions are generally underestimated by about 1.2℃. In this operating range, the first residual tree will output a correction value of 1.2℃.
[0046] After obtaining the first residual prediction model, it is weighted and combined with the initial prediction model to obtain the first-round ensemble model. The weights of the weighted combination are controlled by a preset learning rate, which is another important technical feature of the gradient boosting algorithm. Assuming the learning rate is set to 0.08, for a certain sample, the initial model predicts a value of 83.5℃, and the first residual tree predicts a correction value of 1.5℃, then the prediction value of the first-round ensemble model is 83.5 plus 1.5 multiplied by 0.08, equaling 83.62℃. The role of the learning rate is to limit the influence of each new tree on the final prediction. By using a smaller learning rate and a larger number of trees, the model can be made more robust, avoiding overfitting.
[0047] Next, the operating parameters from the historical data are input into the first-round ensemble model to obtain the current predicted value. Taking the above sample as an example, the current predicted value is 83.62℃, the actual value is 85.0℃, and the calculated current residual is 1.38℃. It is then determined whether this current residual is less than a preset convergence threshold, which can be set to 0.05℃, indicating that the model is considered converged when the average prediction error is less than 0.05℃. In this example, 1.38℃ is significantly greater than 0.05℃, therefore the model has not yet converged and needs further training.
[0048] At this point, the current residual is used as the new fitting target to train the second residual decision tree, resulting in the second residual prediction model. This second residual prediction model is then weighted according to the learning rate and combined with the first-round ensemble model to obtain the second-round ensemble model. This process is repeated continuously, with each round training a new decision tree based on the prediction residuals of the previous round, gradually reducing the model's prediction error. For example, after 50 iterations, the model's average prediction error may decrease to 0.5℃; after 150 iterations, it may decrease to 0.1℃; and after 300 iterations, it may decrease to 0.03℃, satisfying the convergence condition.
[0049] It's important to note that this residual asymptotic fitting training mechanism enables the model to capture the complex nonlinear relationships between operating parameters and runtime parameters. Unlike simple models such as linear regression, which can only fit linear relationships, and unlike single deep decision trees, which are prone to overfitting, gradient boosting decision trees, through the combination of multiple shallow trees, can express complex nonlinear mappings and possess excellent generalization ability. Each tree is limited to a depth of 3 layers, called a weak learner. While their prediction accuracy is not high when used alone, the collaborative work of 300 weak learners achieves very high overall prediction accuracy.
[0050] In another embodiment, to further improve model performance, random sampling techniques can be introduced during training. When training a new residual tree in each round, instead of using all training samples, 80% of the samples are randomly selected for training. This increases model diversity and reduces the risk of overfitting. Simultaneously, when splitting at each node of the decision tree, some operating condition parameters can be randomly selected as candidate splitting variables, such as randomly selecting 4 from 6 operating condition parameters. This technique, called feature subsampling, can reduce the correlation between different trees and improve the ensemble effect.
[0051] Once the model training converges, i.e., when the current residual is less than a preset convergence threshold, the current ensemble model is saved as the final baseline generation model. This model includes all trained decision trees and their corresponding weight coefficients. In subsequent applications, only the real-time collected operating condition parameters need to be input into the model to quickly calculate the reference baseline for the corresponding operating condition. The entire prediction process typically takes only milliseconds, meeting the needs of real-time monitoring. This data-driven adaptive baseline generation method, compared to traditional fixed threshold or physical model methods, has stronger adaptability to operating conditions and higher prediction accuracy, and is a key technological foundation for achieving accurate health monitoring.
[0052] 103. Perform feature extraction and distribution analysis on the differential data to obtain a health feature representation reflecting the degree of equipment degradation. In this embodiment, the difference data contains information about the equipment's deviation from its healthy state. However, these raw difference value sequences are still shallow features and are difficult to use directly to accurately assess the degree of equipment degradation. For example, the difference data may experience brief peaks due to instantaneous fluctuations in operating conditions, or there may be large difference values even when the equipment is healthy under certain operating conditions. These situations can interfere with the accuracy of health assessment. Therefore, it is necessary to perform deep feature extraction on the difference data to uncover the dynamic feature patterns that truly reflect equipment degradation.
[0053] The feature extraction and distribution analysis includes: inputting the differential data into the first encoding layer of a pre-trained feature extraction network to obtain a primary feature representation; performing residual operations on the primary feature representation and the differential data to obtain detrended feature data; sequentially passing the detrended feature data through a multi-layer encoder for layer-by-layer feature compression to obtain a deep feature representation; performing kernel density estimation on the deep feature representation to obtain a real-time state probability density function; and calculating the divergence value between the pre-stored health state probability density function and the real-time state probability density function to obtain the health feature representation.
[0054] Specifically, the feature extraction network employs a stacked sparse autoencoder structure, an unsupervised deep learning algorithm capable of automatically learning the intrinsic representation of data. After the differential data is input into the network's first encoding layer, this layer maps the original differential data to a high-dimensional feature space through weighted connections of neurons and non-linear activation functions, obtaining a primary feature representation. For example, an input differential data sequence with 1000 time points can yield a primary feature vector of dimension 400 after processing by the first encoding layer.
[0055] A key technological innovation of this application lies in the introduction of a residual operation layer after the first encoding layer. This residual layer subtracts the primary feature representation from the original difference data, removing linear trend components and preserving dynamic features. This residual processing filters out trend terms in the difference data caused by slow changes in operating conditions, allowing subsequent layers to focus more on extracting dynamic fluctuation features reflecting equipment degradation. For example, when a vehicle gradually accelerates from low speed to high speed, the difference data may show a slow upward trend, but this trend is caused by normal operating condition changes, not equipment degradation. Through residual operation, this trend interference can be eliminated, highlighting the true degradation features.
[0056] The detrended feature data then enters a multi-layer encoder for layer-by-layer compression. The second encoding layer compresses the 400-dimensional features to 200 dimensions, and the third encoding layer further compresses them to 100 dimensions. Each layer learns a more abstract and higher-level feature representation. During this layer-by-layer compression process, the network automatically identifies and retains the feature information most important for health assessment, while filtering out redundancy and noise. The final deep feature representation is a 100-dimensional feature vector that highly condenses the essential information of the device's health status.
[0057] To quantify health status, distribution analysis of deep feature representations is required. This embodiment employs kernel density estimation to calculate the probability density function of deep features. Kernel density estimation is a nonparametric statistical method that does not assume the data follows a specific distribution, but rather estimates the probability distribution based on the sample data itself. Specifically, for deep feature data collected over a period of time, a Gaussian kernel function is used to calculate the probability density contribution of each feature point, and the contributions of all points are summed to obtain the real-time state probability density function.
[0058] Simultaneously, the health state probability density function has been pre-calculated and stored during the model training phase. This function represents the distribution characteristics of deep features when the device is in a healthy state. By calculating the KL divergence between the real-time state probability density function and the health state probability density function, the degree of deviation between the current state and the healthy state can be quantified. KL divergence is an indicator in information theory that measures the difference between two probability distributions; a larger value indicates a more severe deviation. This divergence value represents the health features, directly reflecting the degree of device degradation, and can be used for subsequent health assessment.
[0059] 104. Determine the health status of the target device based on the health characteristics to obtain the health monitoring results of the target device.
[0060] In this embodiment, the step of determining the health status of the target device based on the health feature representation to obtain the health monitoring result of the target device includes: acquiring a preset health status grading standard, which includes multiple divergence threshold intervals and corresponding health level identifiers; comparing the divergence value in the health feature representation with each of the multiple divergence threshold intervals to determine the threshold interval to which the divergence value belongs, and obtaining the corresponding health level identifier; determining whether a preset alarm threshold is exceeded based on the health level identifier, and if so, generating a device abnormality alarm signal; collecting the device identifier, current timestamp, health level identifier, and divergence value of the target device, and combining them to generate a structured health monitoring record; storing the health monitoring record in a health status database, and outputting the health monitoring record as the health monitoring result of the target device.
[0061] Specifically, the health status grading standard is based on statistical analysis of a large amount of equipment degradation data. Unlike traditional methods that use a single threshold for binary classification, this application adopts a multi-level classification system, which can more precisely reflect the gradual degradation process of equipment from health to failure. The health status grading standard can include four levels: a healthy state corresponds to a divergence value of less than 0.001, a slightly degraded state corresponds to a divergence value between 0.001 and 0.01, a moderately degraded state corresponds to a divergence value between 0.01 and 0.1, and a severely degraded state corresponds to a divergence value greater than or equal to 0.1.
[0062] The core technological innovation of this application lies in the adaptive determination mechanism of the threshold range. Traditional methods usually rely on expert experience or simple statistical methods to set fixed thresholds, which are difficult to adapt to the differences in different equipment types and application scenarios. This embodiment proposes an automatic threshold calibration method based on historical degradation data, specifically including: collecting complete degradation process data of the equipment from health to failure, which covers all stages of the equipment life cycle; calculating the corresponding divergence value for each stage to form a mapping relationship between divergence value and degree of degradation; determining the key quantiles of the divergence value distribution through statistical analysis, for example, arranging the divergence values from smallest to largest, taking the 75th quantile as the boundary between health and slight degradation, the 90th quantile as the boundary between slight and moderate degradation, and the 95th quantile as the boundary between moderate and severe degradation.
[0063] The advantage of this adaptive threshold determination method is that different types of equipment can automatically obtain judgment criteria suitable for their own characteristics. For example, for high-precision CNC machine tools, the divergence value in a healthy state may be concentrated below 0.0001, and the degradation threshold may be set to 0.0005; while for large construction machinery, the normal fluctuation is larger, and the divergence value in a healthy state may be around 0.005, and the degradation threshold may be set to 0.02. Through automatic calibration, the subjectivity and limitations of manually setting thresholds are avoided.
[0064] When determining the health level, the real-time calculated divergence value is compared with an adaptively determined threshold range. Assuming a divergence value of 0.015, the range it belongs to is determined sequentially, and the health level is identified as moderate degradation. Further checks are made to see if an alarm threshold is exceeded; if so, an equipment malfunction alarm signal is generated. The alarm signal can be transmitted to maintenance personnel through various means, including the monitoring interface and notification messages.
[0065] Subsequently, information such as device identification, timestamp, health level, and divergence value are collected and combined to generate a health monitoring record. This record is stored in a health status database using time-series database technology, supporting fast retrieval by time index and multi-device management. Finally, the health monitoring record is output, which can be displayed on a monitoring screen, generate trend reports, push to management software, or be queried on mobile devices, achieving deep integration with the enterprise's digital management system and supporting predictive maintenance decisions.
[0066] In this embodiment, monitoring data is obtained by collecting the operating parameters of the target equipment; a reference benchmark is calculated based on the monitoring data using a pre-trained baseline generation model, and the monitoring data is compared with the reference benchmark to obtain difference data; feature extraction and distribution analysis are performed on the difference data to obtain a health feature representation; and the health status is determined based on the health feature representation to obtain the monitoring result. This method dynamically calculates the reference benchmark based on real-time operating conditions using a baseline generation model, replacing the traditional fixed threshold method. It uses a deep feature extraction network to capture equipment degradation features and uses probability distribution divergence values to quantify the health status, solving the problem of poor adaptability of existing methods under dynamic operating conditions and improving the accuracy and reliability of equipment health monitoring.
[0067] Please see Figure 2 The second embodiment of the device health monitoring method in this application includes: 201. Collect the operating parameters of the target equipment during operation to obtain monitoring data corresponding to the operating status of the target equipment; 202. A reference benchmark is calculated based on the monitoring data using a pre-trained baseline generation model, and the monitoring data is compared with the reference benchmark to obtain difference data characterizing the health status of the target device; In this embodiment, steps 201-202 are similar to steps 101-102 in the first embodiment, and will not be described again here.
[0068] 203. Input the difference data into the first encoding layer of the pre-trained feature extraction network to obtain the primary feature representation; In this embodiment, the feature extraction network adopts a stacked sparse autoencoder architecture, which has learned the intrinsic feature distribution of differential data under healthy conditions during the training phase. The first encoding layer is the starting point of the entire feature extraction process, and its structure contains 400 neuron nodes, each of which is connected to the input differential data through a weight matrix.
[0069] When the differential data is input into the first encoding layer, the network first performs a linear transformation on the data, which involves multiplying the input vector by the pre-trained weight matrix and adding a bias vector. This linear transformation maps the original differential data to a new feature space. The transformation result is then processed by a non-linear activation function, such as ReLU or Sigmoid, to introduce non-linear characteristics, enabling the network to learn non-linear relationships within the data.
[0070] After linear transformation and nonlinear activation, a primary feature representation is obtained. This primary feature is a 400-dimensional vector, with each dimension corresponding to the activation value of a neuron. These activation values reflect the response strength of the differential data across different feature dimensions. For example, some neurons may be sensitive to the peak features of the differential data, while others may be sensitive to low-frequency fluctuations; different neurons capture different aspects of the data's characteristics.
[0071] While the initial feature representation has completed the preliminary mapping from the original data space to the feature space, it still contains a significant amount of raw information and redundant components. These features are still influenced to some extent by the linear trend components in the original variance data, such as the trend changes in variance values caused by slow changes in operating conditions. Therefore, in subsequent steps, residual operations are needed to remove these trend interferences and extract purer dynamic degradation features.
[0072] 204. Perform residual operation on the primary feature representation and the difference data to obtain detrended feature data; In this embodiment, the residual operation layer is designed to eliminate the shallow trend components contained in the primary features, so that subsequent coding layers can focus on extracting deep dynamic features that reflect device degradation.
[0073] The specific process of residual operation is as follows: First, the primary feature representation is inversely transformed. Through a trainable weight matrix and bias vector, the 400-dimensional primary features are mapped back to the dimensional space of the original difference data, resulting in the reconstructed input data. This inverse transformation process is essentially the inverse mapping of the features learned by the first encoding layer to the original data. The reconstructed data can be understood as the first encoding layer's preliminary understanding and expression of the input data.
[0074] The original difference data is then subtracted element-wise from the reconstructed input data to calculate the residual vector. This residual vector represents the information components in the original data that were not fully expressed by the first encoding layer. Since the first encoding layer mainly captures the main trends and large-scale features of the data, the residual vector mainly retains the detailed fluctuations and dynamic changes in the data.
[0075] The residual vectors are normalized to ensure the numerical scale is suitable for subsequent encoding layers. Batch normalization is employed, which calculates the mean and variance of the residual vectors across batches of samples, standardizing the data to a distribution with a mean of 0 and a variance of 1. Batch normalization not only unifies the numerical scale but also accelerates network training convergence and improves model stability. Compared to the original variance data, the resulting detrended feature data filters out linear trend interference caused by slow changes in operating conditions, highlighting the dynamic characteristics of rapid fluctuations in equipment status.
[0076] 205. The detrended feature data is sequentially compressed layer by layer through a multi-layer encoder to obtain a deep feature representation; In this embodiment, the multi-layer encoder adopts a stacked structure, including a second coding layer and a third coding layer, forming a feature extraction path of layer-by-layer compression. This stacked structure follows the hierarchical feature learning principle in deep learning, extracting basic features in shallow layers and abstract features in deep layers.
[0077] The detrended feature data is first fed into the second encoding layer. This layer contains 200 neurons, achieving feature compression compared to the previous layer. The second encoding layer connects to the output of the previous layer through a weight matrix, which is learned during the training phase using a backpropagation algorithm. This weight matrix automatically identifies and retains the most important feature combinations for health assessment while filtering out redundant information.
[0078] In the processing of the second encoding layer, a sparsity constraint mechanism is introduced, which is a core technical feature of stacked sparse autoencoders. The sparsity constraint requires that only a small number of neurons in the network are active, while the majority remain inhibited. Specifically, the average activation rate of each neuron on the training samples is calculated. If the average activation rate exceeds a preset sparsity target of 0.1, a penalty term is added to the loss function, forcing the neuron to reduce its activation frequency. This sparsity constraint makes the features learned by the network more compact and efficient.
[0079] The 200-dimensional features output from the second coding layer are further input into the third coding layer for compression. The third coding layer contains 100 neurons and also uses weighted connections and non-linear activation. After compression by the third layer, the feature dimension is reduced from 200 to 100, achieving a higher degree of information condensation.
[0080] The output of the third encoding layer is the deep feature representation, a 100-dimensional feature vector representing the most compact expression of the differential data after multiple layers of nonlinear transformation and feature compression. This deep feature highly condenses the essential information of the device's health status, achieving tens of times the dimensionality compression compared to the original thousands of dimensions of time series data, while retaining the most critical feature information for health assessment. The layer-by-layer compression network structure follows the information bottleneck theory, forcing the network to learn the essential representation of the data by gradually reducing the feature dimension, discarding noise and redundant information, thereby improving the accuracy of health assessment.
[0081] 206. Perform kernel density estimation on the deep feature representation to obtain the real-time state probability density function; In this embodiment, the step of kernel density estimation of the deep feature representation to obtain the real-time state probability density function includes: calculating the optimal bandwidth parameter of the kernel function based on the number of samples and feature dimension of the deep feature representation; constructing a Gaussian kernel function with each sample point in the deep feature representation as the center and the optimal bandwidth parameter as the scale parameter to obtain the kernel function corresponding to each sample point; setting multiple evaluation grid points in the feature space, and for each evaluation grid point, calculating the function value of the kernel function of all sample points at the corresponding evaluation grid point to obtain the density contribution value of each sample point to the corresponding evaluation grid point; summing the density contribution values of each sample point and dividing by the total number of samples and the feature dimension power of the bandwidth value to obtain the probability density estimate value of the corresponding evaluation grid point; repeating the above density estimation calculation for all evaluation grid points in the feature space to obtain the real-time state probability density function.
[0082] Specifically, kernel density estimation is a nonparametric statistical method used to estimate the probability density distribution of data from a finite sample. Unlike traditional parametric estimation methods, kernel density estimation does not assume that the data follows a specific distribution form such as a normal distribution, but rather estimates the distribution entirely based on the sample data itself, thus possessing greater adaptability and flexibility.
[0083] First, the optimal bandwidth parameter of the kernel function is calculated. This parameter controls the width of the kernel function and directly affects the smoothness of the density estimation. Too small a bandwidth will result in an overly coarse estimation curve with excessive local fluctuations; too large a bandwidth will result in an overly smooth estimation curve, losing detailed data features. This embodiment uses the Silverman criterion to automatically determine the optimal bandwidth. This criterion calculates the bandwidth value based on the number of samples and the feature dimension. The calculation process comprehensively considers the sample standard deviation and the interquartile range, taking the smaller of the two, and then multiplying it by the negative fifth power of the number of samples and an adjustment factor of 0.9. For example, for deep feature data containing 1000 samples and a feature dimension of 100, the bandwidth parameter calculated using the Silverman criterion is approximately 0.25.
[0084] After obtaining the optimal bandwidth parameters, a Gaussian kernel function is constructed for each sample point. The Gaussian kernel function is a commonly used kernel function form with good mathematical properties and computational efficiency. Centered on the sample point and using the bandwidth parameter as the scale, the Gaussian kernel function describes the density contribution of that sample point to the surrounding space. The closer a location is to the sample point, the larger the kernel function value, indicating a greater density contribution from that sample point at that location; the farther away the location, the more exponentially the kernel function value decays, and the contribution gradually decreases. For 1000 sample points, 1000 Gaussian kernel functions need to be constructed, each centered on the corresponding sample point.
[0085] To calculate the probability density distribution of the entire feature space, evaluation grid points need to be set in the space. Since deep features are 100-dimensional vectors, directly setting grid points uniformly in the 100-dimensional space would result in excessive computation. This embodiment adopts an adaptive grid strategy, setting a denser grid in areas with dense sample points and a sparser grid in areas with sparse sample points. For example, 10,000 evaluation grid points can be set in the feature space, and the distribution of these grid points is automatically adjusted according to the distribution of sample points.
[0086] For each evaluation grid point, the kernel function value at that point needs to be calculated for all sample points. Specifically, the Euclidean distance between the evaluation grid point and the sample points is first calculated; this distance reflects the proximity of the two points in 100-dimensional space. Then, the distance is substituted into the Gaussian kernel function to calculate the density contribution of that sample point to the evaluation grid point. Since the Gaussian kernel function has a rapid decay characteristic, the contribution value of distant sample points is close to zero and can be ignored, thus reducing computational cost. For each evaluation grid point, 1000 sample points are traversed, and the density contribution values of all sample points are accumulated.
[0087] After obtaining the summation, normalization is required to ensure that the final integral of the probability density function equals 1, satisfying the definition of probability density. The normalization factor includes the total number of samples and the feature dimension power of the bandwidth value. Dividing the summation by the normalization factor yields the probability density estimate for that evaluation grid point. For example, if the original density value obtained from the summation for an evaluation grid point is 50, the total number of samples is 1000, the bandwidth value is 0.25, and the feature dimension is 100, then the normalization factor is 1000 multiplied by 0.25 to the power of 100, ultimately yielding the probability density estimate for that point.
[0088] The density estimation calculation is repeated for all 10,000 evaluation grid points in the feature space to obtain the probability density estimate for each grid point. These discrete grid points and their corresponding density values together constitute the numerical representation of the real-time state probability density function. This function describes the distribution characteristics of deep features in the feature space under the current device state. Different health states will exhibit different distribution patterns, providing a quantitative distribution basis for subsequent health assessment.
[0089] 207. Calculate the divergence value between the pre-stored health state probability density function and the real-time state probability density function to obtain the health feature representation; In this embodiment, the KL divergence value is calculated using KL divergence as the metric. KL divergence, also known as relative entropy, is a classic method in information theory used to measure the degree of difference between two probability distributions. Unlike simple Euclidean distance or mean square error, KL divergence considers the overall morphological differences between probability distributions and can more accurately reflect the degree of similarity between distributions.
[0090] The pre-stored health status probability density function is calculated and saved during the model training phase. This function represents the standard distribution pattern of deep features when the device is in a healthy state. During training, a large amount of operational data under healthy conditions is collected. After extracting deep features, the probability density function of the health status is obtained through kernel density estimation and stored in the monitoring device as a data file. This health baseline function remains unchanged in subsequent applications, serving as a reference standard for judging whether the device deviates from its healthy state.
[0091] When calculating the KL divergence, the real-time state probability density function and the healthy state probability density function need to be compared at the same evaluation grid point. For each evaluation grid point, the probability density values of the two functions at that point are obtained, denoted as P for the real-time state density and Q for the healthy state density. The ratio of P to Q is calculated, and then the natural logarithm of this ratio is taken to obtain the log-probability ratio for that point. Multiplying the log-probability ratio by the real-time state density value P yields the contribution of that grid point to the KL divergence.
[0092] It should be noted that the KL divergence is asymmetric, meaning that the divergence from P to Q is different from the divergence from Q to P. This embodiment uses the divergence calculation direction from the real-time state distribution to the healthy state distribution. This design focuses more on the degree of deviation of the real-time state from the healthy state, which meets the actual needs of health monitoring.
[0093] The divergence contributions of all evaluation grid points are summed to obtain the final KL divergence value. This divergence value is a non-negative number; a larger value indicates a greater difference between the distribution of the real-time state and the healthy state, and a more severe degree of equipment degradation. For example, when the equipment is in a fully healthy state, the real-time state distribution almost overlaps with the healthy state distribution, and the KL divergence is close to 0; when the equipment experiences slight degradation, the distribution of deep features begins to shift, and the KL divergence may increase to 0.005; when the equipment is severely degraded, the distribution deviates significantly, and the KL divergence may reach above 0.1.
[0094] This divergence value, representing the health characteristic, is a single numerical indicator that highly condenses information about the equipment's health status. Compared to raw multidimensional time series data or multidimensional deep feature vectors, this scalar indicator is more intuitive, easier to understand and use, and facilitates setting thresholds and determining health levels. Furthermore, because the KL divergence is based on the overall difference in probability distribution, it is highly robust to single-point noise and random fluctuations, and can stably reflect the true degradation trend of the equipment.
[0095] 208. Determine the health status of the target device based on the health characteristics to obtain the health monitoring results of the target device.
[0096] In this embodiment, step 208 is similar to step 104 in the first embodiment, and will not be described again here.
[0097] In this embodiment, monitoring data is obtained by collecting the operating parameters of the target equipment; a reference benchmark is calculated based on the monitoring data using a pre-trained baseline generation model, and the monitoring data is compared with the reference benchmark to obtain difference data; feature extraction and distribution analysis are performed on the difference data to obtain a health feature representation; and the health status is determined based on the health feature representation to obtain the monitoring result. This method dynamically calculates the reference benchmark based on real-time operating conditions using a baseline generation model, replacing the traditional fixed threshold method. It uses a deep feature extraction network to capture equipment degradation features and uses probability distribution divergence values to quantify the health status, solving the problem of poor adaptability of existing methods under dynamic operating conditions and improving the accuracy and reliability of equipment health monitoring.
[0098] The above describes the device health monitoring method in the embodiments of the present invention. The following describes the device health monitoring device in the embodiments of the present invention. Please refer to [link to device health monitoring device description]. Figure 3 One embodiment of the device health monitoring device in this invention includes: The data acquisition module 301 is used to collect the operating parameters of the target equipment during operation and obtain monitoring data corresponding to the operating status of the target equipment. The baseline comparison module 302 is used to calculate a reference benchmark based on the monitoring data through a pre-trained baseline generation model, and compare the monitoring data with the reference benchmark to obtain difference data characterizing the health status of the target device. Feature analysis module 303 is used to extract features and perform distribution analysis on the differential data to obtain a health feature representation reflecting the degree of equipment degradation. The status determination module 304 is used to determine the health status of the target device based on the health feature representation, and obtain the health monitoring result of the target device.
[0099] In this embodiment of the invention, the equipment health monitoring device operates the aforementioned equipment health monitoring method. The device obtains monitoring data by collecting operating parameters of the target equipment; calculates a reference benchmark based on the monitoring data using a pre-trained baseline generation model; compares the monitoring data with the reference benchmark to obtain difference data; performs feature extraction and distribution analysis on the difference data to obtain a health feature representation; and determines the health status based on the health feature representation to obtain the monitoring result. This method dynamically calculates the reference benchmark based on real-time operating conditions using a baseline generation model, replacing the traditional fixed threshold method. It employs a deep feature extraction network to capture equipment degradation features and uses probability distribution divergence values to quantify the health status, solving the problem of poor adaptability of existing methods under dynamic operating conditions and improving the accuracy and reliability of equipment health monitoring.
[0100] above Figure 3 The device health monitoring device in the embodiments of the present invention will be described in detail from the perspective of unitized functional entities. The device health monitoring device in the embodiments of the present invention will be described in detail from the perspective of hardware processing.
[0101] Figure 4 This is a schematic diagram of the structure of a device health monitoring device 400 provided in an embodiment of the present invention. The device health monitoring device 400 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 410 (e.g., one or more processors) and a memory 420, and one or more storage media 430 (e.g., one or more mass storage devices) for storing application programs 333 or data 432. The memory 420 and storage media 430 can be temporary or persistent storage. The program stored in the storage media 430 may include one or more units (not shown in the diagram), each unit may include a series of instruction operations on the device health monitoring device 400. Furthermore, the processor 410 may be configured to communicate with the storage media 430 and execute the series of instruction operations in the storage media 430 on the device health monitoring device 400 to implement the steps of the aforementioned device health monitoring method.
[0102] The device health monitoring device 400 may also include one or more power supplies 440, one or more wired or wireless network interfaces 450, one or more input / output interfaces 460, and / or one or more operating systems 431, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 4The illustrated device health monitoring device structure does not constitute a limitation on the device health monitoring device provided by the present invention. It may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0103] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of the device health monitoring method.
[0104] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system, device, or unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0105] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0106] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method of device health monitoring, the method comprising: The device health monitoring method comprises: collecting working condition parameters of a target device during operation to obtain monitoring data corresponding to the running state of the target device; calculating a reference benchmark from the monitoring data by a pre-trained baseline generation model, and comparing the monitoring data with the reference benchmark to obtain difference data representing the health state of the target device; performing feature extraction and distribution analysis on the difference data to obtain a health feature representation reflecting the degradation degree of the device operation; determining the health state of the target device according to the health feature representation to obtain the health monitoring result of the target device.
2. The method of claim 1, wherein, The collecting of working condition parameters of a target device during operation to obtain monitoring data corresponding to the running state of the device comprises: determining a target monitoring parameter of the target device, which is a key operating parameter representing the health state of the device; collecting multiple working condition parameters during the operation of the target device, calculating the correlation between each working condition parameter and the target monitoring parameter to obtain the correlation degree value of each working condition parameter; sorting each working condition parameter according to the correlation degree value, selecting working condition parameters with a correlation degree value greater than a preset threshold to obtain an input parameter set; synchronously collecting the input parameter set and the target monitoring parameter within a preset time window to obtain time series data containing time stamps and parameter values, and normalizing the time series data to obtain the monitoring data.
3. The method of claim 1, wherein, The calculating of a reference benchmark from the monitoring data by a pre-trained baseline generation model, and the comparing of the monitoring data with the reference benchmark to obtain difference data representing the health state of the target device comprises: inputting the working condition parameters in the monitoring data into the baseline generation model to obtain prediction values of each layer by cascading multiple decision trees; performing weighted summation on the prediction values of each layer to obtain a reference benchmark corresponding to the current working condition; performing difference calculation on the actual measurement values in the monitoring data and the reference benchmark at the corresponding time points to obtain the difference data.
4. The method of claim 1, wherein, The baseline generation model is trained by the following steps: obtaining historical operation data of the target device in a healthy state, the historical operation data comprising working condition parameters and corresponding operating parameter true values; training an initial decision tree model by taking the working condition parameters in the historical operation data as input and the corresponding operating parameter true values as labels to obtain an initial prediction model; inputting the working condition parameters into the initial prediction model to obtain initial prediction values, and calculating a residual sequence from the initial prediction values and the operating parameter true values to obtain first-round residual data; training a first residual decision tree by taking the working condition parameters as input and the first-round residual data as a fitting target to obtain a first residual prediction model, and combining the initial prediction model and the first residual prediction model by a preset learning rate to obtain a first-round integrated model; Inputting the working condition parameters in the historical running data into the first round of integrated model to obtain a current prediction value, calculating a current residual according to the current prediction value and the real value of the running parameter, and judging whether the current residual is less than a preset convergence threshold value; If not, repeating the residual decision tree training and model integration steps with the current residual as a new fitting target; if yes, taking the current integrated model as the baseline generation model.
5. The method of claim 1, wherein, The feature extraction and distribution analysis on the difference data to obtain the health feature representation reflecting the equipment running degradation degree comprises: Inputting the difference data into a first encoding layer of a pre-trained feature extraction network to obtain a primary feature representation; Performing residual operation on the primary feature representation and the difference data to obtain de-trended feature data; Sequentially passing the de-trended feature data through multiple layers of encoders for layer-by-layer feature compression to obtain a deep feature representation; Performing kernel density estimation on the deep feature representation to obtain a real-time state probability density function; Calculating a divergence value between a pre-stored health state probability density function and the real-time state probability density function to obtain the health feature representation.
6. The method of claim 5, wherein, The kernel density estimation on the deep feature representation to obtain a real-time state probability density function comprises: Calculating an optimal bandwidth parameter of a kernel function according to the sample number and feature dimension of the deep feature representation; Constructing a Gaussian kernel function with each sample point in the deep feature representation as a center and the optimal bandwidth parameter as a scale parameter to obtain a kernel function corresponding to each sample point; Setting multiple evaluation grid points in the feature space, for each evaluation grid point, calculating the function value of the kernel function of all sample points at the corresponding evaluation grid point to obtain the density contribution value of each sample point to the corresponding evaluation grid point; Summing the density contribution values of all sample points and dividing by the total number of samples and the feature dimension of the bandwidth value raised to the power of the feature dimension to obtain the probability density estimation value of the corresponding evaluation grid point; Repeating the above density estimation calculation for all evaluation grid points in the feature space to obtain the real-time state probability density function.
7. The method of claim 1, wherein, The health state determination on the target equipment according to the health feature representation to obtain the health monitoring result of the target equipment comprises: Obtaining a preset health state grading standard, the health state grading standard comprising multiple divergence threshold intervals and corresponding health level identifiers; Comparing the divergence value in the health feature representation with the multiple divergence threshold intervals one by one to determine the threshold interval to which the divergence value belongs and obtain the corresponding health level identifier; Determining whether the health level identifier exceeds a preset alarm threshold value, and if so, generating an equipment abnormality alarm signal; Collecting the equipment identifier, current timestamp, health level identifier and divergence value of the target equipment to generate a structured health monitoring record; Storing the health monitoring record into a health state database and outputting the health monitoring record as the health monitoring result of the target equipment.
8. A device health monitoring apparatus, characterized by, The equipment health monitoring device comprises: The data collection module is configured to collect working condition parameters of the target device during operation to obtain monitoring data corresponding to an operating state of the target device. The baseline comparison module is configured to calculate a reference benchmark from the monitoring data by using a pre-trained baseline generation model, compare the monitoring data with the reference benchmark, and obtain difference data representing a health state of the target device. The feature analysis module is configured to perform feature extraction and distribution analysis on the difference data to obtain a health feature representation reflecting a degradation degree of the device. The state determination module is configured to determine the health state of the target device according to the health feature representation to obtain a health monitoring result of the target device.
9. A device health monitoring device, characterized by The device health monitoring device comprises a memory and at least one processor, and the memory stores instructions. The at least one processor invokes the instructions in the memory to enable the device health monitoring device to perform the steps of the device health monitoring method according to any one of claims 1-7.
10. A computer-readable storage medium having stored thereon instructions, the computer-readable storage medium comprising: The instructions are executed by the processor to implement the steps of the device health monitoring method according to any one of claims 1-7.
Citation Information
Patent Citations
Equipment abnormity identification method, equipment and medium
CN117591949A
Pump station working condition monitoring method and system based on digital twinning and storage medium
CN120195983A
Method and system for monitoring high-pressure gas relay and storage medium
CN120252854A
Cement kiln alternative fuel consumption prediction method and system based on machine learning
CN120508803A
Remote intelligent control distribution box
CN120546270A
Cited By
A device health index prediction method and device, electronic device, and storage medium
CN122359250A
A method, apparatus, electronic device, and storage medium for predicting equipment health index.
CN122359250B