A weightless health assessment method, device and medium based on kernel density
Through kernel density applicability analysis and boundary correction of the unweighted health assessment method, the weight dependence and data imbalance problems in equipment health assessment are solved, and accurate probabilistic assessment of equipment health status is achieved.
Patent Information
- Application Number
- CN202511006333.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-22
AI Technical Summary
Existing equipment health assessment technology relies on weight analysis, which leads to highly subjective assessment results. In addition, the kernel density estimation method cannot handle the problem of data imbalance, which affects the accuracy of equipment health assessment.
A weightless health assessment method based on kernel density is adopted. Through kernel density applicability analysis, boundary correction and multi-indicator comprehensive scoring, a weightless health assessment model is constructed to achieve probabilistic classification of equipment health status.
It improves the accuracy of equipment health assessment, reduces dependence on weight analysis, solves the problem of data imbalance, and achieves accurate probabilistic assessment of equipment health status.
Smart Images

Figure CN120508916B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet of Things technology, and in particular to a kernel density-based weightless health assessment method, device, and medium. Background Art
[0002] In IoT-driven industrial scenarios, the widespread deployment of sensor networks on key equipment, coupled with real-time data collection, transmission, and aggregation via industrial IoT platforms, creates a comprehensive digital mapping of equipment operating status. Automatically identifying abnormal patterns from complex, multidimensional data, quantifying performance degradation, and assessing potential failure risks, while achieving quantitative scoring and prediction of equipment health, has become a key trend in equipment health analysis in digital industrial environments.
[0003] Existing equipment health assessment technologies have significant flaws. For one thing, the weighted summation model used in existing standards relies on manually set weights. For example, the weight of a centrifugal pump's vibration index must be manually adjusted based on the equipment type, resulting in highly subjective assessment results and significant deviations from actual conditions. Furthermore, existing kernel density estimation methods fail to address the issue of data imbalance. When the difference in indicator data volume exceeds 10 times, the statistical characteristics of small-volume indicators are obscured. Furthermore, significant imbalances exist in industrial field data collection. The conflict between high-frequency sampling of key indicators and low-frequency inspections of secondary indicators results in the inability of existing technologies to accurately capture equipment anomalies, impacting the accuracy of equipment health assessments. Summary of the Invention
[0004] The embodiments of the present application provide a kernel density-based weightless health assessment method, device, and medium, which solve the technical problems that existing equipment health assessment methods rely on weight analysis and kernel density estimation methods cannot handle data imbalance.
[0005] In the first aspect, an embodiment of the present application provides a kernel density-based unweighted health assessment method, characterized in that the method includes: obtaining equipment health score index data, and performing kernel density applicability analysis on the equipment health score index data to determine the index comprehensive score adaptation sample; performing boundary correction on the standard kernel density estimation function to obtain a kernel density estimation model; based on the index comprehensive score adaptation sample, determining the model bandwidth of the kernel density estimation model through differentiated applicability comprehensive score analysis; according to the kernel density estimation model and model bandwidth, obtaining the unweighted health assessment kernel density through multi-key indicator kernel density analysis; performing interval probability value calculation of the health level on the health assessment kernel density to determine the probability of the equipment health status.
[0006] In one implementation of the present application, a kernel density applicability analysis is performed on the equipment health score index data to determine the index comprehensive score adaptation sample, specifically including: performing an index data volume analysis on the equipment health score index data to determine the data verification content; wherein the data verification content includes: single index data volume constraint, multi-index data volume balance requirement; based on the data verification content, a kernel density applicability comprehensive score formula is defined, and the comprehensive score is calculated through the kernel density applicability comprehensive score formula; an applicability threshold judgment is performed on the comprehensive score to determine the data volume optimization requirement of the index data volume; according to the data volume optimization requirement, the equipment health score index data is supplemented with data volume to determine the index comprehensive score adaptation sample.
[0007] In one implementation of this application, the kernel density applicability comprehensive scoring formula is:
[0008]
[0009]
[0010] in, The amount of data for key indicators, is the data volume range ratio, is the single indicator data volume score, is the data volume balance score, is the number of indicators, Provides a comprehensive score for kernel density suitability.
[0011] In one implementation of the present application, the standard kernel density estimation function is subjected to boundary correction to obtain a kernel density estimation model, specifically including: setting the kernel of the kernel function to determine the standard kernel density estimation function, and performing kernel function density calculation of sample points of the standard kernel density estimation function to obtain a kernel function density group; based on the kernel function density group, calculating the kernel function density of the reflection term to determine the boundary correction density estimation value; according to the boundary correction density estimation value, filling the kernel density within the boundary to obtain the kernel density estimation model.
[0012] In one implementation of the present application, based on the indicator comprehensive score adaptation sample, the model bandwidth of the kernel density estimation model is determined through differentiated applicability comprehensive score analysis, specifically including: performing applicability division on the applicability comprehensive score of the indicator comprehensive score adaptation sample, and determining the differentiated bandwidth analysis threshold; based on the differentiated bandwidth analysis threshold, the model bandwidth of the kernel density estimation model is determined through indicator differentiated bandwidth revision.
[0013] In one implementation of the present application, the differentiated bandwidth analysis threshold includes: a first differentiation threshold, a second differentiation threshold, and a third differentiation threshold; based on the differentiated bandwidth analysis threshold, the model bandwidth of the kernel density estimation model is determined through the bandwidth revision of indicator differentiation, specifically including: when the comprehensive applicability score is greater than or equal to the first differentiation threshold, the standard bandwidth formula is subjected to a first differentiation correction to suppress outlier interference; when the comprehensive applicability score is less than the first differentiation threshold and greater than or equal to the third differentiation threshold, the standard bandwidth formula is subjected to a second differentiation correction to retain data details of small sample data and reduce noise of large sample data; the model bandwidth of the kernel density estimation model is determined according to the bandwidth obtained by the first differentiation correction or the second differentiation correction.
[0014] In one implementation of the present application, based on the kernel density estimation model and model bandwidth, a multi-key indicator kernel density analysis is performed to obtain an unweighted health assessment kernel density, specifically including: based on the kernel density estimation model and model bandwidth, a multi-indicator kernel density estimation model is determined through multi-indicator fluctuation revision; wherein, the single indicator related parameters of the multi-indicator kernel density estimation model include: sample size, sample interval boundary, sample bandwidth; the multi-indicator sample is input into the multi-indicator kernel density estimation model, and the unweighted health assessment kernel density is obtained through multi-indicator boundary correction.
[0015] In one implementation of the present application, the interval probability value of the health level is calculated for the health assessment kernel density to determine the probability of the device health status, specifically including: performing interval probability integration on the health assessment kernel density to obtain the interval probability distribution; performing device status matching on the interval probability distribution to determine the probability of the device health status.
[0016] In the second aspect, an embodiment of the present application also provides a kernel density-based unweighted health assessment device, characterized in that the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor so that the at least one processor can: obtain device health score index data, and perform kernel density applicability analysis on the device health score index data to determine the index comprehensive score adaptation sample; perform boundary correction on the standard kernel density estimation function to obtain a kernel density estimation model; based on the index comprehensive score adaptation sample, determine the model bandwidth of the kernel density estimation model through differentiated applicability comprehensive score analysis; based on the kernel density estimation model and model bandwidth, obtain the unweighted health assessment kernel density through multi-key indicator kernel density analysis; calculate the interval probability value of the health level of the health assessment kernel density to determine the probability of the device health state.
[0017] In a third aspect, an embodiment of the present application also provides a non-volatile computer storage medium for unweighted health assessment based on kernel density, which stores computer executable instructions, characterized in that the computer executable instructions are set to: obtain device health score index data, and perform kernel density applicability analysis on the device health score index data to determine the index comprehensive score adaptation sample; perform boundary correction on the standard kernel density estimation function to obtain a kernel density estimation model; based on the index comprehensive score adaptation sample, determine the model bandwidth of the kernel density estimation model through differentiated applicability comprehensive score analysis; according to the kernel density estimation model and model bandwidth, obtain the unweighted health assessment kernel density through multi-key indicator kernel density analysis; perform interval probability value calculation of the health level on the health assessment kernel density to determine the probability of the device health status.
[0018] The embodiments of the present application provide a weightless health assessment method, device and medium based on kernel density. By constructing a weightless health assessment model, correcting the boundaries of the kernel density function and comprehensively scoring the practicality of multiple indicators, the technical problems of existing equipment health assessment methods relying on weight analysis and kernel density estimation methods being unable to handle uneven data volume are solved. The weightless health assessment with kernel density function boundary correction and the construction of a probabilistic health status grading system are realized, the accuracy of equipment health assessment is improved, and the dependence of equipment health on weight analysis is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0020] Figure 1 A flowchart of a weightless health assessment method based on kernel density provided in an embodiment of the present application;
[0021] Figure 2 A schematic diagram showing the effect of bandwidth selection on KDE provided in an embodiment of the present application;
[0022] Figure 3 A schematic diagram showing the effects of different kernel functions on KDE provided in an embodiment of the present application;
[0023] Figure 4 A schematic diagram of boundary-corrected density estimation of an Epanechnikov kernel provided in an embodiment of the present application;
[0024] Figure 5 A schematic diagram of the internal structure of a kernel density-based unweighted health assessment device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0025] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0026] The embodiments of the present application provide a weightless health assessment method, device and medium based on kernel density. By constructing a weightless health assessment model, correcting the boundaries of the kernel density function and comprehensively scoring the practicality of multiple indicators, the technical problems of existing equipment health assessment methods relying on weight analysis and kernel density estimation methods being unable to handle uneven data volume are solved. The weightless health assessment with kernel density function boundary correction and the construction of a probabilistic health status grading system are realized, the accuracy of equipment health assessment is improved, and the dependence of equipment health on weight analysis is reduced.
[0027] The technical solutions proposed in the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0028] Figure 1 This is a flow chart of a weightless health assessment method based on kernel density provided in an embodiment of the present application. Figure 1 As shown, the embodiment of the present application provides a weightless health assessment method based on kernel density, which specifically includes the following steps:
[0029] Step 101: Obtain equipment health score index data, and perform kernel density applicability analysis on the equipment health score index data to determine index comprehensive score adaptation samples.
[0030] For example, since it is necessary to establish key indicators for unweighted analysis, it is necessary to determine whether the indicator data is suitable for kernel density analysis. By performing kernel density applicability analysis on the equipment health score indicator data, the indicator comprehensive score adaptation sample is determined, and different types of data verification are performed on the indicator data of the comprehensive health score. The data adaptation processing of the equipment health score indicator data and the kernel density function is realized, which improves the applicability of the indicator comprehensive score adaptation sample.
[0031] Specifically, a kernel density applicability analysis is performed on the equipment health score indicator data to determine the indicator comprehensive score adaptation sample, including: performing indicator data volume analysis on the equipment health score indicator data to determine the data verification content; wherein, the data verification content includes: single indicator data volume constraint, multi-indicator data volume balance requirement; based on the data verification content, a kernel density applicability comprehensive score formula is defined, and the comprehensive score is calculated through the kernel density applicability comprehensive score formula; a suitability threshold judgment is performed on the comprehensive score to determine the data volume optimization requirement of the indicator data volume; according to the data volume optimization requirement, the equipment health score indicator data is supplemented with data volume to determine the indicator comprehensive score adaptation sample.
[0032] Furthermore, the comprehensive scoring formula for kernel density applicability is:
[0033]
[0034]
[0035] in, The amount of data for key indicators, is the data volume range ratio, is the single indicator data volume score, is the data volume balance score, is the number of indicators, Provides a comprehensive score for kernel density suitability.
[0036] In one embodiment, the comprehensive health scoring index system for a centrifugal pump unit in a waterworks plant includes four primary indicators: vibration status, temperature status, wear and leakage, and motor status. Eight secondary indicators include vibration acceleration / velocity RMS, bearing seat vibration spectrum anomalies, bearing temperature, motor winding temperature, shaft seal leakage, impeller / pump casing wear, three-phase current imbalance, and insulation resistance.
[0037] The score of the first-level indicator is obtained by weighting the second-level indicator, and the comprehensive health score is obtained by weighting the first-level indicator. Therefore, we can think that the comprehensive health score can be regarded as a weighted result obtained directly by weighting the key indicators such as the second-level indicator that cannot be further divided into sub-indicators.
[0038] It should be noted that this application can align the numerical calculation of the health score through technical means such as the threshold segmentation method. This application does not limit the specific calculation method used, and this application does not focus on the frequency of raw data collection when the key indicators are used to score the health, but only focuses on the frequency of the key indicators to derive the health score. For example, using a machine learning scoring method, the raw data collection frequency is 10000HZ, but after processing by a certain algorithm, a health score value is obtained every 1s, so the collection frequency of the health score value is 1HZ.
[0039] The verification content mainly includes the amount of data and the difference in the output volume of each data item. It is necessary to consider the single-indicator data volume constraint and the multi-indicator data volume balance requirement.
[0040] For single-indicator data volume constraints, indicators with different data volumes have different effects on the effect and computational complexity of kernel density estimation (KDE).
[0041] When the amount of data is small, the estimated density curve fluctuates greatly, is easily affected by outliers, and may appear over-smoothed. When the amount of data is less than 20, KDE may misclassify sparse points as multimodal distributions.
[0042] Generally speaking, it is recommended that the amount of data for a single indicator be greater than 30, at which point KDE can better capture the distribution of the data (the central limit theorem holds asymptotically); ideally, the amount of data should be greater than 100.
[0043] When the amount of data is large, the density estimation result is closer to the true distribution, but the computational complexity increases exponentially with the amount of data. For example, when the amount of data is greater than 1000, the bandwidth optimization time will increase significantly.
[0044] Therefore, it can be recommended that the data volume of a single indicator must be greater than 30 and no more than 1000.
[0045] Regarding the requirement for balanced data volumes across multiple indicators, when there are multiple indicators and the difference between their maximum and minimum data volumes exceeds 10 times, the information for some indicators may be "overwhelmed" by the indicators with larger data volumes. Therefore, when using this method, it is important to pay attention to the balance of data volumes across all indicators.
[0046] Table 1 Example of comprehensive health rating indicators for centrifugal pump units in water plants
[0047]
[0048] Table 1 shows an example of comprehensive health score indicators for a centrifugal pump group in a water plant. The indicator with the largest amount of data is vibration acceleration / velocity effective value, with a data volume of 80.
[0049] The indicator with the smallest data volume is insulation resistance, with 6 data volumes. The gap between the largest and smallest data volumes is 13.33 times, which is greater than the constraint of 10.
[0050] In order to comprehensively evaluate whether the data is suitable for the KDE method, this application designs a comprehensive scoring formula for kernel density applicability, which is explained by the following formula.
[0051] (1)
[0052] (2)
[0053] in, The amount of data for key indicators, is the data volume range ratio, is the single indicator data volume score, is the data volume balance score, is the number of indicators, Provides a comprehensive score for kernel density suitability.
[0054] The criteria for judging the applicability threshold are as follows:
[0055] S ≥ 0.7: The data volume of a single indicator meets the requirements, and the data volume of each key indicator is relatively close. It is suitable to use the model proposed in this patent, and the bandwidth is set to the theoretical optimal bandwidth;
[0056] 0.7>S≥0.4: The data volume of a single indicator meets the requirements, but the data volume of each key indicator varies greatly. This is more suitable for the model proposed in this patent, but the bandwidth setting of each indicator should be adaptively adjusted according to the data volume;
[0057] 0.4 > S: The data volume for a single indicator does not meet the requirements, and there is a significant discrepancy between the data volumes of various key indicators. The data volume for the indicator needs to be adjusted. In this case, the data volume for the selected key indicators needs to be adjusted by reselecting the analysis time range, deleting certain indicators with small data volumes, and so on. This will narrow the data volume discrepancies between the key indicators and meet the requirements for using the method described in this patent.
[0058] The data volume of the eight key indicators is as follows: =[6, 15, 20, 25, 30, 65, 70, 80].
[0059] Single indicator data volume score The total score is [0.06, 0.15, 0.2, 0.25, 0.3, 0.65, 0.7, 0.8], and the average score is (0.06+0.15+0.2+0.25+0.3+0.65+0.7+0.8) / 8= 0.38875.
[0060] Data balance score: , , , .
[0061] Overall rating: , it is not suitable to directly use the method proposed in this patent. Therefore, according to the data volume limitation of a single indicator, it is necessary to adjust the indicators with smaller data volumes first and increase the data volume to more than 30.
[0062] After the adjustment, the data volume of the "Insulation Resistance" indicator increased by 24, the data volume of the "Three-phase Current Unbalance" indicator increased by 15, the data volume of the "Impeller Wear" indicator increased by 10, and the data volume of the "Shaft Seal Leakage" indicator increased by 5.
[0063] After the two indicators of "Shaft Seal Leakage" and "Three-Phase Current Unbalance" were updated, the value range of the health score changed. Repeat the calculation process of the KDE applicability comprehensive score formula, and the updated comprehensive score , and therefore meets the applicability criterion.
[0064] Step 102: Perform boundary correction on the standard kernel density estimation function to obtain a kernel density estimation model.
[0065] For example, the traditional weight-based comprehensive health scoring method usually carefully selects several key indicators from the operating status of the equipment, such as vibration amplitude, temperature deviation, energy consumption level, and degree of component wear, to construct a comprehensive health scoring index system. Subsequently, a corresponding weight is set for each indicator, and the size of the weight can intuitively reflect the degree of influence of the indicator on the health of the equipment. Finally, the score value of each indicator is multiplied by the corresponding weight and added up. The resulting weighted sum is the final comprehensive health score of the equipment. This score can clearly and intuitively show the pros and cons of the overall health status of the equipment.
[0066] The core of the unweighted analysis method of this application is to determine the probability of the score of the comprehensive health score within the maximum and minimum value range based on the health scores of each indicator collected. The essence of kernel density estimation is to convert discrete data into a continuous probability density curve by superimposing the kernel function of each data point. However, in KDE, when the data point is close to the boundary (such as near the minimum or maximum value of the sample), part of the kernel function will exceed the data range, resulting in estimation deviation. This phenomenon is called boundary effect. In order to solve this problem, the standard kernel function needs to be boundary corrected. The principle of boundary-corrected kernel density estimation is to compensate for the density deviation at the boundary by introducing a reflection term. Its core is to use mirror data to fill the area outside the boundary so that the weight distribution of the kernel function at the boundary is more reasonable.
[0067] Specifically, the standard kernel density estimation function is subjected to boundary correction to obtain a kernel density estimation model, including: setting the kernel of the kernel function to determine the standard kernel density estimation function, and performing kernel function density calculation of sample points on the standard kernel density estimation function to obtain a kernel function density group; based on the kernel function density group, calculating the kernel function density of the reflection term to determine the boundary correction density estimation value; according to the boundary correction density estimation value, filling the kernel density within the boundary to obtain the kernel density estimation model.
[0068] Figure 2 A schematic diagram illustrating the impact of bandwidth selection on KDE provided in an embodiment of the present application.
[0069] Figure 3 A schematic diagram of the impact of different kernel functions on KDE provided in an embodiment of the present application.
[0070] In one embodiment, for a sample , the kernel density estimation is explained by the following formula.
[0071] (3)
[0072] in, As a kernel function, it needs to meet the conditions of non-negativity, normalization and symmetry; is the bandwidth, which controls the degree of smoothing; is the sample size. The kernel function determines the size of each sample data point. Position The contribution of density.
[0073] In kernel density estimation (KDE), the kernel function and bandwidth are two key parameters. For bandwidth selection, we can use the theoretically optimal bandwidth (Silverman's rule), which is applicable when the data approximately follows a normal distribution.
[0074] For bandwidth selection, you can use the theoretical optimal bandwidth (Silverman's law). This law applies to scenarios where data approximately follows a normal distribution and is explained by the following formula.
[0075] (4)
[0076] in, is the sample standard deviation.
[0077] This application selects the Epanechnikov kernel as the kernel function. The Epanechnikov kernel is a non-parametric kernel function commonly used in kernel density estimation. It is a second-order differentiable symmetric kernel and is explained by the following formula.
[0078] (5)
[0079] in, is the standardized variable, and the formula For bandwidth.
[0080] It should be noted that the reasons for choosing the Epanechnikov kernel as the kernel function are as follows:
[0081] The compact support is non-zero only in the interval [-1, 1], and the weight outside this range is 0, so its computational efficiency is higher than that of the infinite support kernel (such as the Gaussian kernel).
[0082] Second-order optimality: Among all compactly supported kernels, the Epanechnikov kernel can minimize the asymptotic mean square error (AMSE) of the kernel density estimation and is theoretically the optimal kernel function.
[0083] It is suitable for small sample data. Its tight support characteristics can effectively reduce the interference of outliers and is particularly suitable for scenarios with sample size n<50.
[0084] It is suitable for situations where computing resources are limited. Since it only calculates the weights of neighboring points, it is more suitable for fast estimation of large-scale data than the Gaussian kernel.
[0085] It is suitable for non-normal distribution data, especially for data that deviates from normality, such as uniform distribution and symmetrical multimodal distribution.
[0086] Therefore, kernel function and bandwidth are key parameters, and different selections of their parameters will have different effects on KDE.
[0087] Figure 4 A schematic diagram of boundary-corrected density estimation of an Epanechnikov kernel provided in an embodiment of the present application.
[0088] In one embodiment, the reflection term is mathematically expressed as follows:
[0089] Near the left border ( ):Virtual point= (x about symmetry), .
[0090] Near the right border ( ):Virtual point= (x is symmetric about U), .
[0091] Suppose the sample data is [1, 2, 3, 4, 5], L=0.5, U=5.5, and calculate the corrected density estimate of the left boundary at x=0.5 (the right boundary is not considered for the time being).
[0092] First, calculate the bandwidth using Silverman's rule of thumb: ; Calculate each sample point The kernel function density at x=0.5 only considers Points:
[0093] , , , the kernel function density is Based on this reasoning, , the kernel function density is 0.
[0094] Similarly, , the kernel function density is 0. Therefore, the kernel function density of the original data is: 0.099
[0095] Then, the density of the kernel function of the reflection term is calculated to generate virtual data points: the only point that satisfies The point is , the virtual point is 1, calculate the distance from x=0.5 to the virtual point: , the kernel function density is .
[0096] Finally, the boundary-corrected density estimate at x=0.5 is: = 0.099 + 0.099 = 0.198. By comparison, the density estimate of the standard Epanechnikov kernel density estimator at x=0.5 is 0.099. After boundary correction, the estimated value increases to 0.198 due to compensation for the reflection term.
[0097] Similarly, the right boundary correction should also be added in the actual calculation. The specific process is the same as the left boundary correction mentioned above and will not be repeated here.
[0098] After filling the left and right boundaries, for the data sample of a single indicator, in the sample interval The kernel function density within is explained by the following formula.
[0099] (6)
[0100] in:
[0101] (7)
[0102] (8)
[0103] (9)
[0104] in, For standard KDE, the last two and is a reflection term, ensuring that Density outside the interval is reflected back into the interval, To correct bandwidth.
[0105] Step 103: Based on the comprehensive score of the indicators, the sample is adapted and the model bandwidth of the kernel density estimation model is determined through differentiated comprehensive applicability score analysis.
[0106] Illustratively, this application uses differentiated applicability comprehensive scoring analysis to analyze the calculation strategy under different indicator comprehensive scores for one of the key parameters of the kernel density estimation model (i.e., model bandwidth), thereby realizing the construction of a probabilistic health status grading system.
[0107] Specifically, based on the indicator comprehensive score adaptation sample, the model bandwidth of the kernel density estimation model is determined through differentiated applicability comprehensive score analysis, including: performing applicability division on the applicability comprehensive score of the indicator comprehensive score adaptation sample, and determining the differentiated bandwidth analysis threshold; based on the differentiated bandwidth analysis threshold, the model bandwidth of the kernel density estimation model is determined through indicator differentiated bandwidth revision.
[0108] Furthermore, the differentiated bandwidth analysis thresholds include: a first differentiation threshold, a second differentiation threshold, and a third differentiation threshold; based on the differentiated bandwidth analysis thresholds, the model bandwidth of the kernel density estimation model is determined through the bandwidth revision of indicator differentiation, specifically including: when the comprehensive applicability score is greater than or equal to the first differentiation threshold, the standard bandwidth formula is subjected to a first differentiation correction to suppress outlier interference; when the comprehensive applicability score is less than the first differentiation threshold and greater than or equal to the third differentiation threshold, the standard bandwidth formula is subjected to a second differentiation correction to retain data details of small sample data and reduce noise of large sample data; the model bandwidth of the kernel density estimation model is determined according to the bandwidth obtained by the first differentiation correction or the second differentiation correction.
[0109] In one embodiment, differentiated bandwidth calculation strategies are adopted according to different intervals of the comprehensive suitability score S.
[0110] when When , it indicates that the data volume of each indicator is highly balanced and the data volume is sufficient. At this time, the bandwidth calculation needs to focus on suppressing the interference of outliers. The following formula is used for revision:
[0111] (10)
[0112] in, For the The interquartile range (75th percentile) of the indicator Subtract the 25th percentile ), For the The standard deviation of an indicator.
[0113] After data adjustment, the sample size of the eight key indicators is 80. After calculation, The detailed parameters are shown in Table 2. It should be noted that the sample data are for reference only and should be replaced with real health scores in actual applications.
[0114] Table 2 For Bandwidth calculation parameters for each indicator under the circumstances
[0115]
[0116] when When the detailed parameters are shown in Table 2, there are moderate differences in the amount of data for multiple indicators (e.g., a data volume gap of 3-10 times). The traditional unified bandwidth method will cause the bandwidth of small-volume indicators to be too large, masking local features (e.g., failure in outlier identification), and the bandwidth of large-volume indicators to be too small, introducing false fluctuations (overfitting). Therefore, we use a smaller bandwidth for small sample size data (to preserve details) and a larger bandwidth for large sample size data (to reduce noise). Therefore, we use the standard bandwidth formula After revision, this application proposes an adaptive bandwidth formula, which is explained by the following formula.
[0117] (11)
[0118] in, For the The sample size of each indicator; is the average sample size of all indicators, ; is the weight index, balancing the difference in data volume, , usually taken as 0.5.
[0119] Total sample size n=80+70+65+30+30+30+30+30=365, average sample size =365 / 8= 45.625.
[0120] Table 3 For Bandwidth calculation parameters for each indicator under the circumstances
[0121]
[0122] Step 104: Based on the kernel density estimation model and the model bandwidth, a multi-key indicator kernel density analysis is performed to obtain an unweighted health assessment kernel density.
[0123] For example, based on the boundary-corrected kernel density estimation model for a single indicator, an algorithm model for multiple key indicators is designed, and the kernel density analysis of multiple key indicators is performed to obtain the unweighted health assessment kernel density. This realizes the unweighted kernel density calculation under multiple key indicators and the unweighted health assessment with boundary correction of the kernel density function, providing a data basis for determining the probability of the equipment's health status.
[0124] Specifically, based on the kernel density estimation model and model bandwidth, the unweighted health assessment kernel density is obtained through multi-key indicator kernel density analysis, including: based on the kernel density estimation model and model bandwidth, through multi-indicator fluctuation revision, determining the multi-indicator kernel density estimation model; wherein, the single indicator related parameters of the multi-indicator kernel density estimation model include: sample size, sample interval boundary, sample bandwidth; multi-indicator samples are input into the multi-indicator kernel density estimation model, and the unweighted health assessment kernel density is obtained through multi-indicator boundary correction.
[0125] In one embodiment, since all The relevant parameters of the indicators, indicators, , the parameters that have been determined include the following:
[0126] The sample size is , the sample value is ;
[0127] Left boundary of the sample interval , and the right boundary of the sample interval ;
[0128] Sample bandwidth .
[0129] Then, we take the kernel function value of the sample point at x=80 (Table 3) as an example, and first perform The calculation process is explained using vibration acceleration indicators.
[0130] Left boundary of the sample interval , and the right boundary of the sample interval , sample bandwidth , take one of the sample points .
[0131] because , its absolute value , then the kernel function .
[0132] By analogy, the cumulative kernel function values of all 80 sample points of this indicator divided by n=365 (total sample size) are:
[0133]
[0134] Therefore, we simplify the assumptions and calculate the densities of the other seven indicators as follows: , , , , , , .
[0135] The sum of the original densities .
[0136] Similarly, continue to reflect the left boundary calculate
[0137] Vibration acceleration effective value For example, the reflection point is , calculate the sample points Kernel function value at reflection point 64.6 .
[0138] Kernel function value The reflection item density of this indicator is:
[0139]
[0140] Therefore, the left boundary reflection item densities of the other seven indicators are calculated based on simplified assumptions as follows:
[0141] , , , , , , ; The sum of the original densities .
[0142] Continue to reflect the left boundary Calculate the effective value of vibration acceleration For example, the reflection point is , Out of sample range.
[0143] Similarly, if other indicators are also out of range, the kernel function value is 0. It is far from the right boundary, and most of the index reflection items are 0, so
[0144] Finally, calculate the boundary-corrected density estimate
[0145]
[0146] The density estimates after boundary correction for other x values within the value range are calculated using the above steps.
[0147] Step 105: Calculate the interval probability value of the health level based on the health assessment kernel density to determine the health status probability of the device.
[0148] Specifically, the interval probability value of the health level is calculated for the health assessment kernel density to determine the probability of the device health state, including: performing interval probability integration on the health assessment kernel density to obtain an interval probability distribution; and performing device state matching on the interval probability distribution to determine the probability of the device health state.
[0149] In one embodiment, for the sample interval In, suppose there are values a, b, , the probability of being greater than a value a is:
[0150]
[0151] In the range The probability of being within is:
[0152]
[0153] By solving the integral function of the above KDE, we can find the probability value in the corresponding interval. For example, we simplify the assumption that for the warning state in the example of step 5, 8 , with 1 point as one integration step. Therefore, the probability calculation formula for being in the warning state is:
[0154]
[0155] Therefore, it can be calculated , ,..., , and finally .
[0156] Based on the established health model and the collected data for key indicators, we can conclude that there is a 32.1% probability that the centrifugal pump unit will be in a warning state between 8:00 AM and 9:00 AM on Day B of Month A. Similarly, the probabilities of other health states, such as bearing health, can be calculated.
[0157] The above is an embodiment of the method proposed in this application. Based on the same inventive concept, this application embodiment also provides a weightless health assessment device based on kernel density, the structure of which is as follows: Figure 5 shown.
[0158] Figure 5 This is a schematic diagram of the internal structure of a weightless health assessment device based on kernel density provided in an embodiment of the present application. Figure 5 As shown, the equipment includes:
[0159] at least one processor 501;
[0160] and, a memory 502 in communication with the at least one processor;
[0161] The memory 502 stores instructions that can be executed by at least one processor, and the instructions are executed by the at least one processor 501 to enable the at least one processor 501 to:
[0162] Obtain equipment health score indicator data and perform kernel density applicability analysis on the equipment health score indicator data to determine the indicator comprehensive score adaptation sample; perform boundary correction on the standard kernel density estimation function to obtain the kernel density estimation model; based on the indicator comprehensive score adaptation sample, determine the model bandwidth of the kernel density estimation model through differentiated applicability comprehensive score analysis; based on the kernel density estimation model and model bandwidth, obtain the unweighted health assessment kernel density through multi-key indicator kernel density analysis; calculate the interval probability value of the health level of the health assessment kernel density to determine the probability of the equipment health status.
[0163] Some embodiments of the present application provide corresponding Figure 1 A non-volatile computer storage medium for weightless health assessment based on kernel density, storing computer executable instructions, wherein the computer executable instructions are set to:
[0164] Obtain equipment health score indicator data and perform kernel density applicability analysis on the equipment health score indicator data to determine the indicator comprehensive score adaptation sample; perform boundary correction on the standard kernel density estimation function to obtain the kernel density estimation model; based on the indicator comprehensive score adaptation sample, determine the model bandwidth of the kernel density estimation model through differentiated applicability comprehensive score analysis; based on the kernel density estimation model and model bandwidth, obtain the unweighted health assessment kernel density through multi-key indicator kernel density analysis; calculate the interval probability value of the health level of the health assessment kernel density to determine the probability of the equipment health status.
Claims
1. A weightless health assessment method based on kernel density, characterized in that: The method comprises: Obtaining equipment health score indicator data, and performing kernel density applicability analysis on the equipment health score indicator data to determine an indicator comprehensive score adaptation sample; Perform boundary correction on the standard kernel density estimation function to obtain the kernel density estimation model; Based on the comprehensive score of the indicators, the model bandwidth of the kernel density estimation model is determined through differentiated comprehensive applicability score analysis of the adapted samples; According to the kernel density estimation model and the model bandwidth, a weighted health assessment kernel density is obtained through a multi-key indicator kernel density analysis; Calculating the interval probability value of the health level of the health assessment kernel density to determine the health status probability of the equipment; Perform kernel density applicability analysis on the equipment health score indicator data to determine the indicator comprehensive score adaptation sample, specifically including: Performing an indicator data volume analysis on the device health score indicator data to determine data verification content; wherein the data verification content includes: single indicator data volume constraints and multi-indicator data volume balance requirements; Based on the data verification content, a kernel density applicability comprehensive scoring formula is defined, and a comprehensive score is calculated using the kernel density applicability comprehensive scoring formula; Performing a suitability threshold judgment on the comprehensive score to determine the data volume optimization requirements of the indicator data volume; According to the data volume optimization requirements, the device health score index data is supplemented with data volume, and a sample suitable for the comprehensive score of the index is determined; The kernel density applicability comprehensive scoring formula is: in, The amount of data for key indicators, is the data volume range ratio, is the single indicator data volume score, is the data volume balance score, is the number of indicators, Provides a comprehensive score for kernel density suitability.
2. The unweighted health assessment method based on kernel density according to claim 1, characterized in that: The standard kernel density estimation function is subjected to boundary correction to obtain the kernel density estimation model, which includes: Setting the kernel of the kernel function to determine the standard kernel density estimation function, and performing kernel function density calculation of sample points on the standard kernel density estimation function to obtain a kernel function density group; Based on the kernel function density group, calculating the reflection term kernel function density and determining the boundary correction density estimate; The kernel density estimation model is obtained by correcting the density estimation value of the boundary and filling the kernel density within the boundary.
3. The unweighted health assessment method based on kernel density according to claim 1, characterized in that: Based on the comprehensive score of the indicators, the adaptive samples are analyzed through differentiated applicability comprehensive score analysis to determine the model bandwidth of the kernel density estimation model, specifically including: Performing applicability classification on the applicability comprehensive scores of the adaptation samples based on the comprehensive scores of the indicators, and determining a differentiated bandwidth analysis threshold; Based on the differentiated bandwidth analysis threshold, the model bandwidth of the kernel density estimation model is determined by revising the bandwidth of the indicator differentiation.
4. The unweighted health assessment method based on kernel density according to claim 3 is characterized in that: Based on the differentiated bandwidth analysis threshold, the model bandwidth of the kernel density estimation model is determined by revising the bandwidth of the indicator differentiation, specifically including: When the comprehensive applicability score is greater than or equal to a first differentiation threshold, performing a first differentiation correction on the standard bandwidth formula to suppress outlier interference; When the comprehensive usability score is less than the first differentiation threshold and greater than or equal to the second differentiation threshold, performing a second differentiation correction on the standard bandwidth formula to retain data details of the small sample size data and reduce noise of the large sample size data; The model bandwidth of the kernel density estimation model is determined according to the bandwidth obtained by the first differential correction or the second differential correction.
5. The unweighted health assessment method based on kernel density according to claim 1, characterized in that: According to the kernel density estimation model and the model bandwidth, a multi-key indicator kernel density analysis is performed to obtain an unweighted health assessment kernel density, specifically including: Based on the kernel density estimation model and the model bandwidth, a multi-indicator kernel density estimation model is determined through multi-indicator fluctuation revision; wherein the single-indicator related parameters of the multi-indicator kernel density estimation model include: sample number, sample interval boundary, and sample bandwidth; The multi-index samples are input into the multi-index kernel density estimation model, and the unweighted health assessment kernel density is obtained through multi-index boundary correction.
6. The unweighted health assessment method based on kernel density according to claim 1, characterized in that: Calculating the interval probability value of the health level of the health assessment kernel density to determine the health status probability of the device specifically includes: Performing interval probability integration on the health assessment kernel density to obtain interval probability distribution; Device state matching is performed on the interval probability distribution to determine the device health state probability.
7. A weightless health assessment device based on kernel density, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Obtaining equipment health score indicator data, and performing kernel density applicability analysis on the equipment health score indicator data to determine an indicator comprehensive score adaptation sample; Perform boundary correction on the standard kernel density estimation function to obtain the kernel density estimation model; Based on the comprehensive score of the indicators, the model bandwidth of the kernel density estimation model is determined through differentiated comprehensive applicability score analysis of the adapted samples; According to the kernel density estimation model and the model bandwidth, a weighted health assessment kernel density is obtained through a multi-key indicator kernel density analysis; Calculating the interval probability value of the health level of the health assessment kernel density to determine the health status probability of the equipment; Perform kernel density applicability analysis on the equipment health score indicator data to determine the indicator comprehensive score adaptation sample, specifically including: Performing an indicator data volume analysis on the device health score indicator data to determine data verification content; wherein the data verification content includes: single indicator data volume constraints and multi-indicator data volume balance requirements; Based on the data verification content, a kernel density applicability comprehensive scoring formula is defined, and a comprehensive score is calculated using the kernel density applicability comprehensive scoring formula; Performing a suitability threshold judgment on the comprehensive score to determine the data volume optimization requirements of the indicator data volume; According to the data volume optimization requirements, the device health score index data is supplemented with data volume, and a sample suitable for the comprehensive score of the index is determined; The kernel density applicability comprehensive scoring formula is: in, The amount of data for key indicators, is the data volume range ratio, is the single indicator data volume score, is the data volume balance score, is the number of indicators, Provides a comprehensive score for kernel density suitability.
8. A non-volatile computer storage medium for weightless health assessment based on kernel density, storing computer executable instructions, characterized in that: The computer executable instructions are configured to: Obtaining equipment health score indicator data, and performing kernel density applicability analysis on the equipment health score indicator data to determine an indicator comprehensive score adaptation sample; Perform boundary correction on the standard kernel density estimation function to obtain the kernel density estimation model; Based on the comprehensive score of the indicators, the model bandwidth of the kernel density estimation model is determined through differentiated comprehensive applicability score analysis of the adapted samples; According to the kernel density estimation model and the model bandwidth, a weighted health assessment kernel density is obtained through a multi-key indicator kernel density analysis; Calculating the interval probability value of the health level of the health assessment kernel density to determine the health status probability of the equipment; Perform kernel density applicability analysis on the equipment health score indicator data to determine the indicator comprehensive score adaptation sample, specifically including: Performing an indicator data volume analysis on the device health score indicator data to determine data verification content; wherein the data verification content includes: single indicator data volume constraints and multi-indicator data volume balance requirements; Based on the data verification content, a kernel density applicability comprehensive scoring formula is defined, and a comprehensive score is calculated using the kernel density applicability comprehensive scoring formula; Performing a suitability threshold judgment on the comprehensive score to determine the data volume optimization requirements of the indicator data volume; According to the data volume optimization requirements, the device health score index data is supplemented with data volume, and a sample suitable for the comprehensive score of the index is determined; The kernel density applicability comprehensive scoring formula is: in, The amount of data for key indicators, is the data volume range ratio, is the single indicator data volume score, is the data volume balance score, is the number of indicators, Provides a comprehensive score for kernel density suitability.
Citation Information
Patent Citations
Abnormity detection method and device for big data integration host
CN118819994A