Electric power Internet of Things monitoring system based on multi-sensor fusion
The power Internet of Things monitoring system, which integrates multiple sensors, solves the problems of data acquisition from single sensors and unreasonable selection of reference devices. It enables multi-dimensional dynamic status monitoring of power equipment and accurate model construction, thereby improving the stability and security of the power system.
Patent Information
- Application Number
- CN202610037553.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-13
- Publication Date
- 2026-02-10
AI Technical Summary
In existing power Internet of Things (IoT) monitoring systems, data acquisition from single sensors is limited in scope and susceptible to external environmental interference. Inappropriate selection of reference equipment, lack of comprehensiveness and accuracy in data processing, and reliance on human experience in model building lead to biased and inaccurate monitoring results.
The power Internet of Things monitoring system, which adopts multi-sensor fusion, acquires multi-dimensional sensor data through the data acquisition unit. Combined with the reference equipment selection module, the assessment capability calculation module, the correlation analysis module, and the feature analysis module, it constructs a status monitoring model, accurately identifies key data points, and builds a monitoring model.
It enables multi-dimensional and dynamic status monitoring of power equipment, improves the accuracy of data acquisition and the precision of models, and can promptly identify operational anomalies in new equipment, ensuring the stable operation of the power system.
Smart Images

Figure CN121508159A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power internet of things monitoring, in particular to a power internet of things monitoring system based on multi-sensor fusion. BACKGROUND
[0002] Power equipment is an important basis for the stable operation of the power system, and real-time and accurate monitoring of its operating state plays an irreplaceable role in the safety protection of the power system. With the penetration of the internet of things technology in the power industry, the power internet of things monitoring system gradually replaces the traditional monitoring method, but such systems still face many technical bottlenecks in actual application.
[0003] Traditional power equipment monitoring mostly uses a single sensor for data collection, which can only obtain information of a certain dimension of the equipment, such as monitoring only the temperature or voltage of the equipment, and cannot fully capture the multi-dimensional changes in the operation process of the equipment, resulting in one-sidedness in the judgment of the operating state of the equipment. At the same time, the data collected by a single sensor is easily disturbed by external environmental factors, such as temperature sensors being affected by environmental temperature fluctuations and voltage sensors being disturbed by transient impact of the power grid, making it difficult to guarantee the accuracy of the collected data and further affecting the reliability of subsequent state judgment.
[0004] Even if some monitoring systems introduce multi-sensor data collection, there are still obvious defects in the data processing link. In the selection of reference equipment, existing technologies mostly determine the reference equipment according to fixed attributes such as equipment model and installation location, without considering the actual operating state of the target equipment, resulting in a large difference between the operating conditions of the reference equipment and the target equipment. Based on such reference equipment, the analysis work is difficult to fit the real operating conditions of the target equipment, and the accuracy of the analysis results is greatly discounted.
[0005] In terms of state evaluation capability calculation, existing methods mostly perform isolated analysis on single sensor data points without considering the data distribution characteristics of the reference equipment set, which cannot accurately judge the actual contribution of each data point to the evaluation of the equipment state, and is easy to include redundant data or interference data that is meaningless to the state evaluation into the analysis process, resulting in deviation of the state evaluation result.
[0006] In the correlation analysis link, existing technologies usually only calculate the initial state correlation degree without adjusting for factors such as sensor errors and environmental disturbances that may occur during data collection, which cannot truly reflect the state correlation relationship between different power equipment, may cause misjudgment of the state correlation of the equipment, and further affect the grasp of the overall operating state of the equipment.
[0007] In the feature analysis process, existing systems often only focus on features that directly reflect the operating status of the equipment, ignoring the potential impact of non-state features on the equipment operation. This results in incomplete feature extraction, failing to cover key factors that may indirectly affect the equipment status, leaving hidden dangers for subsequent model construction.
[0008] During the model application phase, the selection of key points lacks scientific basis and relies heavily on manual experience to select data points. This easily leads to the inclusion of redundant data or the omission of key data, which not only increases the computational load of model construction but also reduces the efficiency of model operation. At the same time, the construction of the monitoring model does not fully integrate key data points with equipment state variables, resulting in insufficient model accuracy. For newly commissioned power equipment, due to the lack of effective data adaptation and model adjustment, it is difficult to achieve accurate state monitoring and timely detect potential problems in the operation of new equipment, posing risks to the stable operation of the power system. Summary of the Invention
[0009] The purpose of this invention is to provide a power Internet of Things monitoring system based on multi-sensor fusion to solve the problems mentioned in the background art.
[0010] To achieve the above objectives, the present invention provides a power Internet of Things (IoT) monitoring system based on multi-sensor fusion, the system comprising:
[0011] The data acquisition unit includes multiple sensor nodes for acquiring sensor data sequences and equipment status variables for each power device in the same batch. The sensor data sequences include time-series data of temperature, humidity, voltage, and current, and the equipment status variables represent the operating status of the equipment.
[0012] The data processing unit includes a reference device selection module, an evaluation capability calculation module, a correlation analysis module, and a feature analysis module. The reference device selection module determines a set of reference power devices based on the device state variables of the target power device. The evaluation capability calculation module calculates the state evaluation capability for each sensor data point based on the data distribution of the reference power device set. The correlation analysis module calculates the initial state correlation degree based on the data distribution of all power devices and adjusts it to obtain the true state correlation degree. The feature analysis module calculates the probability of non-state features based on the state distribution of the reference power device set.
[0013] The model application unit includes a key point screening module and a monitoring model module. The key point screening module selects key data points based on the probability of non-state characteristics and the correlation with the real state. The monitoring model module uses the sensor data values of the key data points and equipment state variables to construct a state monitoring model and perform state monitoring on the new power equipment.
[0014] Preferably, when the assessment capability calculation module calculates the state assessment capability, it performs the following steps:
[0015] For the target power equipment, obtain the set of sensor data of its reference power equipment at the target data point, calculate the information entropy of the set of values, and perform reverse standardization on the information entropy to obtain the state assessment index.
[0016] Calculate the variance of the equipment state variables of the reference power equipment as the state volatility;
[0017] Construct a state assessment index sequence and a state volatility sequence, calculate the mutual information value between the two sequences, and generate state assessment weights based on the mutual information value;
[0018] The state assessment capability is obtained by multiplying the state assessment weight by the state assessment index.
[0019] Preferably, when the correlation analysis module calculates the initial state correlation, it performs the following steps: obtaining the numerical distribution of sensor data of all power equipment at the target data point, generating a numerical distribution histogram, and extracting the cumulative distribution function sequence of the numerical distribution histogram; obtaining the distribution of equipment state variables of all power equipment, generating a variable distribution histogram, and extracting the cumulative distribution function sequence of the variable distribution histogram; applying the dynamic time warping algorithm to calculate the alignment path length between two cumulative distribution function sequences, and converting the path length into a similarity score as the initial state correlation.
[0020] Preferably, when the correlation analysis module adjusts the initial state correlation to obtain the true state correlation, it performs the following steps: ranking the power equipment according to the magnitude of the equipment state variables to generate a state ranking sequence; ranking the power equipment according to the magnitude of the sensor data values to generate a data ranking sequence; calculating the Kendall rank correlation coefficient between the state ranking sequence and the data ranking sequence as the ranking consistency; after removing power equipment with consistent ranking, extracting the state assessment capability value of the remaining power equipment, constructing an assessment capability distribution map, and calculating the Jaccard similarity coefficient of the assessment capability distribution map; combining the ranking consistency and the Jaccard similarity coefficient to generate an adjustment coefficient; multiplying the adjustment coefficient by the initial state correlation to obtain the true state correlation.
[0021] Preferably, when the feature analysis module calculates the probability of non-state features, it performs the following steps: for each target power device, obtain the frequency distribution of the device state variables of its reference power device and generate a reference frequency distribution curve; obtain the global frequency distribution curve of the device state variables of all power devices; calculate the Earth's movement distance between the reference frequency distribution curve and the global frequency distribution curve as the distribution deviation value; calculate the arithmetic mean of the distribution deviation values of all power devices, and perform an inverse logarithmic transformation on the mean to obtain the probability of non-state features.
[0022] Preferably, when the key point screening module selects key data points, it performs the following steps: calculating the ratio of the true state correlation degree to the probability of non-state features for each data point to obtain the original importance value; performing minimum-maximum normalization on the original importance value to obtain a standardized importance score; setting a dynamic threshold, which is determined based on the percentile of the standardized importance scores of all data points; and selecting data points with standardized importance scores higher than the dynamic threshold as key data points.
[0023] Preferably, when the monitoring model module constructs the state monitoring model, it performs the following steps: collecting sensor data values of each power device at key data points, including the values of temperature, humidity, voltage and current, as input feature matrix; collecting equipment state variables of each power device as output target vector; applying support vector regression algorithm to train the input feature matrix and output target vector, optimizing kernel function parameters and penalty coefficients, and generating state monitoring model.
[0024] Preferably, when the monitoring model module applies the state monitoring model, it performs the following steps: acquiring the sensor data sequence of the power equipment to be monitored, extracting the sensor data values of the power equipment to be monitored at key data points, and forming an input data vector; inputting the input data vector into the state monitoring model, and outputting the predicted value of the equipment state variable through the calculation of the support vector regression model; and performing post-processing on the predicted value, including noise reduction and smoothing, to obtain the final state monitoring result.
[0025] Preferably, when the reference equipment selection module determines the reference power equipment set, it performs the following steps: taking the equipment state variable of the target power equipment as the center value, setting a symmetrical tolerance interval, the width of the symmetrical tolerance interval being adaptively adjusted based on the historical variation coefficient of the equipment state variable; screening all power equipment whose equipment state variables are within the tolerance interval to form a candidate set; randomly sampling from the candidate set or selecting based on geographical distance weighting to finally determine the reference power equipment set.
[0026] Preferably, the data acquisition unit further includes a data preprocessing module; the data preprocessing module is responsible for cleaning, aligning and interpolating the sensor data.
[0027] Compared with the prior art, the beneficial effects of the present invention are:
[0028] In the data acquisition phase, the system employs multiple sensor nodes to collect sensor data sequences and equipment status variables for each power device in the same batch. The sensor data sequences encompass time-series data for temperature, humidity, voltage, and current. Compared to traditional monitoring methods where a single sensor can only acquire static data in one dimension, this data acquisition method obtains multi-dimensional and dynamic operational information about the equipment. It more comprehensively reflects the changes in the equipment's operating status over different time periods, avoiding the biased monitoring caused by single-dimensional or static data failing to capture the dynamic operation of the equipment. This provides richer and more accurate foundational data for subsequent data processing and model building, better reflecting the actual operating conditions of the equipment.
[0029] In the data processing stage, the reference equipment selection module determines the reference power equipment set based on the equipment state variables of the target power equipment. This breaks the limitation of traditional technology that relies on fixed equipment attributes to select reference equipment, ensuring that the reference equipment and the target equipment have a high degree of consistency in their operating states. This allows subsequent analysis based on the reference set to better match the actual operating conditions of the target equipment, reducing analytical bias caused by reference equipment mismatch from the source and laying the foundation for accurate analysis in the future.
[0030] The assessment capability calculation module calculates the condition assessment capability for each sensor data point, combining it with the data distribution of a reference power equipment set. This allows for the accurate identification of the actual contribution of each data point to the equipment condition assessment. In this way, invalid data points that are meaningless or highly interfering with the condition assessment can be effectively eliminated, preventing invalid data from entering subsequent analysis processes. This ensures that the data used in the condition assessment process is highly effective and relevant, improving the accuracy of the condition assessment results.
[0031] The correlation analysis module adjusts the initial state correlation to obtain the true state correlation, effectively eliminating initial correlation bias caused by environmental interference, sensor errors, and other factors during data acquisition. The adjusted true state correlation more objectively and accurately reflects the state relationships between different power devices, avoiding misjudgments of device state correlations due to inaccurate initial correlation. This provides a more reliable correlation basis for subsequent feature analysis and model construction, helping to improve the ability to grasp the overall operating status of the equipment.
[0032] The feature analysis module calculates the probability of non-state features based on the state distribution of a reference set of power equipment. This overcomes the limitations of traditional feature analysis, which focuses solely on direct state features. It can identify features that do not directly belong to the equipment's operating state but may affect it. The inclusion of these non-state features makes feature extraction more comprehensive, covering more potential factors influencing equipment operating state. This provides richer feature support for subsequent key point selection and model construction, helping to improve the model's sensitivity and recognition ability to changes in equipment state.
[0033] In the model application phase, the key point selection module combines the probability of non-state features with the correlation with the actual state to select key data points, replacing the traditional selection method that relies on manual experience. This allows for the accurate selection of data points that are of significant importance to equipment status monitoring. This scientific selection method not only effectively reduces the inclusion of redundant data and lowers the computational load for subsequent model construction, improving the model's operating efficiency, but also avoids the problem of insufficient model monitoring capabilities due to the omission of key data points, ensuring the effectiveness and relevance of the model's input data.
[0034] The monitoring model module utilizes sensor data values from key data points and equipment state variables to construct a condition monitoring model. Compared to traditional models that rely on only partial or irrelevant data, this model uses more targeted and effective input data, and its construction better reflects the actual operating patterns of the equipment, resulting in higher accuracy. When monitoring new power equipment, this model better adapts to the operating characteristics of the new equipment, accurately captures changes in its operating status, and promptly identifies anomalies. This avoids monitoring failures caused by the model's inability to adapt to the new equipment, providing strong support for the safe and stable operation of new power equipment. Attached Figure Description
[0035] Figure 1 This is a schematic diagram illustrating the working principle of the power Internet of Things monitoring system based on multi-sensor fusion as described in this invention.
[0036] Figure 2 A flowchart for calculating the status assessment capability of the capability calculation module;
[0037] Figure 3 A flowchart for adjusting the initial state correlation degree of the correlation analysis module to obtain the true state correlation degree. Detailed Implementation
[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0039] Please see Figure 1 This invention provides a power Internet of Things (IoT) monitoring system based on multi-sensor fusion. The system includes: a data acquisition unit composed of multiple sensor nodes, which are responsible for collecting sensor data sequences and equipment state variables for each power device in the same batch; the sensor data sequences cover time-series data of temperature, humidity, voltage, and current, while the equipment state variables directly reflect the operating status of the power devices. A data processing unit integrates a reference device selection module, an evaluation capability calculation module, a correlation analysis module, and a feature analysis module. The reference device selection module determines a set of reference power devices based on the equipment state variables of the target power device; the evaluation capability calculation module calculates the state evaluation capability for each sensor data point based on the data distribution of the reference power device set; the correlation analysis module calculates the initial state correlation degree based on the data distribution of all power devices and further adjusts it to obtain the true state correlation degree; and the feature analysis module calculates the probability of non-state features based on the state distribution of the reference power device set. A model application unit includes a key point screening module and a monitoring model module; the key point screening module selects key data points based on the probability of non-state features and the true state correlation degree; and the monitoring model module uses the sensor data values of the key data points and equipment state variables to construct a state monitoring model and perform state monitoring on new power devices.
[0040] Example 1: See Figure 2The assessment capability calculation module begins its process by acquiring data from the target power equipment and its reference power equipment set. The reference power equipment set is pre-determined by the reference equipment selection module. The module reads the corresponding sensor values of all equipment within the reference power equipment set at the target data point. These values constitute a numerical set for analysis. The calculation of the information entropy of this numerical set is fundamental to subsequent steps. Information entropy is quantified based on the probability distribution of values appearing in the set. The probability of each value is obtained by dividing its frequency by the total number of values in the set. The negative sum of the products of probabilities and their base-2 logarithms is the information entropy value of the numerical set. The information entropy value reflects the degree of disorder or uncertainty of the data points. Inverse standardization of the information entropy maps the original entropy value to a standardized range of zero to one. Inverse standardization uses a linear transformation method, employing pre-statistically calculated global minimum and maximum information entropies as boundaries. High information entropy is converted to low state assessment indicators, and low information entropy is converted to high state assessment indicators. The state assessment indicators thus characterize the centralization of the data distribution.
[0041] The variance of the equipment state variables of the reference power equipment is calculated as the state volatility. State volatility is obtained by averaging the squared deviations of the equipment state variable values from their average values. The variance calculation process includes several steps: averaging, calculating the difference between each value and the average, squaring these differences, and finally averaging them. The level of state volatility directly reflects the overall stability of the reference power equipment set; a higher state volatility indicates greater state variation among the reference equipment. Constructing the state assessment index sequence and the state volatility sequence requires maintaining a one-to-one correspondence between sequence elements. Each reference power equipment contributes one state assessment index and one state volatility value, and the lengths of the two sequences are strictly equal to the number of reference power equipment. The mutual information value of the state assessment index sequence and the state volatility sequence is calculated using a statistical method based on histogram distribution. The joint probability distribution and individual marginal probability distributions of the two sequences are estimated through binning statistics. The mutual information value is calculated using a formula based on the probability distribution, and it measures the degree of statistical dependence between the two sequences.
[0042] The generation of state assessment weights based on mutual information values involves a linear scaling process. The mutual information value is divided by a theoretically maximum possible value to achieve normalization, and the normalized value is directly used as the state assessment weight. The state assessment weight is multiplied by the previously obtained state assessment index, and the product is the state assessment capability of the target data point. State assessment capability is a comprehensive measure that integrates information on data distribution and state fluctuations. The level of the state assessment capability value indicates the effectiveness of this sensor data point in assessing the state of the equipment. The correlation analysis module calculates the initial state correlation degree independently of the assessment capability calculation module. The correlation analysis module needs to obtain the numerical distribution of sensor data from all power equipment at the target data point. The data point values of the entire power equipment group are collected to generate a numerical distribution histogram. The generation of the numerical distribution histogram adopts an equal-width binning strategy. The number of bins is automatically determined according to Scott's rule, which considers the standard deviation of the data and the total number of equipment. The number of equipment falling within the interval of each bin is counted. Extracting the cumulative distribution function sequence from the numerical distribution histogram is the key transformation from the histogram to the distribution function. The cumulative distribution function sequence is obtained by sequentially summing the number of devices in each box and dividing by the total number of devices. Each point in the sequence represents the proportion of all devices whose value is less than or equal to the right boundary of that box.
[0043] Obtaining the distribution of device state variables for all power equipment and generating variable distribution histograms follows a binning principle similar to that used for numerical distribution histograms. The value range and distribution characteristics of the device state variables determine the specific binning parameters. The method for extracting the cumulative distribution function sequence from the variable distribution histogram is identical to that used when processing sensor data. The cumulative distribution function sequence accurately describes the cumulative probability distribution of the device state variables. The core step is calculating the alignment path length between two cumulative distribution function sequences using a dynamic time warping algorithm. This algorithm constructs a cost matrix to find the optimal curved path between the two sequences; the sum of the costs at each point on the path is the alignment path length, which quantifies the difference in the shapes of the two distributions. Converting the alignment path length to a similarity score uses an exponential decay function. The conversion formula maps longer path lengths to lower similarity scores and shorter path lengths to higher similarity scores. The similarity score is output as the initial state correlation. The initial state correlation reflects the overall similarity between the distribution of sensor values and the distribution of device state variables at the target data point. A higher initial state correlation indicates a strong potential correlation between sensor data and device states.
[0044] Example 2: See Figure 3The correlation analysis module adjusts the initial state correlation to obtain the true state correlation, which is a multi-stage, refined calculation process. The correlation analysis module receives initial state correlation data from preceding modules, along with relevant device state variables and sensor data. Based on the magnitude of the device state variables, power devices are ranked to generate a state ranking sequence. The ranking operation uses an ascending order method, with the power device with the smallest state variable value receiving ranking one, the next smallest value receiving ranking two, and so on, until all devices have a unique ranking position. When power devices have identical state variable values, these devices are assigned the average of their respective ranking positions. Based on the magnitude of the sensor data values, power devices are ranked to generate a data ranking sequence. The generation rules for the data ranking sequence are strictly consistent with those for the state ranking sequence. The data ranking sequence is also based on ascending order, and tied values are averaged for ranking. The state ranking sequence and the data ranking sequence have a clear correspondence in the order of elements, and each power device has a ranking value in both sequences.
[0045] The Kendall's rank correlation coefficient between the state ranking sequence and the data ranking sequence is used as the ranking consistency. The calculation of the Kendall's rank correlation coefficient is based on a consistency check of all possible power device pairings. For any two different power devices forming a pair, the order of this pair in the state ranking sequence and the data ranking sequence is checked. Consistent order means that one device ranks higher than the other in both sequences, while inconsistent order means that one device ranks higher in one sequence and lower in the other. The number of consistent pairings minus the number of inconsistent pairings, divided by the total number of possible pairings, yields the Kendall's rank correlation coefficient. The Kendall's rank correlation coefficient ranges from -1 to +1, where +1 indicates that the two ranking sequences are completely consistent, -1 indicates that they are completely opposite, and zero indicates no correlation. The calculated Kendall's rank correlation coefficient value is directly defined as the ranking consistency. Devices with consistent rankings are removed based on a preset consistency judgment threshold. The pairing consistency information implicit in the ranking consistency calculation is used to identify a subset of power devices whose relative order in the state ranking and data ranking remains consistent. These power devices are temporarily excluded from subsequent analysis, and the analysis focuses on the group of power devices with inconsistent rankings. For the remaining equipment, extract its condition assessment capability values to construct an assessment capability distribution map. The condition assessment capability values are provided in advance by the assessment capability calculation module. Each retained power equipment contributes a specific condition assessment capability value. These values constitute a dataset used to generate the assessment capability distribution map. The generation of the assessment capability distribution map adopts the kernel density estimation method. The kernel density estimation uses a Gaussian kernel function and optimizes the bandwidth parameter to smoothly approximate the probability density function of the condition assessment capability value. The resulting assessment capability distribution map is a continuous probability density curve.
[0046] Calculating the Jaccard similarity coefficient of the assessment capability distribution map requires defining an ideal reference distribution. The ideal reference distribution is set as a normal distribution centered on the global average value of the state assessment capability. The calculation of the Jaccard similarity coefficient is done by discretizing the assessment capability distribution map and the ideal distribution map into multiple intervals in the domain, and counting which intervals the two distribution maps have values in simultaneously (i.e., in the intersection), and at least one of the distribution maps has a value (i.e., in the union). The Jaccard similarity coefficient is equal to the number of intersection intervals divided by the number of union intervals. The range of the Jaccard similarity coefficient is from zero to one. The larger the value, the closer the shape of the assessment capability distribution map is to the ideal distribution map. The adjustment coefficient is generated by combining ranking consistency and Jaccard similarity coefficient using a weighted geometric mean. The ranking consistency needs to be linearly transformed to map its range from negative one to positive one to zero to facilitate subsequent calculations. The transformed ranking consistency and Jaccard similarity coefficient are assigned pre-set weight factors. The values of the weight factors are determined based on historical data analysis to reflect the relative importance of the two factors in contributing to the adjustment coefficient. The adjustment coefficient is equal to the weight raised to the power of the transformed ranking consistency multiplied by the weight raised to the power of the Jaccard similarity coefficient. The weighted geometric mean can balance the influence of ranking consistency and Jaccard similarity coefficient. The true state correlation is obtained by multiplying the adjustment coefficient by the initial state correlation. The multiplication is performed directly between the two values. The initial state correlation is derived from the correlation analysis module in the preliminary calculation. The result, the true state correlation, is a corrected correlation metric. The true state correlation considers not only the global data distribution similarity but also incorporates local information on ranking consistency and the distribution characteristics of state assessment capabilities. Compared to the initial state correlation, the true state correlation more accurately characterizes the true correlation strength between sensor data points and device state variables. The true state correlation is then passed to the key point screening module for subsequent key data point selection decisions. The entire adjustment process, by introducing ranking consistency and distribution similarity metrics, effectively corrects for initial correlation biases that may be caused by heterogeneity of the device group or data noise.
[0047] Example 3: The processing flow for calculating the probability of non-state features in the feature analysis module is based on the determination of a reference power device set, which is generated independently by the reference device selection module for each target power device. For each target power device, the frequency distribution of the device state variables in its reference power device set is obtained. The frequency distribution is constructed by counting the number of devices whose device state variable values fall within several preset equal-width intervals. The width of the interval is determined by the Sturgess rule, which takes into account the number of devices in the reference power device set. Generating the reference frequency distribution curve requires converting the discrete frequency distribution into a continuous representation. The conversion process uses a linear interpolation method to connect the frequency values at the midpoint of each interval, forming a piecewise linear frequency polygon curve. This curve represents the distribution of the target power device's neighboring devices in the state space.
[0048] Obtaining the global frequency distribution curves of all power equipment state variables employs a method similar to that used to generate the reference frequency distribution curve. However, the statistical basis is the complete power equipment population rather than a subset. The interval division rules for the global frequency distribution curves are consistent with those for the reference frequency distribution curves to ensure comparability. The global frequency distribution curves represent the background distribution characteristics of the entire equipment population's state. Before calculating the distribution bias, the equipment state variable data must be preprocessed to eliminate the influence of dimensions. The equipment state variable data is normalized to a closed interval between zero and one. The normalization formula uses the historical minimum and maximum values of the equipment state variables as scaling benchmarks. Calculating the Earth's movement distance between the reference frequency distribution curve and the global frequency distribution curve requires treating the two curves as two probability distributions. The Earth's movement distance is defined as the minimum amount of work required to transform one distribution into another. The measure of this work is the integral of the probability mass multiplied by the movement distance. The Earth's movement distance is calculated by discretizing the continuous curve into probability masses at a large number of equally spaced points, solving a linear programming problem to obtain the optimal transportation scheme, and this distance value is recorded as the distribution bias value.
[0049] Calculating the arithmetic mean of the distribution deviations for all power devices involves iterating through every power device in the dataset. Each power device is treated as a target device, and a distribution deviation value is calculated for each. The sum of all these distribution deviation values is then divided by the total number of power devices to obtain the arithmetic mean. An inverse logarithmic transformation is applied to the arithmetic mean to obtain the probabilities of non-state features. The mathematical expression used for this transformation is:
[0050]
[0051] Where: symbol Representing the calculated probability of non-state characteristics, it is a dimensionless scalar value. (Symbol) Represents the natural logarithm operation. (Symbol) This represents the arithmetic mean of the distribution deviation values of all power equipment calculated above; since the equipment state variables have already been normalized, the distribution deviation values calculated based on them... and its average value All are dimensionless quantities. Constants The logarithmic operation is incorporated into the average to ensure that the logarithmic parameter is always greater than zero, avoiding mathematical undefined cases. Both sides of the formula are dimensionless, maintaining a consistent dimension. Non-state characteristic probability. The arithmetic mean of the numerical values and the distribution deviation values They are inversely proportional. The smaller the value, the more typical the difference between the reference distribution and the global distribution, and the greater the possibility that the sensor data is affected by factors other than the device status.
[0052] The key point selection module considers two input parameters simultaneously: the probability of non-state features and the correlation with the true state. The correlation with the true state is provided by the correlation analysis module. The raw importance value is obtained by calculating the ratio of the true state correlation to the probability of non-state features for each data point. The ratio is calculated using direct division, with the true state correlation as the numerator and the probability of non-state features as the denominator. The raw importance value is a comprehensive indicator that reflects both the correlation strength between the data point and the state, and its vulnerability to interference from non-state factors. The raw importance values are then subjected to min-max normalization, linearly mapping all raw importance values to a closed interval between zero and one. The normalization formula uses the minimum and maximum raw importance values in the dataset as scaling boundaries. Each raw importance value is subtracted from the minimum value and divided by the difference between the maximum and minimum values to obtain the corresponding standardized importance score. A dynamic threshold is set based on the distribution characteristics of the standardized importance scores of all data points. The dynamic threshold is set as a specific percentile of all standardized importance scores, specifically the 75th percentile, or the upper quartile. Data points with standardized importance scores higher than a dynamic threshold are selected as key data points. The screening process iterates through and evaluates each data point, comparing its standardized importance score with the dynamic threshold. Data points with scores greater than the threshold are retained and marked as key data points. The set of key data points constitutes the feature set used for subsequent state monitoring model training. These data points are considered to have high state indication value and low risk of non-state interference. The output of the key point screening module determines the input feature space of the monitoring model module. The screening logic reflects the design principle of seeking a balance between information content and robustness. The entire implementation method, through quantitative distribution comparison and threshold screening, automatically identifies the core data points with the strongest state representation capability from massive sensor data sequences. The cascaded operation of the feature analysis module and the key point screening module constitutes a key step in data refinement, laying the data foundation for building a high-performance state monitoring model. The calculation of the probability of non-state features introduces a global distribution perspective, reducing the selection bias risk that may be caused by local reference sets.
[0053] Example 4: When constructing the state monitoring model, the workflow begins with the collection of sensor data at key data points. The key point screening module pre-determines the locations of data points with high state characterization value. Collecting sensor data values for each power device at key data points requires extracting the corresponding temperature, humidity, voltage, and current values from the storage system. The four sensor readings for each power device at each key data point are organized into a feature vector. Forming the input feature matrix requires arranging the feature vectors of all power devices row-wise. The row index of the matrix corresponds to different power devices, and the column index corresponds to different combinations of key data points and sensor types. The dimension of the input feature matrix is four times the number of power devices multiplied by the number of key data points. Collecting the device state variables for each power device requires reading the state value corresponding to each device from the device state database. Device state variables are typically continuous or discrete numerical indicators. The device state variables of all devices are arranged sequentially to form the output target vector, the length of which is consistent with the number of power devices. The input feature matrix and the output target vector must be strictly aligned in order to ensure that the feature vector of each power device corresponds to its correct device state variable value.
[0054] Training the input feature matrix and output target vector using the Support Vector Regression (SVR) algorithm is an optimization process. The goal of SVR is to find a function such that the deviation between the predicted values of most training samples and the actual device state variables does not exceed a predetermined tolerance. Optimizing the kernel function parameters and penalty coefficients employs a grid search cross-validation strategy. The grid search systematically traverses all possible parameter combinations within a predefined parameter range. For each parameter combination, K-fold cross-validation is used to evaluate the predictive performance of the SVR model, and the parameter combination with the smallest average error in cross-validation is selected as the optimal parameter. The training of the SVR model is based on the principle of minimizing structural risk. The model attempts to achieve a balance between fitting accuracy and model complexity, and the resulting condition monitoring model can learn the complex nonlinear mapping relationship from multi-sensor data to device state variables. When applying the condition monitoring model to monitor new equipment, the first step in the monitoring model module is to obtain the sensor data sequence of the power equipment to be monitored. The sensor data sequence contains complete time-series records of temperature, humidity, voltage, and current. Extracting sensor data values from key data points in the sensor data sequence of the power equipment to be monitored requires data extraction based on the location index provided by the key point filtering module. The values of the four sensors at each key data point are extracted and arranged in the same order as during the training phase. Forming the input data vector requires concatenating all extracted sensor values into a one-dimensional vector. The dimension of the input data vector is exactly the same as the number of columns in the input feature matrix during the condition monitoring model training. Each element in the vector represents the reading of a sensor at a specific key data point.
[0055] The input data vector is fed into the condition monitoring model. The calculation process of the support vector regression model involves mapping the input data vector to a high-dimensional feature space and performing linear regression in that space. The predicted value of the output device state variable is a real number, representing the estimate of the operating state of the monitored power equipment by the support vector regression model based on the input sensor data. Post-processing of the predicted value aims to improve the stability and smoothness of the prediction results. Denoising is achieved using a moving median-based filter to combat impulse interference, and smoothing is achieved using a low-pass digital filter to suppress high-frequency fluctuations in the prediction results. Obtaining the final condition monitoring result marks the completion of the condition monitoring task. The final condition monitoring result is output as a post-processed predicted value of the device state variable for use by system users or upper-level applications (see Table 1).
[0056] Table 1: Training Data Table for Power Equipment Condition Monitoring Model
[0057]
[0058] The choice of kernel function in Support Vector Regression (SVR) affects the model's ability to capture nonlinear relationships. Radial basis function (RBF) kernels can handle complex nonlinear patterns between input features and the output target. The penalty coefficient controls the model's tolerance for outliers in the training data. A larger penalty coefficient forces the model to minimize training error, potentially leading to overfitting, while a smaller penalty coefficient allows for larger training errors but may improve the model's generalization ability. In grid search cross-validation, the K value is typically set to five or ten. The training data is randomly divided into K subsets, each of which is used as the validation set, and the remaining K-1 subsets are used as the training set. This process is repeated K times, and the average performance is used as the evaluation result for this parameter combination. Preprocessing of the input feature matrix includes standardization, where the value of each feature column is subtracted from its mean and then divided by its standard deviation, resulting in all features having zero mean and unit variance. The output target vector is standardized using a similar Z-score standardization method. Standardization helps the SVR optimization process converge to a better solution. After the support vector regression model is trained, the model parameters include the weight coefficients of the support vectors, the bias term, and the parameters of the kernel function. These parameters together define the mapping function from input features to device state variables.
[0059] The preprocessing method for the sensor data sequences of the power equipment to be monitored must be completely consistent with the preprocessing method for the training data, including the same standardized parameters: mean and standard deviation. The construction order of the input data vectors must strictly correspond to the column order of the input feature matrix during the training phase; any inconsistency in order will lead to incorrect predictions from the model. The calculation process of the support vector regression model can be represented as a kernel function transformation of a linear combination of support vectors plus a bias term. The calculation process depends on the support vectors obtained during the training phase and their corresponding weight coefficients. In the post-processing stage, the window size of the moving median filter is set to an odd number, typically three or five. The filter takes the median of the data within the window as the output for the current point each time. The cutoff frequency of the low-pass digital filter is selected based on the actual changing characteristics of the equipment state variables. The cutoff frequency should be set lower than the normal fluctuation frequency of the equipment state variables to effectively smooth random fluctuations. The output format of the final state monitoring result is consistent with the format of the equipment state variables in the training data; it can be a fraction normalized to zero or one, or an engineering unit value with actual physical meaning.
[0060] Example 5: When the reference equipment selection module determines the set of reference power equipment, its workflow is illustrated using a specific power equipment condition monitoring scenario as an example. Assume the current reading of the target power equipment's equipment condition variable is 0.85 (this variable has been normalized to the 0-1 range). A symmetrical tolerance interval is set with the target power equipment's equipment condition variable as the center value, with 0.85 as the midpoint of the interval. The width of the symmetrical tolerance interval is adaptively adjusted based on the historical coefficient of variation of the equipment condition variable. The historical coefficient of variation is calculated by analyzing the data of all power equipment condition variables over the past month, using the formula: standard deviation divided by the mean. Assuming the calculated historical coefficient of variation is 0.12, the width of the symmetrical tolerance interval is set to twice the coefficient of variation, i.e., 0.24. Therefore, the lower bound of the symmetrical tolerance interval centered at 0.85 is 0.85 - 0.24 / 2 = 0.73, and the upper bound is 0.85 + 0.24 / 2 = 0.97. The system filters all power devices whose device state variables fall within the tolerance range of 0.73 to 0.97, forming a candidate set. This filtering process is implemented through database queries, comparing the latest device state variable value of each power device to see if it falls within this range. Assuming there are 200 power devices in the system, after filtering, 45 devices have device state variable values between 0.73 and 0.97, forming the candidate set. A final reference set of power devices is determined by random sampling or a selection based on geographical distance weighting from the candidate set. Random sampling involves randomly selecting 15 devices without replacement from the 45 candidates. The geographical distance weighting method requires obtaining the geographical coordinates (e.g., latitude and longitude) of the target power device and the geographical coordinates of the 45 devices in the candidate set. The spherical distance between the target device and each candidate device is calculated, and the reciprocal of this distance is used as a weighting factor. A roulette wheel selection algorithm is used to probabilistically select 15 devices from the candidate set based on these weights. The final reference set of power devices contains 15 devices that are similar in state and potentially geographically adjacent.
[0061] The data acquisition unit includes a data preprocessing module responsible for cleaning, aligning, and interpolating sensor data. This module performs these operations before the sensor data is input into the data processing unit. Data cleaning removes outliers and missing values. Outlier detection uses the Z-score method. For temperature sensor data sequences, the mean μ and standard deviation σ of the entire sequence are calculated. The difference between each data point and the mean μ is divided by the standard deviation σ to obtain the Z-score. Data points with an absolute Z-score greater than 3 are considered outliers. Missing value handling directly deletes data segments with more than five consecutive missing points. Sporadic missing values are marked with a special identifier. Data alignment synchronizes the time series of different sensors to a unified timestamp. Voltage, current, temperature, and humidity sensors may collect data at slightly different time intervals. The data alignment module uses a standard time series generated by a master clock source (e.g., GPS clock) as a reference, finds the standard timestamp closest to the timestamp of each sensor data point, and resamples all sensor data onto the unified standard time series. Data interpolation is used to fill data gaps caused by resampling or sporadic missing data. Piecewise cubic spline interpolation is selected as the interpolation method. In the case of missing observations at a certain standard timestamp in the temperature sensor sequence, a cubic spline curve is fitted using three valid data points before and after the missing point, and the function value of the curve at that timestamp is calculated as the interpolation result.
[0062] The symmetric tolerance interval adaptive mechanism of the reference equipment selection module can flexibly adjust the interval range according to the overall fluctuation of equipment status. The historical coefficient of variation reflects the dispersion of equipment status variables. A large coefficient of variation indicates large differences in equipment status, requiring a wider tolerance interval to obtain a sufficient number of reference equipment. A small coefficient of variation indicates concentrated equipment status, allowing for a narrower tolerance interval to improve the homogeneity of reference equipment. The formation of the candidate set depends on real-time data of equipment status variables. The system needs to maintain and update an equipment status variable database. Database query operations need to be optimized to ensure rapid screening of qualified candidate equipment from a large number of power devices. The random sampling method is simple and easy to implement, ensuring the random representativeness of the reference equipment set and avoiding systematic bias. The geographical distance weighted selection method is based on the assumption that the clustering of power equipment in physical space may reflect the similarity of their operating environments. Geographically close equipment may experience similar environmental temperatures, humidity, and grid load conditions. Weighted selection makes it more likely that neighboring equipment will be selected into the reference power equipment set.
[0063] The Z-score threshold of 3 in the data cleaning stage of the data preprocessing module is based on the assumption of normal distribution, which means that 99.7% of the data points should fall within three standard deviations above and below the mean. Points outside this range are highly likely to be anomalies. In missing value handling, data segments with more than five consecutive missing points are deleted entirely to avoid unreliable interpolation results; sporadic missing values are left for subsequent interpolation processing. The uniform timestamp sequence interval for data alignment is set according to monitoring requirements, for example, one sampling point per second, with all sensor data aligned to the whole second. Piecewise cubic spline interpolation ensures the smoothness of the interpolation curve and approximates the true change trajectory of sensor data better than linear interpolation. The collaborative work of the reference device selection module and the data preprocessing module constitutes the core of the data preparation stage of the power IoT monitoring system. The dynamic range and flexible selection strategy of the reference device selection module ensure the quality of the reference set, while the cleaning, alignment, and interpolation operations of the data preprocessing module ensure the reliability of the input data. The implementation details of these two modules directly affect the accuracy and stability of subsequent modules such as state assessment capability calculation, correlation analysis, and feature analysis.
[0064] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A power Internet of Things monitoring system based on multi-sensor fusion, characterized in that, The system includes: The data acquisition unit includes multiple sensor nodes for acquiring sensor data sequences and equipment status variables for each power device in the same batch. The sensor data sequences include time-series data of temperature, humidity, voltage, and current, and the equipment status variables represent the operating status of the equipment. The data processing unit includes a reference device selection module, an evaluation capability calculation module, a correlation analysis module, and a feature analysis module. The reference device selection module determines a set of reference power devices based on the device state variables of the target power device. The evaluation capability calculation module calculates the state evaluation capability for each sensor data point based on the data distribution of the reference power device set. The correlation analysis module calculates the initial state correlation degree based on the data distribution of all power devices and adjusts it to obtain the true state correlation degree. The feature analysis module calculates the probability of non-state features based on the state distribution of the reference power device set. The model application unit includes a key point screening module and a monitoring model module. The key point screening module selects key data points based on the probability of non-state characteristics and the correlation with the real state. The monitoring model module uses the sensor data values of the key data points and equipment state variables to construct a state monitoring model and perform state monitoring on the new power equipment.
2. The power Internet of Things monitoring system based on multi-sensor fusion according to claim 1, characterized in that, When calculating the status assessment capability, the assessment capability calculation module performs the following steps: For the target power equipment, obtain the set of sensor data of its reference power equipment at the target data point, calculate the information entropy of the set of values, and perform reverse standardization on the information entropy to obtain the state assessment index. Calculate the variance of the equipment state variables of the reference power equipment as the state volatility; Construct a state assessment index sequence and a state volatility sequence, calculate the mutual information value between the two sequences, and generate state assessment weights based on the mutual information value; The state assessment capability is obtained by multiplying the state assessment weight by the state assessment index.
3. The power Internet of Things monitoring system based on multi-sensor fusion according to claim 2, characterized in that, When calculating the initial state correlation degree, the correlation analysis module performs the following steps: obtain the numerical distribution of sensor data of all power equipment at the target data point, generate a numerical distribution histogram, and extract the cumulative distribution function sequence of the numerical distribution histogram; obtain the distribution of equipment state variables of all power equipment, generate a variable distribution histogram, and extract the cumulative distribution function sequence of the variable distribution histogram; apply the dynamic time warping algorithm to calculate the alignment path length between two cumulative distribution function sequences, and convert the path length into a similarity score as the initial state correlation degree.
4. The power Internet of Things monitoring system based on multi-sensor fusion according to claim 3, characterized in that, When the correlation analysis module adjusts the initial state correlation to obtain the true state correlation, it performs the following steps: ranking the power equipment according to the magnitude of the equipment state variables and generating a state ranking sequence; ranking the power equipment according to the magnitude of the sensor data values and generating a data ranking sequence. Calculate the Kendall rank correlation coefficient between the state ranking sequence and the data ranking sequence as the ranking consistency. After removing power equipment with consistent ranking, extract the state assessment capability value for the remaining power equipment, construct an assessment capability distribution map, and calculate the Jaccard similarity coefficient of the assessment capability distribution map. Combine the ranking consistency and the Jaccard similarity coefficient to generate the adjustment coefficient. Multiply the adjustment coefficient by the initial state correlation to obtain the true state correlation.
5. The power Internet of Things monitoring system based on multi-sensor fusion according to claim 4, characterized in that, When calculating the probability of non-state features, the feature analysis module performs the following steps: For each target power device, obtain the frequency distribution of the device state variables of its reference power device and generate a reference frequency distribution curve; obtain the global frequency distribution curve of the device state variables of all power devices; calculate the Earth's movement distance between the reference frequency distribution curve and the global frequency distribution curve as the distribution deviation value; calculate the arithmetic mean of the distribution deviation values of all power devices, and perform an inverse logarithmic transformation on the mean to obtain the probability of non-state features.
6. The power Internet of Things monitoring system based on multi-sensor fusion according to claim 5, characterized in that, When the key point screening module selects key data points, it performs the following steps: calculates the ratio of the true state correlation degree to the probability of non-state features for each data point to obtain the original importance value; performs min-max normalization on the original importance value to obtain a standardized importance score. A dynamic threshold is set, which is determined based on the percentile of the standardized importance score of all data points; data points with a standardized importance score higher than the dynamic threshold are selected as key data points.
7. The power Internet of Things monitoring system based on multi-sensor fusion according to claim 6, characterized in that, When the monitoring model module constructs the condition monitoring model, it performs the following steps: collecting sensor data values of each power device at key data points, including the values of temperature, humidity, voltage and current, as input feature matrix; Collect the device state variables of each power device as the output target vector; The support vector regression algorithm is applied to train the input feature matrix and the output target vector, optimize the kernel function parameters and penalty coefficient, and generate a state monitoring model.
8. The power Internet of Things monitoring system based on multi-sensor fusion according to claim 7, characterized in that, When the monitoring model module applies the state monitoring model, it performs the following steps: acquiring the sensor data sequence of the power equipment to be monitored, extracting the sensor data values of the power equipment to be monitored at key data points, and forming an input data vector; inputting the input data vector into the state monitoring model, and outputting the predicted value of the equipment state variable through the calculation of the support vector regression model; The predicted values are post-processed, including denoising and smoothing, to obtain the final state monitoring results.
9. The power Internet of Things monitoring system based on multi-sensor fusion according to claim 8, characterized in that, When the reference equipment selection module determines the reference power equipment set, it performs the following steps: taking the equipment state variable of the target power equipment as the center value, setting a symmetrical tolerance interval, the width of the symmetrical tolerance interval is adaptively adjusted based on the historical variation coefficient of the equipment state variable; screening all power equipment whose equipment state variables are within the tolerance interval to form a candidate set; randomly sampling from the candidate set or selecting based on geographical distance weighting to finally determine the reference power equipment set.
10. The power Internet of Things monitoring system based on multi-sensor fusion according to claim 1, characterized in that, The data acquisition unit also includes a data preprocessing module; the data preprocessing module is responsible for cleaning, aligning and interpolating the sensor data.
Citation Information
Patent Citations
Equipment surface interaction and state sensing system based on distributed piezoelectric sensor
CN121113197A