Multi-source heterogeneous environment data credibility evaluation method and system based on kernel density estimation
Through the kernel density estimation method, time synchronization, outlier processing and comprehensive evaluation of multi-source heterogeneous environmental data is solved, and the problem of data fluctuations and correlations not being considered in the traditional method is achieved, more accurate data credibility assessment is achieved, and the operation and management of new energy power generation systems is improved.
Patent Information
- Application Number
- CN202510478470.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-04-16
AI Technical Summary
When traditional data evaluation methods process multi-source heterogeneous environment data, they cannot fully consider the complex fluctuations and correlation of the data, resulting in inaccurate evaluation results.
Using a method based on kernel density estimation, data is collected through multiple sensors and time synchronization, outlier value detection and correction, and data filtering is performed. The known accurate samples are modeled using kernel density estimation, and the accuracy, stability and consistency of the data are calculated, and the credibility of the data is comprehensively evaluated.
It improves the evaluation accuracy and reliability of multi-source heterogeneous environmental data, and is suitable for a variety of environmental monitoring scenarios, especially in the field of new energy power generation, and provides quantitative evaluation indicators.
Smart Images

Figure CN120524409A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of renewable energy power generation environmental monitoring technology, and in particular to a method and system for evaluating the credibility of multi-source heterogeneous environmental data based on kernel density estimation, which performs credibility evaluation on environmental factor data such as temperature, humidity, and wind speed collected by multiple sensors. Background Art
[0002] In the field of renewable energy generation, accurate monitoring and reliability assessment of environmental factor data are crucial for predicting and optimizing power generation. The power generation efficiency of renewable energy sources such as solar and wind energy is directly affected by environmental conditions. For example, changes in environmental factors such as light intensity, temperature, and wind speed can significantly affect the output power of solar panels and wind turbines.
[0003] Therefore, it is necessary to collect these environmental factor data in real time through multiple sensors and accurately assess their credibility. This can help improve the accuracy of power generation forecasts and optimize the operation and management of renewable energy power generation systems. Traditional data evaluation methods have limitations when processing this type of multi-source heterogeneous data. They may not fully account for the complex fluctuations and correlations of the data, resulting in inaccurate evaluation results. Kernel density estimation, as a flexible nonparametric probability density estimation method, can better adapt to the actual distribution of data and provide more reliable support for the credibility assessment of environmental factor data.
[0004] Traditional data evaluation methods have the following shortcomings when dealing with multi-source heterogeneous data:
[0005] Unable to fully consider the complex fluctuations of data: Traditional data evaluation methods may not be able to adapt to the complex fluctuation characteristics of environmental data, resulting in inaccurate evaluation results.
[0006] Failure to fully consider the correlation of data: Traditional data evaluation methods may ignore the correlation between data of different environmental factors, affecting the comprehensiveness and accuracy of the evaluation.
[0007] The evaluation results are not accurate enough: Due to the above two reasons, the accuracy and reliability of the overall evaluation results of traditional data evaluation methods are insufficient when processing multi-source heterogeneous environmental data. Summary of the Invention
[0008] The present invention aims to solve the problem that traditional data evaluation methods cannot fully consider the complex fluctuations and correlations of data when processing multi-source heterogeneous environmental data, resulting in inaccurate evaluation results. A multi-source heterogeneous environmental data credibility evaluation method and system based on kernel density estimation is proposed. By utilizing kernel density estimation, a flexible non-parametric probability density estimation method, it can better adapt to the actual distribution of data, thereby providing more reliable support for the credibility evaluation of environmental factor data.
[0009] In order to achieve the above purpose, the technical solutions adopted are:
[0010] The present invention provides a method for evaluating the credibility of multi-source heterogeneous environmental data based on kernel density estimation, comprising the following steps:
[0011] Step 1: Collect temperature T, humidity H, and wind speed W environmental data through multiple deployed sensors to ensure the time synchronization and format uniformity of the data;
[0012] Step 2: Perform outlier detection, outlier correction and data filtering on the collected data;
[0013] Step 3: Use the kernel density estimation method to model the known accurate samples to obtain the probability density distribution of accurate data, and calculate the density value of the collected new data under this distribution as the accuracy indicator;
[0014] Step 4: Calculate the variance of each type of environmental data series as a stability indicator;
[0015] Step 5: Use Kullback-Leibler divergence to evaluate data consistency;
[0016] Step 6: Comprehensively consider the accuracy, stability, and consistency indicators, and calculate the credibility of each specific data value for each type of data according to the weight distribution.
[0017] According to the multi-source heterogeneous environment data credibility assessment method based on kernel density estimation of the present invention, further, in step 1, the time synchronization of the data is ensured by using a precise time protocol clock synchronization server, and the time synchronization between sensors is achieved by receiving an external high-precision clock source signal to ensure the consistency of the data timestamp.
[0018] According to the credibility assessment method of multi-source heterogeneous environmental data based on kernel density estimation of the present invention, further, in step 2, a statistical method is used to detect outliers in the collected data, specifically: the mean and standard deviation of the data sequence are calculated, and data that exceeds the range of ±3 times the standard deviation of the mean is regarded as an outlier.
[0019] According to the credibility assessment method of multi-source heterogeneous environmental data based on kernel density estimation of the present invention, further, in step 2, the outlier correction of the collected data is performed using the adjacent data point mean method, specifically: for the detected outlier, take the first two and last two data points, calculate the mean of these four data points, and replace the outlier with the mean; if the first data is abnormal, the first two data are considered to be 0; if the last data is abnormal, the last two data are considered to be 0.
[0020] According to the multi-source heterogeneous environmental data credibility assessment method based on kernel density estimation of the present invention, further, the data filtering in step 2 adopts the median filtering method, specifically: taking the median of the data in the sliding window as the value of the noise point to remove impulse noise.
[0021] According to the method for credibility assessment of multi-source heterogeneous environmental data based on kernel density estimation of the present invention, step 3 specifically includes:
[0022] Use kernel density estimation to model known accurate historical data and obtain the probability density function X is any sensor data sequence;
[0023] The current data point X to be evaluated i , calculate its probability density in the distribution
[0024] The calculated density value Divide by the maximum density value of all data points in the current series The normalized accuracy is obtained. The closer the value is to 1, the more the data point conforms to the normal distribution.
[0025] According to the multi-source heterogeneous environment data credibility assessment method based on kernel density estimation of the present invention, further, step 5 uses Kullback-Leibler divergence to assess data consistency as follows:
[0026] Perform kernel density estimation on temperature, humidity, and wind speed data to obtain univariate distribution
[0027] Calculate the KL divergence between two distributions
[0028] Take the average KL divergence of the same sensor to get the average KL value of the sensor data consistency evaluation X ; The smaller the value, the more consistent the sensor data is with other sensor data.
[0029] According to the multi-source heterogeneous environment data credibility assessment method based on kernel density estimation of the present invention, further, the credibility calculation formula in step 6 is:
[0030]
[0031] Among them, w1, w2, w3 are weight coefficients, satisfying w1+w2+w3=1; Represents the current data point X i The stability of the data series to which it belongs. All data points in the same series share the same stability.
[0032] According to the credibility assessment method of multi-source heterogeneous environmental data based on kernel density estimation of the present invention, further, the weight coefficient is determined by expert scoring method, accuracy weight w1=0.4, stability weight w2=0.3, consistency weight w3=0.3.
[0033] Furthermore, the present invention also provides a system for credibility assessment of multi-source heterogeneous environmental data based on kernel density estimation, which is used to implement the above-mentioned credibility assessment method of multi-source heterogeneous environmental data based on kernel density estimation. The system comprises:
[0034] The data acquisition module is used to collect temperature T, humidity H, and wind speed W environmental data through multiple deployed sensors to ensure the time synchronization and format uniformity of the data;
[0035] The preprocessing module is used to perform outlier detection, outlier correction and data filtering on the collected data;
[0036] The accuracy estimation module is used to model the known accurate samples using the kernel density estimation method to obtain the probability density distribution of accurate data, and calculate the density value of the collected new data under this distribution as the accuracy indicator;
[0037] The stability evaluation module is used to calculate the variance of each type of environmental data series as a stability indicator;
[0038] The consistency assessment module is used to assess data consistency using Kullback-Leibler divergence;
[0039] The credibility calculation module is used to comprehensively consider the accuracy, stability and consistency indicators, and calculate the credibility of each specific data value of each type of data according to the weight distribution.
[0040] The beneficial effects achieved by adopting the above technical solution are:
[0041] 1. Improve the accuracy and reliability of data credibility assessment
[0042] The present invention adopts non-parametric kernel density estimation (KDE) to model data distribution, which can better adapt to the actual distribution of environmental data, effectively capture the true characteristics of complex fluctuations and multi-source heterogeneous data, overcome the limitations of traditional parametric methods, and thus more accurately evaluate the credibility of multi-source heterogeneous environmental data.
[0043] 2. Versatile and scalable
[0044] The present invention is applicable to a variety of environmental monitoring scenarios, especially in the field of renewable energy generation. Its methods and models are universal and can be extended to other similar multi-source data assessment scenarios (such as traffic), showing strong adaptability.
[0045] 3. Optimize the operation and management of new energy power generation systems
[0046] By accurately evaluating the credibility of environmental factor data, it can help improve the accuracy of power generation forecasts, and thus optimize the operation and management of new energy power generation systems, which has practical application value.
[0047] 4. Provide quantitative evaluation indicators
[0048] The present invention not only evaluates the credibility of data, but also provides quantifiable evaluation indicators, making the evaluation results more intuitive, objective and easy to apply.
[0049] 5. Comprehensive consideration of multiple factors
[0050] The data is evaluated from multiple dimensions such as accuracy, stability and consistency, and the credibility value is calculated comprehensively through weight distribution to ensure the comprehensiveness and scientific nature of the evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings of the embodiments of the present invention. The drawings are only used to illustrate some embodiments of the present invention, but not to limit all embodiments of the present invention thereto.
[0052] Figure 1 1 is a flow chart of a method for credibility assessment of multi-source heterogeneous environmental data based on kernel density estimation according to an embodiment of the present invention;
[0053] Figure 2 It is a schematic diagram of the data transmission process according to an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The following will be combined with the accompanying drawings of specific embodiments of the present invention to clearly and completely describe the exemplary embodiments of the present invention. Unless otherwise defined, technical or scientific terms used in the present invention should be given the common meanings understood by people with ordinary skills in the relevant field.
[0055] This embodiment discloses a method for evaluating the credibility of multi-source heterogeneous environmental data based on kernel density estimation to improve the evaluation accuracy and reliability. Figure 1 As shown, the method includes steps such as data collection, preprocessing, accuracy assessment, stability assessment, consistency assessment, credibility calculation and result output. Specifically, it includes steps S1-S7:
[0056] Step S1: Collect temperature T, humidity H, and wind speed W environmental data through multiple deployed sensors to ensure the time synchronization and format uniformity of the data.
[0057] like Figure 2As shown, the time synchronization server is used to synchronize the time between sensors, maintaining the timescale consistency of various sensor values. Sensor data is transmitted to the host computer via a switch. The host computer software runs on a general-purpose computer and is responsible for outlier detection, correction, filtering, and credibility assessment of the data. Step S1 includes substeps S101-S104.
[0058] Step S101: Sensor deployment
[0059] Sensor deployment: Temperature, humidity, and wind speed sensors are rationally deployed in the environmental monitoring area of new energy photovoltaic power plants to ensure that the data is representative.
[0060] For example, the temperature sensor should be installed about 0.5 meters below the photovoltaic panel; the humidity sensor should be placed in the ventilation channel between the photovoltaic panels, about 1.5 meters from the ground, and avoid direct sunlight; the wind speed sensor should be installed in an open, unobstructed area within the photovoltaic power station, at a height of about 10 meters from the ground.
[0061] Step S102: Sensor clock synchronization
[0062] Data synchronization: Time synchronization technology is used to ensure the consistency of data collection time of multiple sensors, thereby ensuring the synchronization of timestamps in reported value messages.
[0063] The PTP (precise time protocol) clock synchronization server serves as the reference clock source for the entire system. It generates a precise reference clock by receiving time signals from high-precision external clock sources (such as GPS and BeiDou). The server periodically sends time synchronization messages to each sensor in the network, following the IEEE 1588 standard. These messages contain precise timestamp information, which the sensors use to adjust their clocks.
[0064] After receiving a time synchronization message from a PTP clock synchronization server, the sensor calculates the time offset based on the timestamp in the message and the time difference between its own clock and uses this offset to adjust its own clock to achieve synchronization with the reference clock. To ensure the accuracy and reliability of synchronization, the sensor must implement a precise timestamp function to accurately record timestamps during data acquisition.
[0065] Step S103: Data collection format
[0066] Data format: The collected data should include information such as sensor identification, environmental factor type, data value and timestamp.
[0067] The data collection example is shown below, taking temperature as an example:
[0068]
[0069]
[0070] Step S104: Data collection content
[0071] Temperature, humidity, and wind speed data are collected through three different sensors, and q data are collected for each type of data. Each sensor collects data once every millisecond, and the temperature data sequence obtained within q milliseconds is T = {T1, T2, ... T q}、Humidity data sequence is H={H1,H2,......H q}, the wind speed data sequence is W={W1,W2,......W q}.
[0072] Step S2: Perform preliminary processing on the collected data, including outlier detection, outlier correction, and data filtering. Step S2 includes sub-steps S201-S203.
[0073] Step S201: Outlier detection
[0074] Statistical methods were used to detect outliers, calculate the mean and standard deviation, and consider data outside the range of mean ± 3 times standard deviation as outliers.
[0075] Take temperature as an example:
[0076] For the temperature data series T, calculate its mean μ T and standard deviation σ T , judge each data point T i Is it in [μ T -3μ T , μ T +3σ T ] range.
[0077] The formula for calculating the mean is:
[0078]
[0079] The formula for calculating the standard deviation is:
[0080]
[0081] Step S202: Outlier correction
[0082] Take the temperature data sequence T as an example:
[0083] The outliers are corrected by the neighboring data point mean method. i, take the first two and last two data points (a total of four data points), calculate the mean of these four data points, and use the mean to replace the outlier T i The correction formula is:
[0084]
[0085] If the first data is abnormal, the first two data are considered to be 0; if the last data is abnormal, the last two data are considered to be 0.
[0086] Since the first data point has no predecessor data (that is, there is no data point in front of it), it is impossible to obtain the "first two data points". Therefore, the "first two data points" are regarded as 0, and only the mean of the last two data points is used to correct the outliers. Therefore, the correction formula degenerates into:
[0087]
[0088] Since the last data point has no subsequent data (i.e., there are no data points behind it), it is impossible to obtain the "last two data points". Therefore, the "last two data points" are regarded as 0, and only the mean of the first two data points is used to correct the outliers. The correction formula degenerates to:
[0089]
[0090] It should be noted that when a piece of data is replaced due to an anomaly, the credibility of other source data at the same time is 0. Therefore, even if the data of other sensors (such as humidity and wind speed) appear to be normal, their credibility is forced to zero to avoid potential erroneous data from contaminating subsequent analysis.
[0091] Step S203: numerical filtering
[0092] The median filter method is used to take the median of the data within a period of time to remove the impulse noise. The formula is:
[0093] T i '=median(T i-N / 2 ,T i-N / 2+1 ,......T i+N / 2 )
[0094] Among them, T i ' is the filtered data, N is the filter window size, usually an odd number, and in this embodiment, N=5.
[0095] It should be noted that when a piece of data is filtered out, the credibility of other source data at the same time is considered to be 0.
[0096] Step S3: Use the kernel density estimation method to model the known accurate samples to obtain the probability density distribution of the accurate data, and calculate the density value of the collected new data under this distribution as the accuracy indicator. Step S3 includes sub-steps S301-S305.
[0097] Step S301: Data modeling
[0098] Use kernel density estimation to model known accurate historical data and obtain the probability density function X is an arbitrary sensor data sequence.
[0099] Taking the temperature sensing data sequence T as an example, h is the bandwidth, which controls the width of the kernel function and thus affects the smoothness of the estimation.
[0100] Probability density of bandwidth h as follows:
[0101]
[0102] Among them, K(·) is the kernel function, and the Gaussian kernel function is selected:
[0103]
[0104] If T1=20 in the temperature sensor data sequence T, the bandwidth h=3, and the total amount of sequence data is 100, then the probability density of T1=20 is:
[0105]
[0106] Step S302: Bandwidth h determination method
[0107] The bandwidth h has a great influence on the results. The preliminary estimation formula is:
[0108]
[0109] Among them, σ is the standard deviation of the data, q is the sample size, and then optimized through cross-validation.
[0110] Cross-validation is performed by dividing the data set into a training set and a validation set (taking the temperature sensor data sequence T as an example, the first 50% of the data is selected as the training set and the last 50% of the data is selected as the validation set). A series of candidate bandwidth values h1, h2, ..., h are randomly selected based on experience. m For each candidate bandwidth h s , use the training set data to perform kernel density estimation and obtain the density estimation function
[0111] Calculate the mean square error on the validation set and select the bandwidth that minimizes the mean square error. The formula is:
[0112]
[0113] Among them, q is the number of samples, 0.5q is the validation set size, The bandwidth is h s The estimated density at time , f(T i ) is the true density.
[0114] Step S303: Calculate single point density value
[0115] The current data point T to be evaluated i , calculate its probability density in the distribution The higher the density value, the higher the T i The more likely it is to belong to a normal data distribution.
[0116] Step S304: Normalize to accuracy
[0117] The calculated density value Divide by the maximum density value of all data points in the current series Get the normalized accuracy:
[0118]
[0119] The closer the value is to 1, the higher the accuracy of the data, which means that the data conforms to the normal distribution more closely.
[0120] Step S305: Sliding window updates historical data sequence and refits To adapt to environmental changes.
[0121] Step S4: Calculate the variance of each type of environmental data sequence as a stability indicator.
[0122] The variance of each type of data series is calculated as a stability indicator. The smaller the value, the more stable the data.
[0123] Take the temperature sensing data sequence T as an example:
[0124] Variance calculation formula:
[0125]
[0126] Among them, μ T is the data mean:
[0127]
[0128] It should be noted that, when calculating a data point T i When the reliability is high, the variance of the entire temperature sensor data sequence T is directly used as its stability input. For example, if the sequence variance σ T 2 =4, then all Ti In the credibility calculation, the stability terms are
[0129] Step S5: Use Kullback-Leibler divergence to evaluate data consistency, or calculate the density value of each sensor data under the joint distribution of kernel density estimation after mixing different sensor data. The closer the value is, the higher the data consistency.
[0130] This example uses the Kullback-Leibler divergence to evaluate consistency:
[0131] First, the KL divergence between the probability distributions of each two data sequences is calculated, and then the average KL divergence value is calculated.
[0132] The probability density function of the temperature sensing data sequence T is and the probability density function of the wind speed sensor data sequence W and They are the kernel density estimation functions of different sensor data. The smaller the KL divergence, the more similar the distribution is and the higher the data consistency is.
[0133] and The KL divergence between is calculated as follows:
[0134]
[0135] Calculated and The KL dispersion between them is then calculated, and the KL dispersion between the probability density functions of the temperature sensor data sequence T and the humidity sensor data sequence H is calculated. The weighted average is used to obtain the consistency evaluation of the temperature sensor data.
[0136] is the humidity probability density function, KL T This is the average value for the consistency evaluation of temperature sensing data. The closer the value is to 1, the higher the consistency of the data.
[0137]
[0138] It is understandable that the consistency indicators of each sensor type, such as the temperature consistency KL T , for all data points T in the temperature sensing data sequence T i Credibility calculation.
[0139] Step S6: Comprehensively consider the accuracy, stability, and consistency indicators, and calculate the credibility of each specific data value of each type of data according to the weight distribution.
[0140] Comprehensively consider the accuracy, stability and consistency indicators of the three types of data: temperature, humidity and wind speed, and calculate the credibility value C of each specific data value of each type of data according to the weight distribution. For example, when calculating the temperature T i of Value formula:
[0141]
[0142] Among them, w1, w2, and w3 are weight coefficients, satisfying w1+w2+w3=1, such as w1=0.4, w2=0.3, and w3=0.3. The weights are adjusted according to the actual scenario and needs, and the expert scoring method can be used to determine the weight of each indicator. Represents the current data point T i The stability of the data sequence, the closer the value is to 1, the more stable the data is. Ensure that stability is positively correlated with credibility (the smaller the fluctuation, the larger the value of this term, and the higher the credibility). The maximum value of the denominator of the stability term max(σ T 2 ) is based on historical data. The maximum value of the denominator of the accuracy term It is the maximum density value based on the data series currently being analyzed (new data).
[0143] Step S7: Result output and application
[0144] Output the credibility value of each specific data value of each type of data, filter, integrate or support decision-making according to needs, such as selecting highly credible data for environmental analysis and decision-making to optimize new energy power generation forecast and system operation.
[0145] Corresponding to the above method, this embodiment further discloses a system for evaluating the credibility of multi-source heterogeneous environmental data based on kernel density estimation, comprising:
[0146] The data acquisition module is used to collect temperature T, humidity H, and wind speed W environmental data through multiple deployed sensors to ensure the time synchronization and format uniformity of the data.
[0147] The preprocessing module is used to perform outlier detection, outlier correction and data filtering on the collected data.
[0148] The accuracy estimation module is used to model known accurate samples using the kernel density estimation method to obtain the probability density distribution of accurate data, and calculate the density value of the collected new data under this distribution as the accuracy indicator.
[0149] The stability evaluation module is used to calculate the variance of each type of environmental data sequence as a stability indicator.
[0150] The consistency assessment module is used to assess data consistency using the Kullback-Leibler divergence.
[0151] The credibility calculation module is used to comprehensively consider the accuracy, stability and consistency indicators, and calculate the credibility of each specific data value of each type of data according to the weight distribution.
[0152] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0153] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0154] The units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation is not considered to be beyond the scope of the present invention.
[0155] Those skilled in the art will appreciate that all or part of the steps in the above method can be performed by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disk. Alternatively, all or part of the steps in the above embodiment can be implemented using one or more integrated circuits. Accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or software functional modules. The present invention is not limited to any specific combination of hardware and software.
[0156] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A method for credibility assessment of multi-source heterogeneous environmental data based on kernel density estimation, characterized in that: The following steps are involved: Step 1: Collect temperature T, humidity H, and wind speed W environmental data through multiple deployed sensors to ensure the time synchronization and format uniformity of the data; Step 2: Perform outlier detection, outlier correction and data filtering on the collected data; Step 3: Use the kernel density estimation method to model the known accurate samples to obtain the probability density distribution of accurate data, and calculate the density value of the collected new data under this distribution as the accuracy indicator; Step 4: Calculate the variance of each type of environmental data series as a stability indicator; Step 5: Use Kullback-Leibler divergence to evaluate data consistency; Step 6: Comprehensively consider the accuracy, stability, and consistency indicators, and calculate the credibility of each specific data value for each type of data according to the weight distribution.
2. The method for evaluating the credibility of multi-source heterogeneous environmental data based on kernel density estimation according to claim 1 is characterized in that: In step 1, the time synchronization of the data is ensured by using a precision time protocol clock synchronization server. By receiving an external high-precision clock source signal, the time between sensors is synchronized to ensure the consistency of the data timestamp.
3. The method for evaluating the credibility of multi-source heterogeneous environmental data based on kernel density estimation according to claim 1 is characterized in that: In step 2, the statistical method is used to detect outliers in the collected data. Specifically, the mean and standard deviation of the data series are calculated, and data outside the range of the mean ± 3 times the standard deviation are considered outliers.
4. The method for evaluating the credibility of multi-source heterogeneous environmental data based on kernel density estimation according to claim 1 is characterized in that: In step 2, the outlier correction of the collected data is performed using the adjacent data point mean method. Specifically, for the detected outlier, take the first two and last two data points, calculate the mean of these four data points, and replace the outlier with the mean; if the first data is abnormal, the first two data are considered to be 0; if the last data is abnormal, the last two data are considered to be 0.
5. The method for evaluating the credibility of multi-source heterogeneous environmental data based on kernel density estimation according to claim 1 is characterized in that: In step 2, the data filtering adopts the median filtering method, specifically: taking the median of the data in the sliding window as the value of the noise point to remove the impulse noise.
6. The method for evaluating the credibility of multi-source heterogeneous environmental data based on kernel density estimation according to claim 1 is characterized in that: Step 3 specifically includes: Use kernel density estimation to model known accurate historical data and obtain the probability density function X is any sensor data sequence; The current data point X to be evaluated i , calculate its probability density in the distribution The calculated density value Divide by the maximum density value of all data points in the current series The normalized accuracy is obtained. The closer the value is to 1, the more the data point conforms to the normal distribution.
7. The method for evaluating the credibility of multi-source heterogeneous environmental data based on kernel density estimation according to claim 6 is characterized in that: Step 5 uses Kullback-Leibler divergence to evaluate data consistency as follows: Perform kernel density estimation on temperature, humidity, and wind speed data to obtain univariate distribution Calculate the KL divergence between two distributions Take the average KL divergence of the same sensor to get the average KL value of the sensor data consistency evaluation X ; The smaller the value, the more consistent the sensor data is with other sensor data.
8. The method for evaluating the credibility of multi-source heterogeneous environmental data based on kernel density estimation according to claim 7 is characterized in that: The calculation formula for credibility in step 6 is: Among them, w1, w2, w3 are weight coefficients, satisfying w1+w2+w3=1; Represents the current data point X i The stability of the data series to which it belongs. All data points in the same series share the same stability.
9. The method for evaluating the credibility of multi-source heterogeneous environmental data based on kernel density estimation according to claim 8, characterized in that: The weight coefficients are determined by an expert scoring method, with the accuracy weight w1 = 0.4, the stability weight w2 = 0.3, and the consistency weight w3 = 0.
3.
10. A credibility assessment system for multi-source heterogeneous environmental data based on kernel density estimation, characterized in that: The system is used to implement the method for evaluating the credibility of multi-source heterogeneous environmental data based on kernel density estimation according to any one of claims 1 to 9, comprising: The data acquisition module is used to collect temperature T, humidity H, and wind speed W environmental data through multiple deployed sensors to ensure the time synchronization and format uniformity of the data; The preprocessing module is used to perform outlier detection, outlier correction and data filtering on the collected data; The accuracy estimation module is used to model the known accurate samples using the kernel density estimation method to obtain the probability density distribution of accurate data, and calculate the density value of the collected new data under this distribution as the accuracy indicator; The stability evaluation module is used to calculate the variance of each type of environmental data series as a stability indicator; The consistency assessment module is used to assess data consistency using Kullback-Leibler divergence; The credibility calculation module is used to comprehensively consider the accuracy, stability and consistency indicators, and calculate the credibility of each specific data value of each type of data according to the weight distribution.
Citation Information
Patent Citations
WiFi (Wireless Fidelity) abnormal link detection method based on Kullback-Leibler divergence
CN110856201A
Quality evaluation method and system for multi-source heterogeneous data
CN111639850A
Fault diagnosis method based on optimized kernel density estimation and JS divergence
CN113051092A
Power quality monitoring reliability evaluation method for intelligent fusion terminal in distribution network area
CN113344406A
Multi-sensor fusion method and system based on multi-dimensional attribute correlation analysis
CN113761705A