Time sequence data truth value discovery method based on kernel density estimation

By adopting a kernel density estimation method in time series data processing, combining Gaussian kernel function and Kalman filtering technology, the problem of difficulty in capturing nonlinear patterns and evaluating the credibility of data sources is solved in the existing technology, and more accurate prediction and estimation of missing values ​​of time series data is achieved.

CN120180266APending Publication Date: 2025-06-20SOUTHEAST UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510251171.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When processing time series data, it is difficult to effectively capture nonlinear patterns and evaluate the credibility of data sources, resulting in insufficient prediction of missing data.

Method used

A method based on kernel density estimation is adopted, combining Gaussian kernel function and Kalman filtering technology to deal with missing data from the perspectives of time attributes and confidence attributes. This method optimizes parameters and estimates the most likely truth value through the kernel density estimation phase and the truth value estimation phase.

Benefits of technology

This method not only takes into account the temporal dynamics of the data, but also evaluates the credibility of the data source, and can more accurately predict and estimate missing values ​​in the time series, improving the accuracy of truth value discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180266A_ABST
    Figure CN120180266A_ABST
Patent Text Reader

Abstract

The invention discloses a time sequence data truth value discovery method based on kernel density estimation, which comprises the following steps of: initializing necessary parameters such as bandwidth and credibility weight for each object and time point; a kernel density estimation stage: performing kernel density estimation on data of each data source at each time point by using a Gaussian kernel function, updating kernel density estimation by using an exponentially weighted recursive formula so as to reflect time dependence of time sequence data, calculating a joint probability of given known data for a missing value, and calculating a kernel density estimation result; kernel density estimation is used to assess the likelihood of missing data. Kernel estimation of multi-source data is carried out, and the credibility of different source data sets is considered; in the truth value estimation stage, an expectation maximization algorithm and a kernel Kalman filtering technology are applied, parameters are optimized, and the most probable truth value is estimated. According to the method, the kernel function and the Kalman filtering algorithm are used, the problem of time sequence data truth value discovery is solved, the truth value discovery effect is optimized, and the method has wide application value and use prospects in the field of big data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a true value discovery method for time series data based on kernel density estimation, belonging to the technical fields of big data, machine learning, and true value discovery. Background Art

[0002] In recent years, with the development of Internet of Things technology, time series data, as one of the main data types of the Internet of Things, has been used more and more widely. For example, weather data, carbon dioxide concentration data, etc. Currently, a large amount of time series data is collected through sensors. However, on the one hand, the data generation methods of some sensors have low credibility. On the other hand, the sensors themselves have inaccuracies. Therefore, in the real world, the time series reported by sensors always deviates from the real data. Directly using inaccurate data for time series data analysis may lead to incorrect decisions, restricting the development of related fields. Therefore, to solve the above problems, it is considered to adopt a true value discovery method to recover real data from multiple inaccurate information.

[0003] From the perspective of time attributes, the observed values change continuously over time. According to the central limit theorem, the sum of independent and identically distributed random variables will approximately follow a normal distribution when the sample size is large enough, even if the original variables themselves are not normally distributed. Therefore, in the process of true value discovery, it is necessary to combine the central limit theorem to mine the information of time attributes in the time series.

[0004] Many researchers have proposed true value discovery methods for time series data in different scenarios, but these methods all have certain limitations. In the prior art, an online true value discovery algorithm has been proposed. Using the SARIMA model to model time series data can accurately estimate the true value of time series data in the environment of data flow, but its ability to capture non-linear patterns is weak. A multi-source dynamic data fusion algorithm combining linear regression, AHP (Analytic Hierarchy Process) and dynamic data sample similarity has been proposed, but the AHP depends on subjective assignment, may be affected by personal preferences, and has a large computational amount when dealing with a large amount of data. A new synthetic time series data fusion prediction model has been proposed, which uses the fusion concept of original data and synthetic data to generate metadata as enhanced data, improving the model learning ability and prediction results, but the quality of the augmented data cannot be judged. A method combining batch processing sliding window and attention turning network has been proposed for true value discovery of low-correlation multi-source time series data, but it has high requirements for computing resources.

[0005] A kernel function is a method in machine learning that can map data into a higher-dimensional feature space, making data that is linearly inseparable in the original space linearly separable in the high-dimensional space. This mapping can simplify complex data structures and thus improve the performance of the model. The truth discovery method proposed in this invention is based on the kernel function and considers the problem of missing data from two perspectives: the time attribute and the credibility attribute. In terms of the time attribute, the Gaussian kernel function is used to estimate the probability density function of time series data, mapping the true value from the true value / vector to the function; in terms of the credibility attribute, the Gaussian kernel is combined with the credibility of each source to obtain the value of the missing data. Summary of the Invention

[0006] To overcome the deficiencies of existing truth discovery techniques for time series data, the present invention provides a truth discovery method for time series data based on kernel density estimation for time series data from different data sources, considering the time attribute and the credibility attribute, and applying the Gaussian kernel function and the Kalman filtering technique.

[0007] To achieve the above object, the technical solution adopted by the present invention is: a truth discovery method for time series data based on kernel density estimation: the method includes the following steps:

[0008] Including the following stages:

[0009] A. Preprocessing stage: Initialize necessary parameters for each object and time point, such as the bandwidth and the discount factor;

[0010] B. Kernel density estimation stage: Use the Gaussian kernel function to perform kernel density estimation on the data of each data source at each time point, and update the kernel density estimation using the exponentially weighted recursive formula to reflect the time dependence of the time series data. For missing values, calculate the joint probability of the given known data and use kernel density estimation to evaluate the possibility of the missing data;

[0011] C. Truth value estimation stage: Apply the expectation maximization algorithm and the kernel Kalman filtering technique to optimize the parameters and estimate the most likely truth value. In addition, define an objective function to improve the accuracy of all missing data and minimize the loss. Represent the bandwidth, the weight of the time series data, and the reliability of the source through a parameter matrix, and transform the problem into an unconstrained optimization problem.

[0012] Among them, the steps of the preprocessing stage are as follows:

[0013] A1. Set the initial bandwidth;

[0014] A2. Allocate the initial discount factor.

[0015] Among them, the steps of the kernel density estimation stage are as follows:

[0016] B1. At the first time point of the time series, use all available data points to calculate the initial kernel density estimate, which is done by applying the kernel function to each data point and accumulating the results;

[0017] B2. For each subsequent time point in the time series, use a recursive formula to update the kernel density estimate, which typically involves taking into account new observations and giving more recent observations higher weights, which is achieved through a discount factor;

[0018] B3. Use an exponentially weighted mechanism to combine the previous kernel density estimate with the new observations;

[0019] B4. If there are missing values at a certain time point, the kernel density estimate can estimate the probability distribution of the missing values by considering the known data points. For the missing data points, the joint probability distribution given the known data points can be calculated;

[0020] B5. For each object, consider the set of sources that provide values and use credibility weights to map the true values and claims to a function.

[0021] Among them, the steps of the true value estimation stage are as follows:

[0022] C1. Apply the EM algorithm to optimize the bandwidth and discount factor. The EM algorithm finds the parameter values that maximize the data likelihood by iterating two steps: the expectation (E) step and the maximization (M) step;

[0023] C2. Combine the kernel Kalman filtering technique to estimate and predict missing values. The kernel Kalman filter uses the kernel trick to handle non-linear and non-Gaussian cases and updates the state estimate by minimizing the error between the observations and the estimates;

[0024] C3. Use the optimized parameters and the kernel density estimation results to estimate the most likely true value of each object at each time point

[0025] The specific steps of the preprocessing stage are as follows:

[0026] A1. According to the existing references, set the initial bandwidth to

[0027] A2. The initialization of the discount factor specifically includes the following steps:

[0028] A2.1 According to the Karush-Kuhn-Tucker conditions, the derivative of the discount factor of object i in data source j at time t is calculated as where L is the Lagrangian function, α i , β i , λ iis the coefficient of the Lagrangian function, and l is the objective function;

[0029] A2.2 According to

[0030] it can be calculated that wherein, is the kernel function, v i is the value of the i-th data source, is at time t and scenario s j under the observation value of the i-th data source. is the kernel estimator of the given data;

[0032] A2.3 The derivative of can be calculated as Furthermore, it can be obtained that wherein, is the weight coefficient of the i-th data source at the j-th moment, is the set of numerical indexes of the i-th data source at time t;

[0034] A2.4 Let be 0, and then the optimal estimated value of can be obtained of,

[0035] A2.5 Let then The optimal estimated value of is calculated through for calculation,

[0036] A2.6 The calculation method of is as follows: Based on this, the discount factor can be calculated.

[0037] The specific steps in the kernel density estimation stage include the following:

[0038] B1. Assume that the collected data follows a Gaussian distribution at time point t, forming a Gaussian process. To map the true value from the true value / vector to the function, the specific steps are as follows:

[0039] B1.1 Use the Gaussian kernel to estimate the probability density function as follows:

[0040] B1.2 For the observation sample T, on the data source s j the object o i in, the value is v i the kernel estimator of is wherein, K() is the kernel function and h is the bandwidth.

[0041] B2. A weighted scheme is adopted in the kernel estimator to estimate the time-varying density through smoothing and filtering. The transformation of the above equation is as follows:

[0042] B3. By using exponential weighting, the above formula is expressed in a recursive form as:

[0043] B4. To sample the missing values of the objects from the data source s j we assume that is known and is the missing value. The joint probability of the missing value is expressed as Mapping the quantity to the function specifically includes the following steps:

[0044] B5. Depending on the actual scenario, the trust in the source is different for different objects. Transfer the true value and the statement from the real value / to B5.1

[0045] denotes the credible weights of all sources on the object o and i the mapping of the i-th object in the source s

[0046] B5.2 is defined as j in

[0047] B5.3 To simulate the uncertainty of the true value, the function that uses the Gaussian kernel to map the true and declared true values to the i-th object is where

[0048] The truth discovery phase includes the following steps:

[0049] C1. Maximize the likelihood estimates of the bandwidth j and the discount factor of the i-th object in the source s specifically including the following steps:

[0050] C1.1 Define its log-likelihood function as

[0051] C1.2 The log-likelihood functions of all sources are calculated as follows:

[0052] C1.3 To discover the true value among various data sources, the loss function is defined to minimize the gap between the predicted value and the true value, that is

[0053] C2. Combine the kernel Kalman filtering technique to estimate and predict missing values. Specifically, it includes the following steps:

[0054] C2.1 Define the data set as where y t is the observer, x t is the state value, the previous state value at time t in this data source.

[0055] C2.2 Define the respective feature matrices as

[0056] Φ∶=[φ(y1),……,φ(y n )]. The covariance operator is denoted as

[0057] C2.3 The transfer operator that propagates the posterior belief state at time t to the prior belief state at time t + 1 can be described as

[0058] as

[0059] C3. The representation of the posterior mean is as follows:

[0060] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for constructing a large model for chronic disease health management with enhanced uncertainty knowledge described above.

[0061] A computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, they implement the method for constructing a large model for chronic disease health management with enhanced uncertainty knowledge described above.

[0062] A method for discovering the true value of time series data based on kernel density estimation provided by the present invention has the following beneficial effects compared with the prior art:

[0063] (1) The present invention not only considers the time dynamics of the data but also the credibility of the data source;

[0064] (2) The present invention combines the kernel method and the Kalman filtering technique to provide a powerful tool for estimating and predicting missing values in time series. When dealing with the problem of discovering the true value of time series data, the present invention not only considers the time dynamics of the data but also the credibility of the data source. This means that when dealing with time series data, not only the changing trend of the data over time is concerned, but also the reliability of different data sources is evaluated, and the non-linear and non-Gaussian situations and state estimation optimization are taken into account. The combination of the two makes the prediction of missing data more accurate. Description of the Drawings

[0065] Figure 1 It is the architecture diagram of the method of the present invention.

[0066] Figure 2 It is the flow schematic diagram of the specific implementation algorithm of the method of the present invention.

[0067] Figure 3 It is the structure schematic diagram of the data set used in the truth discovery problem. Detailed implementation manners

[0068] The present invention will be further clarified below in conjunction with the accompanying drawings and specific implementation cases. It should be understood that these cases are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification by those skilled in the art fall within the scope defined by the appended claims of this application.

[0069] Example: As Figure 1 shown is a method for truth discovery of time series data based on kernel density estimation, which comprehensively considers the time dynamics of data and the credibility of data sources. In the process of obtaining the truth value, kernel methods and Kalman filtering techniques are used to optimize the effect of truth discovery.

[0070] Assume n = |O|. For each object The relevant information is collected from the source set S = {s1, s2 ……, s m}. Then the method for truth discovery of time series data based on kernel density estimation includes the following stages:

[0071] A. Preprocessing stage: A1. Initial bandwidth Set to A2. According to the Karush-Kuhn-Tucker conditions, the derivative of the discount factor of object i at time t in data source j is calculated as In this formula, according to can be calculated as

[0072] Similarly, the derivative of can be calculated as Let be 0, then the optimal estimated value of can be obtained. Let then the optimal estimated value of is calculated through The calculation method of Based on this, the discount factor can be calculated;

[0076] B. Kernel density estimation stage: B1. Use a Gaussian kernel to estimate the probability density function as follows: For the observed sample T, the object o j on the data source s i where the value is v i the kernel estimator is

[0078] B2. When using a weighted scheme in the kernel estimator to estimate the time-varying density. The transformation of the above equation is as follows: B3. By using exponential weighting, the above formula is expressed in a recursive form as: B4. The joint probability of missing values is expressed as B5. The function that uses a Gaussian kernel to map the true and claimed true values to the i-th object is where

[0079] C. Ground truth discovery stage: Define the log-likelihood function of the bandwidth and discount factor as The log-likelihood function of all sources is calculated as follows:

[0081] Define the loss function Let; C2. Define the data set as where y t is the observer, x t is the state value, the previous state value at time t in this data source. Define the respective feature matrices as Υ x ∶=[φ(x1),……,φ(x n )], Φ∶=[φ(y1),……,φ(y n )]. The covariance operator is denoted as The transition operator that propagates the posterior belief state at time t to the prior belief state at time t+1 can be described as

[0084] C3. The representation of the posterior mean is as follows:

[0086] It should be noted that the above embodiments are not intended to limit the protection scope of the present invention. Any equivalent transformation or substitution made on the basis of the above technical solutions falls within the protection scope of the claims of the present invention.

Claims

1. A method for discovering true values ​​of time series data based on kernel density estimation, characterized by: The method The following steps are involved: The following phases are included: A. Preprocessing phase: Initialize necessary parameters for each object and time point, B. Kernel density estimation stage: Use the Gaussian kernel function to perform kernel density estimation on the data of each data source at each time point, and use the exponentially weighted recursive formula to update the kernel density estimation to reflect the time dependence of time series data. For missing values, calculate the joint probability of given known data, and use kernel density estimation to evaluate the possibility of missing data; C. True value estimation stage: Apply the expectation maximization algorithm and kernel Kalman filtering technology to optimize parameters and estimate the most likely true value.

2. The method for discovering true value of time series data based on kernel density estimation according to claim 1, characterized in that: The specific steps of the pre-processing stage are as follows: A1. Set the initial bandwidth. A2. Assign an initial discount factor.

3. The method for discovering true value of time series data based on kernel density estimation according to claim 1, characterized in that: The specific steps of the kernel density estimation stage are as follows: B1. At the first time point of the time series, use all available data points to calculate the initial kernel density estimate. This is done by applying the kernel function to each data point and accumulating the results. B2. For each subsequent time point in the time series, use a recursive formula to update the kernel density estimate. This typically involves taking new observations into account and giving more recent observations a higher weight, achieved through a discount factor. B3, using an exponential weighting mechanism to combine the previous kernel density estimate with the new observations; B4. If there are missing values ​​at a certain time point, kernel density estimation estimates the probability distribution of the missing values ​​by considering the known data points. For the missing data points, the joint probability distribution given the known data points is calculated. B5. For each object, consider the set of all sources that provide values ​​and use credibility weights to map true values ​​and claims to functions.

4. The method for discovering true value of time series data based on kernel density estimation according to claim 3 is characterized in that: The specific steps of the true value estimation stage are as follows: C1. Use the EM algorithm to optimize the bandwidth and discount factor. The EM algorithm finds the parameter value that maximizes the data likelihood by iterating two steps: the expectation (E) step and the maximization (M) step. C2. Combine the kernel Kalman filter technique to estimate and predict missing values. The kernel Kalman filter uses kernel techniques to handle nonlinear and non-Gaussian situations and updates the state estimate by minimizing the error between observation and estimation. C3. Use the optimized parameters and kernel density estimation results to estimate the most likely true value of each object at each time point.

5. The method for discovering true value of time series data based on kernel density estimation according to claim 2, characterized in that: Step A2. Discount factor initialization specifically includes the following steps: A2.1 According to the Karush-Kuhn-Tucker condition, the discount factor of object i in data source j at time t is The derivative of is calculated as Among them, L is the Lagrangian function, α i ,β i ,λ i is the coefficient of the Lagrangian function, l is the objective function; A2.2 Based on Calculate in, is the kernel function, v i is the value of the ith data source, is at time t and scene s j Next, the observation value of the i-th data source, is the kernel estimator for the given data; A2.3 The derivative of can be calculated as It can be concluded that in, is the weight coefficient of the i-th data source at the j-th moment, is the set of numerical indexes of the ith data source at time t; A2.4 Order If is 0, we can get The best estimate of A2.5 Order but The best estimate of Perform calculations, A2.6 The calculation method is as follows: From this the discount factor can be calculated.

6. The method for finding true value of time series data based on kernel density estimation according to claim 1, characterized in that: The kernel density estimation stage is as follows: B1. Assume that the collected data follows a Gaussian distribution at time point t, forming a Gaussian process. In order to map the true value from the true value / vector to the function, the following steps are specifically included: B1.1 uses the Gaussian kernel to estimate the probability density function as follows: B1.2 For the observation sample T, the data source s j Object o i The value is v i The kernel estimator is Among them, K() is the kernel function, h is the bandwidth, B2. Through smoothing and filtering, a weighting scheme is adopted in the kernel estimator to estimate the time-varying density. The transformation of the above equation is as follows: B3. By using exponential weighting, the above formula can be expressed recursively as follows: B4. To obtain the data from the data source j Sampling Objects The missing values ​​of is known, is a missing value, and the joint probability of missing values ​​is expressed as B5. According to the actual scenario, the trust of the source is different for different objects. Mapping the truth value and the statement from the real value / vector to the function specifically includes the following steps: B5.1 Represented as object o i The trust weight of all sources above, B5.2 Source s j The mapping for the i-th object in is defined as B5.3 To model the uncertainty of the true value, a Gaussian kernel is used to map the true and declared true values ​​to the function of the i-th object: in 7. The method for discovering true value of time series data based on kernel density estimation according to claim 1, characterized in that: The truth discovery phase includes the following steps: C1. Make the source j The bandwidth of the i-th object in and discount factor The likelihood estimate of is maximized, which includes the following steps: C1.1 defines its log-likelihood function as The log-likelihood function of all sources in C1.2 is calculated as follows: C1.3 In order to find the true value in various data sources, the loss function is defined to minimize the gap between the predicted value and the true value, that is, C2. Combine the kernel Kalman filter technology to estimate and predict missing values, including the following steps: C2.1 defines the dataset as where y t For the observer, x t is the status value, The previous state value at time t in the data source, C2.2 defines their respective feature matrices as Φ∶=[φ(y1),……,φ(y n )], the covariance operator is denoted as C2.3 describes the transfer operator that propagates the posterior belief state at time t to the prior belief state at time t+1 as C3. The posterior mean is expressed as follows:

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements a method for constructing a large model of chronic disease health management enhanced by uncertainty knowledge as described in any one of claims 1 to 7 above.

9. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instruction is executed by the processor, a method for constructing a large model of chronic disease health management enhanced by uncertainty knowledge is implemented as described in any one of claims 1-7.

Citation Information

Cited By

  • Adaptive localization method based on ensemble paleoclimate data assimilation framework

    CN121996893A