Efficient processing method for real-time acquisition and synchronization of multi-source heterogeneous data
Patent Information
- Application Number
- CN202510327850.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art faces challenges such as time stamp synchronization, data missing and noise interference in the real-time acquisition and synchronization of multi-source heterogeneous data, resulting in problems in data processing.
Multi-step processing methods are adopted, including data acquisition and multi-source synchronization, data cleaning and abnormal detection, and multi-source data fusion and unified output. The specific steps include pre-processing of the data through the timestamp alignment algorithm, cleaning and abnormality detection through the similarity analysis algorithm, and generating more expressive and consistent comprehensive data through multi-level data fusion.
Effectively eliminate redundant and abnormal data, ensure that the data is consistent in time and format, reduce data storage and transmission pressure, improve system operation efficiency, and provide more comprehensive information support for subsequent data analysis and decision-making.
Smart Images

Figure CN120234530A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and specifically relates to an efficient method for real-time acquisition and synchronization of multi-source heterogeneous data. Background Art
[0002] With the rapid development of the Internet of Things (IoT) technology, the diversity and heterogeneity of data sources have increased significantly, involving various sources such as intelligent sensors, embedded systems, industrial control devices, environmental monitoring networks, real-time stream data, and historical databases. The acquisition and synchronization of these data face challenges such as timestamp asynchronization, data loss, and noise interference. Therefore, an efficient data acquisition and synchronization method is proposed to ensure that multi-source heterogeneous data can be uniformly processed on the same time axis, providing high-quality data input for subsequent digital twin modeling and simulation. Summary of the Invention
[0003] The purpose of the present invention is to provide an efficient method for real-time acquisition and synchronization of multi-source heterogeneous data to solve the problems in the prior art mentioned in the background art, where the acquisition and synchronization face challenges such as timestamp asynchronization, data loss, and noise interference, resulting in problems during data processing.
[0004] To solve the above technical problems, the technical solution adopted by the present invention is:
[0005] An efficient method for real-time acquisition and synchronization of multi-source heterogeneous data, comprising the following processing steps:
[0006] Step S1, data acquisition and multi-source synchronization: By collecting inputs from multiple different data sources and preprocessing the data through a timestamp alignment algorithm, synchronized data is obtained;
[0007] Step S2, data cleaning and anomaly detection: Perform data cleaning and anomaly data detection on the preprocessed data. By analyzing the spatial and temporal characteristics of different data sources, detect data redundancy, eliminate unnecessary duplicate data, and reduce the pressure of data storage and transmission;
[0008] Step S3, multi-source data fusion and unified output: Standardize the processed data and perform multi-level data fusion on the processed data to generate more expressive and consistent comprehensive data.
[0009] According to the above technical solution, in step S1, the data sources include: data collected by IoT devices, sensor network data, real-time stream data, historical databases, and network log data.
[0010] According to the above technical solution, the preprocessing of the data by the timestamp alignment algorithm is specifically as follows:
[0011] Suppose there are two data sources A and B, and their timestamp sequences are T A ={t A1 , t A2 ,..., t An} and T B ={t B1 , t B2 ,..., t Bm}, and the corresponding data values are X A ={x A1 , x A2 ,..., x An} and X B ={x B1 , x B2 ,..., x Bm}; Timestamp alignment is performed by linear interpolation, and the calculation formula is as follows:
[0012]
[0013] In the formula, x A (t) and x B (t) respectively represent the interpolation results of data sources A and B at time t; The goal of alignment is to generate a new timestamp sequence T so that the values of data A and B can correspond to the same time point, and align the data to a unified time axis; Ensure that data from different sources is processed on the same time axis.
[0014] According to the above technical solution, after timestamp alignment, clock synchronization of each data source is performed through a clock synchronization algorithm, specifically:
[0015] Receive GPS time signal: The device of each data source needs to receive the time signal from the GPS satellite to provide high-precision standard time;
[0016] Local clock deviation correction: By comparing the GPS signal time and the local clock time, calculate the clock deviation ΔT;
[0017] ΔT = T GPS - T local
[0018] Where, T GPS is the time of the GPS signal, and T local is the time of the local clock; Used to calculate the time difference (ΔT) between the local time (T local ) and the GPS time (T gps );
[0019] Synchronization correction: Adjust the local clock to align it with the GPS time, that is, update the local time, specifically:
[0020] T corrected = T local + ΔT
[0021] Continuous calibration: Regularly or continuously receive GPS time signals to dynamically correct clock deviations to cope with clock drift caused by factors such as temperature changes; This GPS-based clock synchronization algorithm can achieve time synchronization between multiple devices.
[0022] According to the above technical solution, in step S2, data cleaning and abnormal data detection are performed on the data. By the similarity analysis algorithm, the redundancy of the data is detected by comparing the spatial and temporal characteristics of the data sources, and duplicate data is removed. Specifically:
[0023] For two data points A = (x1, x2,..., x n ) and B = (y1, y2,..., y n ), their Euclidean distance can be expressed as:
[0024]
[0025] In the formula, if d(A, B) is less than the set threshold ∈, then the two data points are considered similar and redundancy detection can be performed.
[0026] According to the above technical solution, redundancy detection is performed by calculating the cosine similarity. Specifically:
[0027] For two data vectors A and B, the cosine similarity is expressed as:
[0028]
[0029] In the formula, x i and y i are used to calculate the dot product of vectors and the vector norm.
[0030] If the similarity is greater than the set threshold, it is determined that the two data sources have a high similarity and one of them needs to be removed.
[0031] According to the above technical solution, after the similarity calculation is completed, dynamic time warping needs to be performed on the data. By performing non-linear time alignment on the sequences, the minimum time transformation distance is calculated:
[0032]
[0033] In the formula, i and j represent different time points of the sequences. In the dynamic time warping formula DTW(A, B), a i and y i represent the elements in the two time series A and B respectively.
[0034] According to the above technical solution, in step S3, performing multi-level data fusion includes the following steps:
[0035] Step S301, standardize the data; standardize each feature to obtain the standardized data, so that the mean of each feature is 0 and the standard deviation is 1, eliminating the dimensional differences between different features;
[0036] Step S302, extract the features of each data source, including time features, spatial features, and business features;
[0037] Step S303, feature concatenation or weighted combination; form a unified high-dimensional data matrix through the concatenation or weighted combination of the features of each data source; after data fusion, each row represents a sample and each column represents a feature;
[0038] Step S304, spatial dimension data fusion; determine the geographical location weights of each data source, and the weighted fused data x fusion ,x fusion can be expressed as:
[0039]
[0040] where z i is the data of the i-th data source, and w i is its corresponding weight;
[0041] Step S305, perform PCA dimensionality reduction on the multi-source fused data matrix, project the original high-dimensional data into a low-dimensional space, so as to retain the most important data information.
[0042] According to the above technical solution, the specific operation of standardizing the data in step S301 is:
[0043]
[0044] According to the above technical solution, project the original data into the subspace formed by the selected k eigenvectors to complete the dimensionality reduction operation. The projection calculation formula is:
[0045] Y = ZW
[0046] In the formula, Z is the original data matrix after standardization; W is the matrix composed of the selected k eigenvectors; Y is the data matrix after dimensionality reduction; in this way, PCA can reduce the dimension of the data while retaining most of the information of the original data.
[0047] Compared with the prior art, the present invention has the following beneficial effects:
[0048] In the present invention, through data cleaning and anomaly detection, redundant and abnormal data can be effectively removed, ensuring that subsequent analysis and decision-making are based on a high-quality data foundation. Through timestamp alignment and standardization processing, data from different sources can be made consistent in terms of time and format, reducing the analysis difficulties caused by inconsistent data formats. By removing redundant data and optimizing the data structure, unnecessary data storage requirements are reduced, the bandwidth occupancy during data transmission is lowered, and the overall operating efficiency of the system is improved. Multi-level data fusion enables data from different sources to be combined to generate more expressive and consistent comprehensive data, providing more comprehensive information support for subsequent data analysis and decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a flowchart of the data processing of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0051] Embodiment 1
[0052] As Figure 1 shown, an efficient processing method for real-time acquisition and synchronization of multi-source heterogeneous data includes the following processing steps:
[0053] Step S1, data acquisition and multi-source synchronization: By collecting inputs from multiple different data sources and preprocessing the data through a timestamp alignment algorithm, synchronized data is obtained.
[0054] Step S2, data cleaning and anomaly detection: Perform data cleaning and anomaly data detection on the preprocessed data. By analyzing the spatial and temporal characteristics of different data sources, the redundancy of the data is detected, unnecessary duplicate data is removed, and the pressure on data storage and transmission is reduced.
[0055] Step S3, multi-source data fusion and unified output: Perform standardization processing on the processed data and perform multi-level data fusion on the processed data to generate more expressive and consistent comprehensive data.
[0056] In the present invention, through data cleaning and anomaly detection, redundant and abnormal data can be effectively eliminated, ensuring that subsequent analysis and decision-making are based on a high-quality data foundation; through timestamp alignment and standardization processing, data from different sources can be kept consistent in terms of time and format, reducing the analysis difficulties caused by inconsistent data formats. By eliminating redundant data and optimizing the data structure, unnecessary data storage requirements are reduced, the bandwidth occupancy during data transmission is decreased, and the overall operating efficiency of the system is improved. The multi-level data fusion enables data from different sources to be combined to generate more expressive and consistent comprehensive data, providing more comprehensive information support for subsequent data analysis and decision-making.
[0057] The method in the present invention emphasizes real-time data collection and processing, can respond in a timely manner in a rapidly changing environment, supports real-time decision-making and dynamic adjustment, can adapt to various different types of data sources, has good scalability, and can cope with new data sources and new requirements that may appear in the future. The optimized comprehensive data provides a good foundation for intelligent applications such as machine learning and data mining, promoting the realization of intelligent decision-making.
[0058] In summary, in the present invention, through an efficient data processing process, not only the quality and consistency of the data are improved, but also the real-time performance and adaptability of the system are enhanced, providing strong support for the application of multi-source heterogeneous data.
[0059] Embodiment 2
[0060] This embodiment provides a specific implementation manner.
[0061] Step 1, data collection and multi-source synchronization; the system of the present invention can efficiently receive and process inputs from a variety of different data sources, including but not limited to:
[0062] Internet of Things devices: such as intelligent sensors, embedded systems, industrial control devices, which collect various physical quantities such as device status, environmental temperature, humidity, and pressure.
[0063] Sensor networks: such as environmental monitoring sensors, position sensors, flow sensors, etc., which are used to obtain spatial and status information.
[0064] Real-time stream data: such as video surveillance, vehicle trajectory data, etc., which are used to capture dynamic changes.
[0065] Historical databases: including historical data, business data, user data, etc., which serve as a reference basis.
[0066] Network logs: record various types of log data generated during system operation.
[0067] Timestamp alignment and synchronization mechanism Since the timestamps of each data source may not be synchronized, the present invention adopts a unified timestamp alignment algorithm.
[0068] Overview of Timestamp Alignment Algorithm: The purpose of the timestamp alignment algorithm is to unify the timestamps of different data sources so that the data can be analyzed and processed on the same timeline. Alignment can solve the problems of different data collection frequencies and inconsistent collection times. The algorithm processes the data through interpolation, imputation, or time window methods to align the collection time points of different data sources.
[0069] Calculation Formula of Timestamp Alignment Algorithm:
[0070] Suppose there are two data sources A and B, and their timestamp sequences are T A ={t A1 , t A2 ,..., t An} and T B ={t B1 , t B2 ,..., t Bm}, and the corresponding data values are X A ={x A1 , x A2 ,..., x An} and X B ={x B1 , x B2 ,..., x Bm}. The goal of alignment is to generate a new timestamp sequence T so that the values of data A and B can correspond to the same time point.
[0071] Timestamp alignment can be performed by linear interpolation, and the formula is as follows:
[0072]
[0073] where x A (t) and x B (t) represent the interpolation results of data sources A and B at time t, respectively.
[0074] Through this formula, the interpolation results of the two data sources at the aligned timestamp T can be calculated, thereby aligning the data to the same timeline. Ensure that data from different sources is processed on the same timeline.
[0075] Algorithm Processing Flow:
[0076] Data Preprocessing: First, sort the timestamps of different data sources to form their respective time series.
[0077] Generate a Unified Timeline: Generate a unified timestamp sequence T based on the timestamps of all data sources. This sequence can take the smallest time interval as the time step and cover the time range of all data sources.
[0078] Interpolation calculation: Use linear interpolation or other interpolation methods for the data of each data source to align the data to a unified time axis.
[0079] Data merging: Merge the interpolation results into a unified data set.
[0080] The alignment algorithm can ensure that when fusing multi-source data, the data can be aligned to the same time point, making further analysis more accurate.
[0081] GPS signal-based clock synchronization algorithm: The core of the proprietary time synchronization algorithm is to use the globally unified timestamp provided by the GPS signal to synchronize the clocks of each data source. The GPS signal can provide high-precision time information for correcting local clock deviations.
[0082] Algorithm steps:
[0083] Receive GPS time signal: The devices of each data source need to receive the time signal from GPS satellites to provide a high-precision standard time.
[0084] Local clock deviation correction: Calculate the clock deviation ΔT by comparing the GPS signal time and the local clock time.
[0085] ΔT = T GPS -T local
[0086] where T GPS is the time of the GPS signal, and T local is the time of the local clock.
[0087] In the GPS signal-based clock synchronization algorithm, the GPS time is a high-precision globally unified time reference. The time of local devices (such as sensors, computers, etc.) may deviate from the GPS time due to reasons such as clock drift. Through this formula, the deviation value between the local time and the GPS time can be obtained. Once this deviation value is obtained, the local time can be calibrated to be consistent with the GPS time. This calibration is very important for applications that require high-precision time synchronization (such as multi-source data fusion, distributed systems, etc.), and can ensure that the timestamps of different data sources are aligned on the same time axis, thereby improving the accuracy of data processing and analysis. For example, in multi-source data fusion, the timestamps of different data sources may be inconsistent due to their respective clock deviations. By calibrating with the GPS time, these data can be aligned in time for subsequent data analysis and processing.
[0088] Synchronization correction: Adjust the local clock to align it with the GPS time, that is, update the local time to
[0089] T corrected = T local + ΔT
[0090] Continuous calibration: Regularly or continuously receive GPS time signals and dynamically correct clock deviations to cope with clock drift caused by factors such as temperature changes. This GPS-based clock synchronization algorithm can achieve time synchronization between multiple devices and reach sub-millisecond accuracy.
[0091] Delay compensation strategy: When there is delay in some data sources, the present invention uses a prediction model or machine learning algorithm to compensate for the delayed data to reduce the impact of data asynchronization on subsequent processing. For the case where there is data delay in some data sources, a delay compensation algorithm is used for adjustment to reduce the impact of the delay on timestamp alignment.
[0092] (1) Historical delay model: By analyzing the delay patterns of historical data, a delay prediction model is established. When data is synchronized, the delayed data is predicted and compensated in advance.
[0093] (2) Real-time delay adjustment: A fast adaptive algorithm is used to monitor and compensate for data delay in real time, which is suitable for cases where the delay changes greatly.
[0094] Step two: Data cleaning and anomaly detection:
[0095] Redundant data detection and elimination:
[0096] Similarity analysis algorithm: By analyzing the spatial and temporal characteristics of different data sources, the redundancy of data is detected, unnecessary duplicate data is removed, and the pressure of data storage and transmission is reduced.
[0097] The similarity analysis algorithm detects the redundancy of data by comparing the spatial and temporal characteristics of data sources and eliminates duplicate data to reduce the pressure of data storage and transmission. This algorithm is based on the calculation of data similarity and uses methods such as Euclidean distance, cosine similarity, or dynamic time warping (DTW) to judge the similarity of data.
[0098] Calculation of similarity analysis:
[0099] 1. Euclidean distance calculation
[0100] For two data points A = (x1, x2,..., x n ) and B = (y1, y2,..., y n ), its Euclidean distance can be expressed as:
[0101]
[0102] If d(A, B) is less than the set threshold ∈, then the two data points are considered similar and redundancy detection can be performed.
[0103] 2. Cosine similarity calculation:
[0104] For two data vectors A and B, the cosine similarity can be expressed as:
[0105]
[0106] If the similarity is greater than a certain threshold, it is determined that the two data sources have a high degree of similarity.
[0107] 3. Dynamic Time Warping (DTW)
[0108] DTW is used to measure the similarity of two time series. By performing non - linear time alignment on the sequences, the minimum time transformation distance is calculated:
[0109]
[0110] In the formula, i and j represent different time points of the sequences. In the dynamic time warping formula DTW(A, B), a i and y i represent the elements in two time series A and B respectively.
[0111] The role of DTW: Dynamic Time Warping (DTW) is mainly used to compare the similarity of two time series, especially in cases where the time series may have stretching or bending on the time axis. For example, in speech recognition, the speaking speeds of different people may be different, and DTW can help find the best matching path to more accurately compare the similarity of speech signals.
[0112] DTW: Different from Euclidean distance and cosine similarity, DTW does not require two time series to be strictly aligned on the time axis. It matches two time series by finding the path of the minimum cumulative distance and is suitable for dealing with cases where the lengths of time series are different or the time axes are misaligned.
[0113] Through these calculation methods, the redundancy of different data sources can be detected, and duplicate data can be removed to reduce storage and transmission pressure.
[0114] Noise filtering and data smoothing
[0115] Noise detection algorithm: For the random noise in sensor data, the present invention adopts Kalman filtering, weighted moving average or a noise detection algorithm based on signals to remove unstable data.
[0116] Kalman filtering is a recursive algorithm used to estimate the state of a system and remove noise. The update equation for state estimation is:
[0117]
[0118] Wherein: is the estimated value of the current state, K k is the Kalman gain, z k is the observed value, and H is the observation matrix.
[0119] Weighted moving average: The weighted moving average sums the data with weights, giving larger weights to the most recent observed values. The formula is:
[0120]
[0121] Where w i represents the weight of the i-th data point, and x t-i represents the (t - i)-th value in the time series.
[0122] Data smoothing algorithm:
[0123] Low-pass filter: The low-pass filter is used to smooth the signal and eliminate high-frequency noise. Its formula is:
[0124] y t = α · x t + (1 - α) · y t-1
[0125] Where: y t is the smoothed data; x t is the original data; α is the smoothing factor (0 < α < 1)
[0126] Missing data handling:
[0127] Interpolation algorithm: For continuous data, linear interpolation, Lagrange interpolation, or spline interpolation methods are used to fill in the missing data.
[0128] Prediction model based on historical data: When the time interval of the data is large or the data pattern changes rapidly, machine learning models (such as time series prediction models or regression algorithms) are used to predict the possible values of the missing data.
[0129] Missing data handling is one of the important steps in data cleaning. Especially in multi-source data fusion, ensuring the integrity and consistency of the data is crucial for subsequent modeling and simulation. The present invention proposes a comprehensive missing data handling process, including multiple methods. The detailed specific steps are as follows:
[0130] 1. Missing data detection
[0131] Before processing, it is necessary to first detect whether there are missing values in the data, as well as the distribution and proportion of these missing values. The specific methods for missing data detection include:
[0132] Null value detection: Scan for null values (such as empty strings, empty arrays, NaN values) in the data table and mark their positions.
[0133] Outlier detection: Data points outside the outlier range (such as negative temperatures, invalid GPS coordinates) can be treated as missing values.
[0134] Time series breakpoint detection: If data is missing at consecutive time points, breakpoint detection can identify abnormal time intervals in the data and mark them as missing data.
[0135] Strategies for handling missing data
[0136] Depending on the type and proportion of missing data, select appropriate strategies to handle the problem of data missing. The following are the specific methods and steps:
[0137] 1. Linear interpolation:
[0138] For cases where there is less missing data and the data changes smoothly, use linear interpolation:
[0139]
[0140] where t missing is the time point of the missing data.
[0141] 2. Polynomial interpolation
[0142] Suitable for cases where the data changes greatly, and fit with a high-degree polynomial:
[0143] P(t) = a0 + a1t + a2t 2 + … + a n t n
[0144] where a0, a1, ..., a n are polynomial coefficients, obtained by the least squares method.
[0145] 3. Spline interpolation
[0146] Suitable for cases where the data changes complexly, and use a cubic spline function for interpolation:
[0147] S i (t) = a i (t - t i ) 3 + b i (t - t i ) 2 + c t (t - t i ) + d i
[0148] where a i , b i , c i , d i are the coefficients of the spline function.
[0149] Prediction Algorithm for Missing Data
[0150] When the proportion of missing data is large or the time series change of the data is irregular, the accuracy of the interpolation algorithm may not be sufficient. The present invention proposes to use a machine learning algorithm to predict missing data. The prediction algorithm for missing data can use machine learning methods to predict missing values, such as regression analysis or time series prediction models.
[0151] 1. Linear Regression Prediction
[0152] For time series data, linear regression can be used to predict missing values:
[0153]
[0154] where is the predicted missing value, and w0 and w1 are regression coefficients.
[0155] 2. Autoregressive Moving Average (ARMA) Model: The ARMA model is used to predict missing values in a time series, and its formula is:
[0156]
[0157] where: φ is the autoregressive coefficient; θ is the moving average coefficient; ∈ is the error term; through the above algorithms, the problem of missing data can be efficiently processed, and the accuracy of data analysis can be improved.
[0158] Filling with Mean or Median: For cases where the proportion of missing values is small and the data distribution is relatively stable, simple statistical methods can be used for filling:
[0159] Mean Filling: Replace the missing value with the mean of the data, which is applicable when the central tendency of the data is obvious.
[0160] Median Filling: Replace the missing value with the median of the data, which is applicable when there are outliers in the data distribution and can effectively avoid the influence of outliers on the filling result.
[0161] Step Three: Multi-source Data Fusion and Unified Output:
[0162] 1. Time Feature Extraction
[0163] For time series data, extract time-related features, such as timestamps, time intervals, periodic features, etc.
[0164] For example, the time difference between adjacent data points can be calculated, or periodic features in the data can be extracted by methods such as Fourier transform.
[0165] 2. Spatial feature extraction
[0166] Extract spatial features based on the geographical location information of the data source. For example, if the data source is a sensor network, the spatial features may include the coordinates of the sensors, the distances between sensors, the spatial distribution pattern, etc.
[0167] Geographic Information System (GIS) technology can be used to assist in extracting spatial features.
[0168] 3. Business feature extraction
[0169] Extract relevant business features according to specific business requirements. For example, if it is financial data, the business features may include transaction amount, transaction frequency, transaction type, etc.; if it is meteorological data, the business features may include temperature, humidity, air pressure, etc.
[0170] Feature concatenation or weighted combination:
[0171] Feature concatenation: Concatenate the features extracted from each data source in order into a high-dimensional vector. For example, if there are three data sources and m1, m2, and m3 features are extracted respectively, the length of the concatenated high-dimensional vector is m1 + m2 + m3.
[0172] Let A, B, and C be the feature matrices extracted from three data sources. The concatenated matrix D can be expressed as D = [A|B|C] (here the vertical bar represents matrix concatenation).
[0173] Weighted combination: If the weighted combination method is adopted, the weight of each feature needs to be determined. The weights can be determined according to factors such as the importance and reliability of the features.
[0174] Let w1, w2, and w3 be the weight vectors of the features from three data sources. The combined feature vector x can be expressed as x = w1A + w2B + w3C (here it is assumed that A, B, and C are feature vectors in column vector form).
[0175] PCA dimensionality reduction
[0176] Calculate the covariance matrix: For the fused data matrix, first calculate its covariance matrix, then calculate the eigenvalues and eigenvectors, and then perform eigenvalue decomposition on the covariance matrix to obtain the eigenvectors corresponding to the eigenvalues. The eigenvalues and eigenvectors satisfy the equation.
[0177] Select principal components: Sort the eigenvectors according to the magnitudes of the eigenvalues, and select the eigenvectors corresponding to the top eigenvalues (less than the original feature dimension). These eigenvectors form the projection matrix.
[0178] Data projection: Project the original data matrix Z onto the low-dimensional space Y to obtain the dimensionality-reduced data.
[0179] Spatial dimensional data fusion
[0180] 1. Weighted fusion based on spatial weights
[0181] Determine the geographical location weights of each data source. For example, if there are multiple sensors distributed at different locations, the sensors closer to the target area have higher weights.
[0182] The weighted-fused data x fusion can be expressed as:
[0183]
[0184] Let x i be the data of the i-th data source, and w i be its corresponding weight.
[0185] The weight w i can be determined according to methods such as the reciprocal of the spatial distance, Gaussian kernel function, etc.
[0186] 2. Multi-scale data processing
[0187] For data sources at different spatial scales, adopt a hierarchical fusion strategy. For example, there is satellite remote sensing data at a large scale and ground sensor data at a small scale.
[0188] First, perform coarse-grained processing on the large-scale data to extract global features, then perform fine-grained processing on the small-scale data to extract local features, and finally fuse the global features and local features.
[0189] Temporal dimensional data fusion
[0190] 1. Synchronous interpolation and prediction algorithm
[0191] For data with different sampling frequencies, use interpolation techniques to generate data at intermediate time points. For example, if one data source samples once an hour and another data source samples once a minute, linear interpolation, spline interpolation, etc. can be used to interpolate the data sampled once an hour to align it with the data sampled once a minute in terms of time.
[0192] 2. Time series smoothing
[0193] Smooth continuous data using smoothing algorithms (such as moving average, exponential smoothing, etc.) to remove outliers.
[0194] For example, the calculation formula for simple moving average is:
[0195]
[0196] where x i is the original data, and y t is the smoothed data.
[0197] Attribute dimension fusion
[0198] 1. Feature extraction and dimensionality reduction
[0199] Use dimensionality reduction algorithms (such as principal component analysis PCA, linear discriminant analysis LDA, etc.) to reduce the dimensionality of high-dimensional data.
[0200] For example, in PCA, the principal components are determined by calculating the eigenvalues and eigenvectors of the covariance matrix, and the original data is projected onto the principal component space to reduce data redundancy.
[0201] 2. Multi-attribute correlation analysis
[0202] Conduct correlation analysis on the fused data to ensure logical consistency between different attributes.
[0203] Methods such as Pearson correlation coefficient and Spearman rank correlation coefficient can be used to calculate the correlation between attributes. For example, the calculation formula for Pearson correlation coefficient is:
[0204]
[0205] where x i and y i are the data of two attributes, and are their means.
[0206] The principle and steps of principal component analysis (PCA).
[0207] PCA is a statistical dimensionality reduction technique that projects the original high-dimensional data onto a low-dimensional space through orthogonal transformation to find the principal components that can best represent the data variance. The basic idea is to find the main directions (i.e., the directions with the largest variance) in the dataset, transform the original data onto these directions, thereby reducing the dimension while retaining the main features of the data.
[0208] The specific steps are as follows:
[0209] First, standardize the original data to ensure that each feature has the same scale. This step is very important because PCA is sensitive to the scale of the data. After standardization, the mean of the data is zero and the variance is one.
[0210] Data standardization
[0211] 1. Before performing PCA, it is usually necessary to standardize the data. Since different features of the original data may have different dimensions and ranges, standardization can ensure that all features have the same dimension.
[0212] 2. A commonly used standardization method is to set the mean of the data to zero and normalize the standard deviation, that is:
[0213]
[0214] Calculate the covariance matrix
[0215] Calculate the covariance matrix of the standardized data to describe the relationship between features.
[0216]
[0217] where, x i represents the feature vector of the sample, is the mean of the feature.
[0218] Calculate the eigenvalues and eigenvectors of the covariance matrix: Perform eigenvalue decomposition on the covariance matrix to obtain the eigenvalues and corresponding eigenvectors: Let the covariance matrix be C, then the eigenvalues and eigenvectors satisfy:
[0219] Cv = λv
[0220] where: v is the eigenvector; λ is the corresponding eigenvalue; the eigenvalue represents the variance size of the data in the direction of the eigenvector. The sorted eigenvalues and their corresponding eigenvectors are used to determine the direction and importance of the principal components.
[0221] Select the principal components
[0222] Sort according to the size of the eigenvalues and select the eigenvectors corresponding to the top k larger eigenvalues. These eigenvectors are the new coordinate axes after data dimensionality reduction and are called principal components.
[0223] The formula for calculating the percentage of explained variance is:
[0224] where:
[0225]
[0226] where: λ i is the selected i-th eigenvalue; m is the total number of eigenvalues
[0227] Select the top k eigenvectors (principal components) such that the cumulative percentage of variance they explain reaches a preset threshold (e.g., 95%) to ensure the retention of information.
[0228] Transform the data into the principal component space (project it onto a low-dimensional space):
[0229] Project the original data onto the subspace formed by the selected k eigenvectors to complete the dimensionality reduction operation. The projection calculation formula is:
[0230] Y = ZW
[0231] Where: Z is the matrix of the original data after standardization; W is the matrix composed of the selected k eigenvectors; Y is the matrix of the data after dimensionality reduction.
[0232] In this way, PCA can reduce the dimensionality of the data while retaining most of the information of the original data.
[0233] Specific data processing process:
[0234] To apply PCA to the scenario of multi-source data fusion, the specific data processing process can include the following steps:
[0235] Application of the dimensionality-reduced data:
[0236] Use the dimensionality-reduced data for modeling and simulation. The dimensionality-reduced data not only reduces the computational complexity but also reduces the interference of redundant information on the model.
[0237] The dimensionality-reduced data can be further subjected to clustering analysis, classification modeling, or time series prediction.
[0238] Optimization and extension of PCA:
[0239] In practical applications, PCA can be combined with other algorithms to improve the effect of data processing:
[0240] Principal component analysis combined with kernel method (KPCA): Map the data into a high-dimensional space through the kernel trick and then perform PCA to handle non-linear data.
[0241] Principal component analysis combined with independent component analysis (ICA): ICA can extract the non-Gaussian independent components of the data and is suitable for processing data sets with non-linear characteristics.
[0242] Dimensionality reduction methods based on deep learning (such as autoencoders): Achieve non-linear dimensionality reduction of the data through neural network models to obtain better data representation.
[0243] For example: Suppose a system integrates multi-source data from a large general hospital, such as the temperature and humidity of sensors, the working status data of devices, historical fault records, etc. These data have a high dimension and redundancy, and can be processed according to the following steps:
[0244] Data preprocessing: Standardize the data of each sensor.
[0245] Feature extraction: Extract time features from the data of each sensor, such as mean, maximum value, frequency features.
[0246] Multi-source data integration: Concatenate the features from different sources into a high-dimensional matrix.
[0247] PCA dimensionality reduction: Perform PCA on the concatenated matrix and select the first 5 principal components as the new feature input.
[0248] Apply to the prediction model: Use the data after dimensionality reduction for fault prediction.
[0249] By this method, the training time of the model can be significantly reduced, while the prediction accuracy and stability are improved.
[0250] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0251] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An efficient processing method for real-time acquisition and synchronization of multi-source heterogeneous data, characterized by: The processing steps include: Step S1, data collection and multi-source synchronization: by collecting inputs from multiple different data sources and pre-processing the data through a timestamp alignment algorithm, synchronized data is obtained; Step S2, data cleaning and anomaly detection: data cleaning and anomaly detection are performed on the preprocessed data. By analyzing the spatial and temporal characteristics of different data sources, data redundancy is detected, unnecessary duplicate data is eliminated, and the pressure of data storage and transmission is reduced; Step S3, multi-source data fusion and unified output: by standardizing the processed data and performing multi-level data fusion on the processed data, more expressive and consistent comprehensive data is generated.
2. According to claim 1, a highly efficient processing method for real-time acquisition and synchronization of multi-source heterogeneous data, characterized in that: In step S1, the data sources include: IoT device collection data, sensor network data, real-time stream data, historical database and network log data.
3. The method for efficiently processing multi-source heterogeneous data real-time acquisition and synchronization according to claim 2 is characterized in that: The timestamp alignment algorithm preprocesses the data as follows: Assume there are two data sources A and B, and their timestamp sequences are T A ={t A1 , t A2 , ..., t An } and T B ={t B1 , t B2 , ..., t Bm }, the corresponding data value is X A ={x A1 , x A2 , ..., x An } and X B ={x B1 , x B2 , ..., x Bm }; The timestamps are aligned by linear interpolation, and the calculation formula is as follows: In the formula, x A (t) and x B (t) respectively represent the interpolation results of data sources A and B at time t; the goal of alignment is to generate a new timestamp sequence T so that the values of data A and B can correspond to the same time point and align the data to a unified time axis; ensuring that data from different sources are processed on the same time axis.
4. The method for efficiently processing multi-source heterogeneous data real-time acquisition and synchronization according to claim 3 is characterized by: After the timestamp alignment is completed, the clock synchronization of each data source is performed through the clock synchronization algorithm, specifically: Receive GPS time signals: Each data source device needs to receive time signals from GPS satellites to provide high-precision standard time; Local clock deviation correction: By comparing the GPS signal time and the local clock time, the clock deviation ΔT is calculated; ΔT=T GPS -T local Among them, T GPS is the time of the GPS signal, T local is the time of the local clock; used to calculate the local time (T local ) and GPS time (T gps ) between the time difference (ΔT); Synchronization correction: adjust the local clock to align it with the GPS time, that is, update the local time, specifically: T corrected =T local +ΔT Continuous calibration: Regularly or continuously receive GPS time signals and dynamically correct clock deviations to account for clock drift caused by factors such as temperature changes; this GPS-based clock synchronization algorithm can achieve time synchronization between multiple devices.
5. The method for efficiently processing multi-source heterogeneous data real-time acquisition and synchronization according to claim 4 is characterized in that: In step S2, data cleaning and abnormal data detection are performed on the data by comparing the spatial and temporal characteristics of the data sources through a similarity analysis algorithm, detecting data redundancy and removing duplicate data, specifically: For two data points A = (x1, x2, ..., x n ) and B = (y1, y2, ..., y n ), its Euclidean distance can be expressed as: In the formula, if d(A, B) is less than the set threshold ∈, the two data points are considered to be similar and redundancy detection can be performed.
6. The method for efficiently processing multi-source heterogeneous data real-time acquisition and synchronization according to claim 5, characterized in that: Redundancy detection is performed through cosine similarity calculation, specifically: For two data vectors A and B, the cosine similarity is expressed as: In the formula, x i and i Used to calculate vector dot products and vector magnitudes. If the similarity is greater than the set threshold, it is determined that the two data sources have a high similarity and one of them needs to be removed.
7. The method for efficiently processing multi-source heterogeneous data real-time acquisition and synchronization according to claim 6 is characterized in that: After completing the similarity calculation, the data needs to be dynamically time-warped by performing nonlinear time alignment on the sequence to calculate the minimum time transformation distance: In the formula, i and j represent different time points of the sequence. In the dynamic time warping formula DTW(A,B), a i and i Represent the elements in two time series A and B respectively.
8. The method for efficiently processing multi-source heterogeneous data real-time acquisition and synchronization according to claim 7, characterized in that: In step S3, multi-level data fusion includes the following steps: Step S301, standardizing data; standardizing each feature to obtain standardized data, so that the mean of each feature is 0 and the standard deviation is 1, eliminating the dimensional differences between different features; Step S302, extracting features of each data source, including time features, spatial features, and business features; Step S303, feature concatenation or weighted combination forms a unified high-dimensional data matrix by concatenating or weighted combining the features of each data source; after data fusion, each row represents a sample, and each column represents a feature; Step S304: spatial dimension data fusion; determine the geographical location weight of each data source, and weight the fused data x fusion , x fusion It can be expressed as: Among them, z i is the data of the ith data source, w i is its corresponding weight; Step S305, PCA dimensionality reduction is performed on the data matrix after multi-source fusion, and the original high-dimensional data is projected into a low-dimensional space, thereby retaining the most important data information.
9. The method for efficiently processing multi-source heterogeneous data real-time acquisition and synchronization according to claim 8, characterized in that: The standardized data in step S301 is specifically:
10. The method for efficiently processing multi-source heterogeneous data real-time acquisition and synchronization according to claim 9, characterized in that: The original data is projected into the subspace formed by the selected k eigenvectors to complete the dimensionality reduction operation. The projection calculation formula is: Y=ZW Where Z is the normalized original data matrix; W is the matrix consisting of the selected k eigenvectors; Y is the reduced-dimensional data matrix; in this way, PCA can reduce the dimensionality of the data while retaining most of the information of the original data.
Citation Information
Cited By
Data processing method based on Internet of Things multi-source information
CN120492815A
Intelligent fusion and cleaning method for multi-source heterogeneous data
CN120744330A
A Smart Fusion and Cleaning Method for Multi-Source Heterogeneous Data
CN120744330B
Remote sensing satellite network security situation awareness method and system with dual prevention mechanisms
CN121151901A
Multi-sensor monitoring system and method based on time alignment
CN121667653A