Multi-data source automatic integration method and system based on deep learning
Through a deep learning-based method, the DTW algorithm is used to quantify the trend consistency of multi-source data and perform weighted fusion, which solves the problem of insufficient data accuracy and reliability in multi-source data integration, and achieves high-quality data fusion effect.
Patent Information
- Application Number
- CN202511080128.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-08-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The multi-source data integration method in the prior art fails to effectively quantify the trend consistency and delay amplitude between data sources, resulting in insufficient accuracy and reliability of data fusion, and insufficient consideration of spatiotemporal characteristics and sensor reliability differences, affecting the data integration effect.
A deep learning-based method is adopted to calculate the shortest path distance and matching path smoothness through the DTW algorithm, quantify the trend consistency between data sequences, and weighted fusion based on confidence to optimize the data fusion process.
The quality and reliability of multi-source data integration are significantly improved, and by quantifying the shape similarity and time alignment quality between data sequences, it provides more reliable confidence indicators and improves the accuracy and robustness of data fusion.
Smart Images

Figure CN120561883A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and more particularly to a method and system for automatically integrating multiple data sources based on deep learning. Background Art
[0002] With the advancement of automation and intelligent industrial production, numerous sensors are installed on various equipment in industrial production workshops to monitor their operating status in real time, such as temperature, pressure, vibration, and other parameters. The data collected by these sensors is stored in different databases, creating a multi-data source landscape. However, to optimize production processes, predict faults, and manage maintenance, this dispersed data needs to be integrated for comprehensive analysis and decision support. Data integration not only provides a more comprehensive view of equipment operations but also helps identify potential production issues, improving production efficiency and product quality.
[0003] During the data fusion process, sensors in different locations may collect inconsistent data due to environmental differences, measurement errors, or equipment failures. These conflicting values make it difficult to determine which data is closer to the true value during data fusion, thus affecting the accuracy and reliability of the fusion results. Furthermore, due to factors such as sensor placement, measurement frequency, and data transmission delay, time alignment of different data sources becomes a challenge, further complicating data fusion.
[0004] Furthermore, existing approaches to multi-source data integration still have shortcomings. For example, there is a lack of effective means to quantify trend consistency and data latency across different data sources, making it impossible to accurately assess the credibility of each data source. Furthermore, existing methods often employ simple averaging or voting mechanisms for data fusion, failing to fully consider the spatiotemporal characteristics of the data and differences in sensor reliability. This makes it difficult to achieve high-quality data fusion, impacting the overall effectiveness of data integration and the value of subsequent applications. Summary of the Invention
[0005] To address the problems of low data fusion quality caused by the inability to quantify trend consistency and delay amplitude between data sources, making it difficult to accurately assess credibility, and the simple fusion method that does not fully consider spatiotemporal characteristics and reliability differences, the present invention provides solutions in the following aspects.
[0006] In a first aspect, a method for automatic integration of multiple data sources based on deep learning includes: obtaining data sequences of the same parameter of each sensor in an industrial production workshop within a preset time period, and preprocessing the data sequences; calculating the shortest path distance between the preprocessed data sequences corresponding to any two combinations of sensors and the smoothness of the matching paths, and then analyzing the trend consistency between the two preprocessed data sequences; calculating the confidence level of each sensor data based on the trend consistency, and fusing the multi-sensor data sequences to obtain a fused parameter data sequence, thereby achieving high-quality multi-source data integration.
[0007] Preferably, the trend consistency is calculated by: Use the DTW algorithm to calculate the shortest path distance between the data sequences corresponding to any two combinations of sensors, and divide the shortest path distance by the total number of data to obtain the normalized shortest distance; Calculate the time difference of each matching point pair in any two-to-two sensor data collection sequence, calculate the difference of the time differences of adjacent matching point pairs and sum them up to obtain the smoothness of the matching path; The product of the normalized shortest distance and the smoothness of the matching path is taken as the trend consistency between any two combinations of sensor data series.
[0008] By quantifying the shape similarity and time alignment quality between data sequences, the trend consistency of sensor data sequences is accurately evaluated, providing a reliable confidence indicator for subsequent data fusion, thereby significantly improving the quality and reliability of multi-source data integration.
[0009] Preferably, another calculation method of the trend consistency includes: Use the DTW algorithm to calculate the shortest path distance between the data sequences corresponding to any two combinations of sensors, and divide the shortest path distance by the total number of data to obtain the normalized shortest distance; The standard deviation of the slope between each pair of adjacent matching points on the matching path is calculated, and the product of the normalized shortest distance and the standard deviation of the slope is used as the trend consistency between any two combinations of sensor data series.
[0010] By comprehensively considering the shape similarity of data sequences and the global fluctuation characteristics of matching paths, it can not only adapt to data sequences of different lengths and frequencies, but also effectively identify and quantify subtle differences between data sequences, thereby providing a more reliable basis for confidence calculation in the data fusion process, and thus significantly improving the accuracy and robustness of multi-source data integration.
[0011] Preferably, the step of obtaining the confidence level includes: Taking any sensor as the target sensor, traverse and calculate the ratio between the total number of matching point pairs in the data sequences collected by the target sensor and other sensors and the total number of data contained in the data sequences, use the negative exponential function to perform exponential mapping on the comparison values, and use the mapping result as the matching complexity; The matching complexity is used as the weight of the trend consistency between the data sequences of the target sensor and other sensors, and the weighted sum and normalization are performed to obtain the confidence of the target sensor.
[0012] By quantifying the matching complexity between the target sensor and other sensors and incorporating it as a weight in the trend consistency assessment, the credibility of each sensor's data can be more accurately reflected. By comprehensively considering the similarity between data sequences and the complexity of temporal alignment, a reasonable confidence level is assigned to each sensor, improving the ability to identify and eliminate anomalous data during the data fusion process, and further enhancing the quality and reliability of multi-source data integration.
[0013] Preferably, the step of obtaining the fused parameter data sequence includes: Taking any sensor as the target sensor, the data sequence collected by the target sensor is weightedly fused based on the confidence of the target sensor, and the fused data sequence is divided by the comprehensive confidence of all sensors to obtain the fused data sequence.
[0014] By introducing confidence levels to weightedly fuse the data sequences collected by each sensor, the weights in the fusion process can be dynamically adjusted based on the reliability of the sensor data. This not only fully utilizes information from high-quality data sources, but also effectively reduces the interference of low-quality data sources on the fusion results, significantly improving the accuracy and reliability of the fused data sequences. Furthermore, normalizing the fusion results ensures comprehensive consideration of data from different sensors, further optimizing the effectiveness of multi-source data integration and providing high-quality data support for subsequent analysis and decision-making.
[0015] Preferably, the step of obtaining the matching path includes: Initialize the cumulative distance matrix, set the non-zero row or column values of the cumulative distance matrix to infinity, calculate the cumulative distance of each point in the matrix point by point according to the recursive formula, start from the end point of the matrix, and trace back to the starting point along the direction of the minimum cumulative distance to obtain the matching path.
[0016] The effect is that by initializing the cumulative distance matrix and setting non-zero rows or columns to infinity, the continuity and uniqueness of the matching path are ensured, preventing the occurrence of illegal paths. By calculating the cumulative distance point by point according to a recursive formula, the optimal alignment between two data sequences can be accurately found. Tracing back from the end point of the matrix to the starting point not only ensures the optimality of the path but also provides accurate matching point pair information for subsequent trend consistency analysis. This effectively solves the data alignment issues caused by different acquisition frequencies and time delays, providing a solid foundation for high-quality integration of multi-source data.
[0017] Preferably, the pretreatment step comprises: The data series with the same parameters within a preset time period are denoised, and outliers are removed from the denoised data series. The LSTM model is used to time-align data series with different acquisition frequencies to ensure the accuracy and continuity of timestamps.
[0018] In a second aspect, a multi-data source automatic integration system based on deep learning includes: a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned multi-data source automatic integration method based on deep learning is implemented.
[0019] The present invention has the following effects: 1. This invention addresses the data fusion quality issue inherent in existing technologies, which suffers from an inability to accurately assess credibility, by quantifying trend consistency and delay magnitude between data sources. By using the DTW algorithm to calculate the shortest path distance and the smoothness of matching paths, this method comprehensively assesses the similarity between data sequences, providing a more reliable basis for data fusion and significantly improving the accuracy and reliability of the fusion results.
[0020] 2. This invention provides two methods for calculating trend consistency: a smoothness assessment based on time differences and a global volatility assessment based on slope standard deviation. These two methods are respectively suitable for data series with large local and overall fluctuations, and can flexibly adapt to the needs of different data characteristics. By mapping the matching complexity using a negative exponential function, the confidence calculation is further optimized, making the data fusion process more efficient and adaptable. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a method flow chart of steps S1 to S3 in a method for automatic integration of multiple data sources based on deep learning in an embodiment of the present invention.
[0022] Figure 2 This is a structural block diagram of a multi-data source automatic integration system based on deep learning in an embodiment of the present invention. DETAILED DESCRIPTION
[0023] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, but not all of the embodiments.
[0024] Reference Figure 1 A method for automatic integration of multiple data sources based on deep learning includes steps S1 to S3, which are specifically as follows: S1: Obtain the data sequence of the same parameter of each sensor in the industrial production workshop within a preset time period and preprocess the data sequence.
[0025] Denoising is performed on data sequences with the same parameters within a preset time period to eliminate noise and interference in the data and improve the accuracy and reliability of the data. Denoising methods may include low-pass filtering, moving average method, or wavelet transform, and outliers are removed from the denoised data sequence. Outliers can be detected and corrected or removed through statistical methods such as Z-score and IQR (Interquartile Range) to ensure the rationality of the data. A long short-term memory network (LSTM) model is used to time-align data sequences with different acquisition frequencies to ensure the accuracy and continuity of timestamps. The specific steps are as follows: each data sequence is input into a pre-trained LSTM model. The LSTM model time-aligns the data sequences based on the timestamp and acquisition frequency of the data to ensure the accuracy and continuity of the timestamps, and outputs the aligned data sequence.
[0026] For example, in an industrial production workshop, there are multiple devices, each of which is equipped with a temperature sensor. Within a preset time period (e.g., 1 hour), these sensors record the temperature data of the device at different collection frequencies.
[0027] For example, device A's sensor collects temperature data every 10 seconds, device B's every 15 seconds, and device C's every 20 seconds. The temperature data collected by each sensor is arranged in chronological order to form an independent data series. These data series represent the collection results of the same parameter (device temperature) from multiple data sources within a preset time period. Ensuring these data series are temporally aligned provides high-quality input data for subsequent data fusion processing.
[0028] S2: Calculate the shortest path distance between the preprocessed data sequences corresponding to any two combinations of sensors and the smoothness of the matching paths, and then analyze the trend consistency between the two preprocessed data sequences.
[0029] It should be noted that after preprocessing and time alignment, the data sequences collected by each data source are obtained, among which The data sequence collected by the sensor is expressed as , The first After alignment, the data points of each data sequence correspond one to one and the number of data is the same. During the data integration process, in order to make the fusion result closer to the true value, it is necessary to calculate the confidence level of each sensor data and perform data fusion based on it.
[0030] The steps to obtain the matching path include: Initialize the cumulative distance matrix, set the non-zero row or column values of the cumulative distance matrix to infinity, calculate the cumulative distance of each point in the matrix point by point according to the recursive formula, start from the end point of the matrix, and trace back to the starting point along the direction of the minimum cumulative distance to obtain the matching path.
[0031] Use the DTW (Dynamic Time Warping) algorithm to calculate the shortest path distance between the data sequences corresponding to any two combinations of sensors. Divide the shortest path distance by the total number of data to obtain the normalized shortest distance. Calculate the time difference of each matching point pair in any two-to-two sensor data collection sequence, calculate the difference of the time differences of adjacent matching point pairs and sum them up to obtain the smoothness of the matching path; The product of the normalized shortest distance and the smoothness of the matching path is taken as the trend consistency between any two combinations of sensor data series.
[0032] Specifically, trend consistency satisfies the following relationship: ; Where, Indicates the and The trend consistency between the data series obtained by the sensors, Indicates the and The shortest path distance is calculated by using the DTW algorithm to collect data sequences from sensors. Indicates the total number of data contained in the sensor data collection sequence, Indicates the and The total number of matching point pairs in the data sequence collected by the sensor, Indicates the sequence number of the matching point pair, Indicates the and The first sensor in the data collection sequence The time difference of the matching point pairs, Indicates the and The first sensor in the data collection sequence The time difference between the matching pairs.
[0033] The DTW algorithm calculates the shortest path distance between any two sensor data sequences to measure their shape similarity, and normalizes them to eliminate the effects of dimension and data length. Furthermore, the time difference between matching point pairs and the sum of the differences between adjacent matching point pairs are used to quantify the smoothness of the matching paths, thereby comprehensively evaluating the trend consistency of the two data sequences.
[0034] If the changing trends of the two data series are relatively consistent, the matching path is relatively smooth. If the time differences of all matching points are roughly the same, the smoothness of the matching path is relatively small. Conversely, the greater the smoothness, the less consistent the changing trends of the two data series are.
[0035] Embodiment 1 not only considers the overall length of the matching path, but also focuses on the changes in the time difference in the path. It can carefully reflect the impact of local fluctuations on trend consistency and is suitable for scenarios that require accurate evaluation of local details.
[0036] In addition, in the second embodiment, the DTW algorithm is used to calculate the shortest path distance between the data sequences corresponding to any two combinations of sensors, and the shortest path distance is divided by the total number of data to obtain the normalized shortest distance; The standard deviation of the slope between each pair of adjacent matching points on the matching path is calculated, and the product of the normalized shortest distance and the standard deviation of the slope is used as the trend consistency between any two combinations of sensor data series.
[0037] It should be noted that the slope standard deviation can reflect the degree of change in the slope of each segment in the matching path. A small change in slope indicates a smooth path with high trend consistency. Conversely, a large change in slope indicates a drastic path fluctuation with low trend consistency.
[0038] Specifically, trend consistency satisfies the following relationship: ; Where, Indicates the and The trend consistency between the data series obtained by the sensors, Indicates the and The shortest path distance is calculated by using the DTW algorithm to collect data sequences from sensors. Indicates the total number of data contained in the sensor data collection sequence, Indicates the and The total number of matching point pairs in the sensor data sequence The standard deviation of the slope of the paths between them.
[0039] The standard deviation can reflect the degree of change in the slope of each segment in the matching path. A small change in slope indicates a smooth path with high trend consistency; a large change in slope indicates a drastic path fluctuation with low trend consistency.
[0040] The second embodiment provides a global smoothness evaluation that can quickly measure the overall fluctuation of the path. It is suitable for scenarios where the overall fluctuation of the matching path is large and the smoothness needs to be quickly evaluated. The calculation process is relatively simplified, facilitating efficient processing.
[0041] Because multiple sensors collect the same parameter and suffer from systematic errors and data delays, confidence calculation requires evaluating the trend consistency of each sensor's data with that of other sensors. The DTW algorithm is used to calculate trend consistency. Because the DTW algorithm can adapt to uneven variations in the time axis of two time series, allowing one series to be locally stretched or compressed relative to the other, it better matches the shape and trend of the two series. It can flexibly align local fluctuations, reflecting overall consistency while preserving local matching details, making it suitable for accurately quantifying the similarity between sensor data.
[0042] S3: Calculate the confidence of each sensor data based on trend consistency, and fuse the multi-sensor data series to obtain the fused parameter data series to achieve high-quality multi-source data integration.
[0043] Taking any sensor as the target sensor, traverse and calculate the ratio between the total number of matching point pairs in the data sequences collected by the target sensor and other sensors and the total number of data contained in the data sequences, use the negative exponential function to perform exponential mapping on the comparison values, and use the mapping result as the matching complexity; The matching complexity is used as the weight of the trend consistency between the data sequences of the target sensor and other sensors, and the weighted sum and normalization are performed to obtain the confidence of the target sensor.
[0044] Specifically, the confidence satisfies the following relationship: ; Where, Indicates the The confidence level of each sensor, Indicates the and The trend consistency between the data series collected by the sensors, Indicates the number of sensors collecting the same parameter, Indicates the and The total number of matching point pairs in the data sequence collected by the sensor, Indicates the total number of data contained in the sensor data collection sequence, represents the normalization function, Represented by natural numbers An exponential function with base .
[0045] It should be noted that the The higher the trend consistency between the data series collected by the first sensor and other sensors, the smaller the possibility of abnormality in the data collected by the sensor. The more accurate the data sequence collected by each sensor, the higher the confidence level in the fusion process.
[0046] This quantifies the time delay between data collected by two sensors, reflecting the ratio of the number of matching point pairs required to align the two sequences to the total number of individual sequence data. A ratio close to 1 indicates that the data collected by the two sensors are naturally aligned in time, with a small time delay and high trend consistency. Conversely, a ratio far from 1 indicates a large data delay, reduced naturalness of the time alignment, and decreased confidence in trend consistency. The steps to obtain the fused parameter data sequence include: Taking any sensor as the target sensor, the data sequence collected by the target sensor is weightedly fused based on the confidence of the target sensor, and the fused data sequence is divided by the comprehensive confidence of all sensors to obtain the fused data sequence.
[0047] Specifically, the fused data sequence satisfies the following relationship: ; Where, represents the fused data sequence, Indicates the The data sequence collected by the sensor, Indicates the The confidence level of the data sequence collected by each sensor in the fusion process, Indicates the number of sensors collecting the same parameter.
[0048] Each sensor data sequence is assigned a confidence value, the denominator The sum of the confidence weights is made to be 1 for normalization processing to avoid the absolute size of the confidence affecting the fusion result.
[0049] The present invention also provides a multi-data source automatic integration system based on deep learning. Figure 2 As shown, the system includes a processor and a memory. The memory stores computer program instructions. When the computer program instructions are executed by the processor, a method for automatic integration of multiple data sources based on deep learning according to the first aspect of the present invention is implemented. The system also includes other components familiar to those skilled in the art, such as a communication bus and a communication interface. The configuration and functions of these components are well known in the art and are therefore not described in detail here.
[0050] It should be noted that those skilled in the art may make various modifications and improvements without departing from the scope of the present invention, and these modifications and improvements fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be based on the appended claims.
Claims
1. A method for automatic integration of multiple data sources based on deep learning, characterized in that: include: Obtain the data sequence of the same parameter of each sensor in the industrial production workshop within a preset time period and preprocess the data sequence; Calculate the shortest path distance between the pre-processed data series corresponding to any two combinations of sensors and the smoothness of the matching path, and then analyze the trend consistency between the two pre-processed data series; The confidence of each sensor data is calculated based on trend consistency, and the multi-sensor data series are fused to obtain the fused parameter data series, thus achieving high-quality multi-source data integration.
2. The method for automatic integration of multiple data sources based on deep learning according to claim 1, characterized in that: The calculation method of the trend consistency includes: Use the DTW algorithm to calculate the shortest path distance between the data sequences corresponding to any two combinations of sensors, and divide the shortest path distance by the total number of data to obtain the normalized shortest distance; Calculate the time difference of each matching point pair in any two-to-two sensor data collection sequence, calculate the difference of the time differences of adjacent matching point pairs and sum them up to obtain the smoothness of the matching path; The product of the normalized shortest distance and the smoothness of the matching path is taken as the trend consistency between any two combinations of sensor data series.
3. The method for automatic integration of multiple data sources based on deep learning according to claim 1, characterized in that: Another way to calculate the trend consistency includes: Use the DTW algorithm to calculate the shortest path distance between the data sequences corresponding to any two combinations of sensors, and divide the shortest path distance by the total number of data to obtain the normalized shortest distance; The standard deviation of the slope between each pair of adjacent matching points on the matching path is calculated, and the product of the normalized shortest distance and the standard deviation of the slope is used as the trend consistency between any two combinations of sensor data series.
4. The method for automatic integration of multiple data sources based on deep learning according to claim 1, characterized in that: The step of obtaining the confidence level includes: Taking any sensor as the target sensor, traverse and calculate the ratio between the total number of matching point pairs in the data sequences collected by the target sensor and other sensors and the total number of data contained in the data sequences, use the negative exponential function to perform exponential mapping on the comparison values, and use the mapping result as the matching complexity; The matching complexity is used as the weight of the trend consistency between the data sequences of the target sensor and other sensors, and the weighted sum and normalization are performed to obtain the confidence of the target sensor.
5. The method for automatic integration of multiple data sources based on deep learning according to claim 1, characterized in that: The step of obtaining the fused parameter data sequence includes: Taking any sensor as the target sensor, the data sequence collected by the target sensor is weightedly fused based on the confidence of the target sensor, and the fused data sequence is divided by the comprehensive confidence of all sensors to obtain the fused data sequence.
6. The method for automatic integration of multiple data sources based on deep learning according to claim 1, characterized in that: The step of obtaining the matching path includes: Initialize the cumulative distance matrix, set the non-zero row or column values of the cumulative distance matrix to infinity, calculate the cumulative distance of each point in the matrix point by point according to the recursive formula, start from the end point of the matrix, and trace back to the starting point along the direction of the minimum cumulative distance to obtain the matching path.
7. The method for automatic integration of multiple data sources based on deep learning according to claim 1, characterized in that: The pre-processing steps include: The data series with the same parameters within a preset time period are denoised, and outliers are removed from the denoised data series. The LSTM model is used to time-align data series with different acquisition frequencies to ensure the accuracy and continuity of timestamps.
8. A multi-data source automatic integration system based on deep learning, characterized by: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the method for automatic integration of multiple data sources based on deep learning according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-sensor fusion method and system based on multi-dimensional attribute correlation analysis
CN113761705A
Aquaculture environment monitoring method based on multi-source data fusion
CN118171135A
Construction method of regional basic education comprehensive evaluation large model
CN120258648A
Noise map dynamic deduction method and system based on multi-source heterogeneous data fusion
CN120354064A