Multi-source time sequence test data time alignment method and storage medium
By calculating the maximum data acquisition frequency and determining the reference time axis, the multi-source timing data in the multi-domain joint analysis and evaluation test of underwater navigation bodies was initially aligned and refined, which solved the problems of inconsistent data timing, missing values and field values, significantly reduced the time complexity and improved the accuracy of data evaluation.
Patent Information
- Application Number
- CN202411321161.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-23
- Publication Date
- 2025-06-03
AI Technical Summary
In the multi-domain joint analysis and evaluation test of underwater navigation bodies, there are problems of inconsistent timing, missing values and field values in the time alignment process of multi-source timing data, which affects the authenticity and accuracy of subsequent test evaluations. The time complexity of existing methods is relatively high, making it difficult to meet the continuous improvement of the test methods and the increase in the number of participating products.
By calculating the maximum data acquisition frequency of all products and determining the unique reference time axis, each product is initially aligned in time, and according to the characteristics of static or dynamic products, different methods are used to perform time alignment, including filtering algorithms and interpolation methods, to eliminate wild points and fill in missing data.
The time complexity of time alignment of multi-source timing data is reduced, the efficiency of data preprocessing is improved, and subsequent data analysis is more convenient and accurate.
Smart Images

Figure CN120086504A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a time alignment method in the multi-domain joint analysis and evaluation test of underwater vehicles, and specifically to a time alignment method and storage medium for test data with multiple sources of time series. Background Art
[0002] With the continuous improvement of the test methods for the multi-domain joint analysis and evaluation test of underwater vehicles, the tests often cover various products on land, sea, and air. Different products often carry different data acquisition devices for data acquisition. Due to reasons such as network communication delay and the settings of the data acquisition devices themselves, the product data collected in the same test often has problems such as inconsistent time series, missing values, and outliers, which affect the authenticity and accuracy of the subsequent test evaluation. Therefore, time alignment of multi-source time series data has become an essential part of the evaluation process.
[0003] In a method for matching time series position data of a surface ship and a buoy, a method for cleaning and time series matching of multi-source time series data, and a fuzzy time alignment method for multi-source heterogeneous data for equipment status analysis disclosed in CN114416818A, CN115640284A, and CN117370415A respectively, the time series of the remaining products are matched with the time axis of a certain product as the reference time. If product data collected with a known acquisition frequency is added, it may be necessary to re-match all the data for time series. With the continuous improvement of the test methods, the number of participating products is gradually increasing, and the test time is gradually increasing. The time complexity required by the above matching methods is relatively high when performing time alignment.
[0004] Therefore, it is necessary to provide a method to reduce the time complexity of time alignment of multi-source time series data and perform data preprocessing to facilitate subsequent data analysis. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to overcome the above deficiencies and provide a time alignment method for test data with multiple sources of time series, including:
[0006] Obtain the data of each product in the test, and calculate the maximum data acquisition frequency and the time series data reference time axis of all products;
[0007] Perform a preliminary time alignment for each product;
[0008] Distinguish static products and dynamic products, and use different methods for different types of products to perform time alignment according to the time series data reference time axis.
[0009] Furthermore, the method of the present invention includes the following steps:
[0010] Step 1, obtain the original data, and record the relevant information of each product j in the test as follows:
[0011] Product attributes: divided into dynamic products (the position of the product changes over time during the test) and static products (the position of the product remains basically unchanged over time during the test).
[0012] The data acquisition frequency is denoted as: HZ_j,
[0013] The time series data set is denoted as:
[0014] The time axis of the time series data set is denoted as:
[0015] Step 2, calculate the reference time axis for unifying the time axes of all products, including:
[0016] (1) Calculate the maximum data acquisition frequency HZ_max = max({HZ_j} j=1,...,m ), that is, HZ_max time point information is collected per second under the reference time axis,
[0017] (2) Calculate the range [t_start, t_end] to which the start time and end time of the reference time axis belong, where:
[0018] t_start = max({tt_j_start} j=1,...,m ), t_end = min({tt_j_end} j=1,...,m ),
[0019] (3) Calculate the reference time axis [t 1 , t n , where:
[0020] t i+1 = t i + interval, i = 1,..., n - 1,
[0021] where:
[0022]
[0023] mod(t 1 , interval) = 0, (mod is the modulo operation),
[0024] {t i} i=1,...,n ∈[t_start, t_end],
[0025] n is the number of time points of the reference time axis;
[0026] Step 3, perform preliminary time alignment, including:
[0027] Calculate the relationship between each point tt_j_i on the time axis of each product and the reference time axis [t 1 ,t n ], take t_i as the reference time at time tt_j_i, take the time series data data_j_i at time tt_j_i as the time series data at time t_i, and record the time series data and its corresponding time axis data set after preliminary alignment as
[0028] {data_new_j}({data_new_j}∈{data_j_i}) and {t j,i}, {t j,i The data in} can be repeated;
[0029] Step 4, determine whether each product is a dynamic product, if the product is a dynamic product, go to step 5, otherwise go to step 6;
[0030] Step 5: Perform the following steps on the dynamic product:
[0031] (1) For {t j,i} corresponds to multiple data in the time series data set at the same time, the first complete data is taken as the data at the current reference time. At this time, the time points in {data_new_j} are unique.
[0032] (2) According to {data_new_j} and {t j,i}, and use the filtering algorithm to remove the wild points. The time series data set after removing the wild points and the corresponding time axis set are {data_new2_j}, {t j,m},
[0033] (3) At this time, [t 1 ,t n ] does not belong to {t j,m} i The time series data at the moment is missing, so we use
[0034] The collected information in {data_new2_j} is filled in through interpolation method or transformer model calculation;
[0035] Step 6: Perform the following steps on the static product:
[0036] (1) When there are multiple data_j_i at the same time in {data_new_j}, take the mean of data_j_i (non-empty) as the data at the current reference time. At this time, the time points in {data_new_j} are unique.
[0037] (2) Calculate the mean square error of the position data (x 1 , x 2 ,...) over a period of time and eliminate outliers based on this. That is: for time t j,i , take the position data within the time period [t j,i-3 , t j,i+3 and calculate the mean square error:
[0038]
[0039] Where: is the mean value in the x j,i-3 , t j,i+3 time period, m direction,
[0040] If the mean square error is greater than a certain threshold, it is considered that there are outliers at this moment, and all data at this moment are eliminated. Denote the time series data set after eliminating outliers and the corresponding time axis set as {data_new2_j}, {t j,m},
[0041] (3) At this time, the time series data at time t 1 , t n in [t j,m that does not belong to {t i} is missing, and it is filled with the mean value of the time series data corresponding to the moments with data before and after it;
[0042] Step 7, loop the above steps to unify the time series data of all products under the same time axis.
[0043] A computer-readable storage medium, on which a computer program is stored, and the computer program can be executed by a processor to implement the steps of a method for time alignment of multi-source time series experimental data described in the present invention.
[0044] The beneficial effects of the present invention are:
[0045] Currently, the acquisition frequencies of the devices for collecting product position data are mostly 1Hz, 5Hz, 10Hz, and 20Hz. However, the existing time-series data matching method in the field of experimental technology performs a large number of unnecessary search and matching calculations. When performing time matching for the data at each moment, its time complexity is O(n), and there is a large amount of redundant calculation. The present invention calculates the maximum data acquisition frequency, calculates a unique reference time axis, and the time complexity when performing time matching for the data at each moment is O(1). When the data volume is too large, the preliminary time alignment work of the time-series data can be completed at a faster speed. Under the condition of knowing the maximum acquisition frequency of the acquisition device in an experiment, for the time-series data subsequently added to the time alignment calculation, it will not affect the time-series data that has already completed the time alignment. At the same time, the present invention performs data preprocessing work according to the data characteristics of different data sources, making the time alignment work of multi-source time-series data more complete. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a schematic diagram of the overall flow of the method of the present invention.
[0047] Figure 2 It is a main algorithm flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific implementation methods and with reference to the accompanying drawings. It should be understood that these descriptions are exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention. As Figure 2 shown, the present invention provides a method for time alignment of multi-source time-series experimental data, including the following steps:
[0049] Step 1: Obtain the original data. Record the relevant information of product j in the experiment as follows:
[0050] Product attribute: dynamic product (the position of the product changes over time during the experiment) or static product (the position of the product remains basically unchanged over time during the experiment),
[0051] The data acquisition frequency is denoted as: HZ_j, that is, HZ time points of information of product j are collected per second,
[0052] The time-series data set is denoted as: data_j_i is the information related to the navigation state of product j collected at the i-th moment (where n i is the number of time-series data collected for product j),
[0053] The time axis of the time-series data set is denoted as: where \(tt_{j\_start}\) is the start time of the time series data of product \(j\) collected, and \(tt_{j\_end}\) is the end time of the time series data of product \(j\) collected;
[0054] Step 2, calculate the reference time axis for unifying the time axes of all products, including:
[0055] (1) Calculate the maximum data collection frequency \(HZ_{max}=\max(\{HZ_j\}\) j=1,...,m ), (where \(m\) is the number of products for which product data needs to be collected in the same experiment), that is, \(HZ_{max}\) time point information is collected per second under the reference time axis,
[0056] (2) Calculate the range \([t_{start},t_{end}]\) to which the start time and end time of the reference time axis belong, where:
[0057] \(t_{start}=\max(\{tt_{j\_start}\}\) j=1,...,m ), \(t_{end}=\min(\{tt_{j\_end}\}\) j=1,...,m ),
[0058] (3) Calculate the reference time axis \([t\) 1 ,t\) n , where:
[0059] t\) i+1 =t\) i +interval, i = 1,..., n - 1,
[0060] where:
[0061]
[0062] \(\text{mod}(t\) 1 ,interval)=0, (\text{mod} is the modulo operation),
[0063] \{t\) i}\) i=1,...,n \(\in[t_{start},t_{end}]\),
[0064] n is the number of time points of the reference time axis;
[0065] Step 3, perform preliminary time alignment, including:
[0066] Calculate the closest point \(t_i\) on the reference time axis \([t\) 1 ,t\) n for each point \(tt_{j\_i}\) on the time axis of each product, where \(t_i\in[t\) 1 ,t\) n, taking \(t_i\) as the reference time at the moment \(tt_{j_i}\), regarding the time-series data \(data_{j_i}\) at the moment \(tt_{j_i}\) as the time-series data at the moment \(t_i\), and denoting the time-series data and its corresponding time-axis data set after preliminary alignment as \(\{data_{new_j}\}(\{data_{new_j}\}\in\{data_{j_i}\})\) and \(\{t\) j,i},\(\{t\) j,i} where the data can be repeated. The specific calculation steps are as follows:
[0067] (1) Calculate the smallest interval \([t\) j_i_start , \(t\) j_i_end to which \(tt_{j_i}\) belongs to the reference axis, where:
[0068] \(t\) j_i_start = \(tt_{j_i}-\text{mod}(tt_{j_i},\text{interval})\),
[0069] \(t\) j_i_end = \(tt_{j_i}+(\text{interval}-\text{mod}(tt_{j_i},\text{interval}))\),
[0070] (2) Calculate the point \(t_i\) closest to \(tt_{j_i}\), where:
[0071] \(t\) j_i_start_error = \(\vert tt_{j_i}-t\) j_i_start \vert\),
[0072] \(t\) j_i_end_error = \(\vert tt_{j_i}-t\) j_i_end \vert\),
[0073] \(t_i=\min(t\) j_i_start_error , \(t\) j_i_end_error ),
[0074] (3) Record the time-series data and time axis after preliminary alignment:
[0075] The time-series data \(data_{j_i}(data_{j_i}\in\{data_{new_j}\})\) corresponding to the moment \(tt_{j_i}\) is used as the time-series data at the moment \(t_i\) after matching, that is, there is the following corresponding relationship:
[0076] Moment on the product timeline Moment on the reference timeline Corresponding timing data tt_j_i t_i data_j_i
[0077] Denote the set of corresponding time-series data obtained after matching as \(\{data_{new_j}\}\), and the set of moments on the reference time axis as \(\{t\) j,i};
[0078] Step 4, determine whether each product is a dynamic product. If the product is a dynamic product, go to Step 5; otherwise, go to Step 6;
[0079] Step 5: Perform the following steps on the dynamic product:
[0080] (1) For {t j,i} corresponds to multiple data in the time series data set at the same time, the first complete data is taken as the data at the current reference time. At this time, the time points in {data_new_j} are unique.
[0081] (2) According to {data_new_j} and {t j,i}, the filtering algorithm is used to remove the wild points. The time series data set and the corresponding time axis set after removing the wild points are recorded as {data_new2_j}, {t j,m},
[0082] (3) At this time, [t 1 ,t n ] does not belong to {t j,m} i The time series data at the moment is missing. If there is a small amount of missing time series data, the missing values can be filled by interpolating the data before and after the missing value. However, if there is a large amount of missing time series data, only using the interpolation method will cause the data results to be distorted. It is recommended to use the collected information in {data_new2_j} to calculate and fill it through the transformer model;
[0083] Step 6: Perform the following steps on the static product:
[0084] (1) When there are multiple data_j_i at the same time in {data_new_j}, take the mean of data_j_i (non-empty) as the data at the current reference time. At this time, the time points in {data_new_j} are unique.
[0085] (2) According to {data_new_j}, calculate the position data (x 1 ,x 2 ,...) and remove outliers accordingly. That is: for t j,i Time, take [t j,i-3 ,t j,i+3 ]The position data within the time period is used to calculate the mean square error:
[0086]
[0087] in: for [t j,i-3 ,t j,i+3 ]x in the time period m The mean of the directions,
[0088] If the mean square error is greater than a certain threshold, it is considered that there are outliers at this moment, and all data at this moment are removed. Denote the set of time series data and the corresponding time axis set after removing outliers as {data_new2_j} and {t j,m},
[0089] (3) At this time, the time series data at the t 1 ,t n that do not belong to {t j,m} is missing, and it is filled with the mean of the time series data corresponding to the moments with data before and after it; i
[0090] Step 7: Loop the above steps to unify the time series data of all products under the same time axis.
[0091] In particular, compared with the previous method, this method has a lower time complexity when performing preliminary time alignment. The following uses an example to verify this. The data acquisition frequency of the time series data to be matched is 10Hz, and the length of the time axis is 53092.
[0092] Example of time series data to be matched:
[0093] Time Longitude Latitude 2023-10-10 16:07:52.005 121.69 38.80 2023-10-10 16:07:52.006 121.69 38.81 2023-10-10 16:07:52.397 121.72 38.95
[0094] Example of the reference time axis:
[0095]
[0096]
[0097] Example of data after preliminary time alignment:
[0098] Time Longitude Latitude 2023-10-10 16:07:52.000 121.69 38.80 2023-10-10 16:07:52.000 121.69 38.81 2023-10-10 16:07:52.400 121.72 38.95
[0099] Comparison of the time used for preliminary time alignment of data with a time axis length of 53092:
[0100] Patent usage method mentioned in the background art This method 11.4 seconds 0.4 seconds
[0101] The above time comparison results are obtained based on MATLAB running calculations.
Claims
1. A method for time alignment of test data of multi-source time series, characterized in that: include: Obtain data from each product in the test, and calculate the maximum data collection frequency and the benchmark time axis of the timing data among all products; Perform preliminary time alignment for each product; Differentiate between static products and dynamic products, and use different methods to align different types of products based on the benchmark time axis of the time series data.
2. The method for time alignment of test data of multi-source time series according to claim 1, characterized in that: The steps include: Step 1: Obtain product original data, where the relevant information of each product j in the test in the original data includes: Product attributes: divided into dynamic products and static products, The data collection frequency is recorded as: HZ_j, The time series data set is recorded as: The time axis of the time series data set is recorded as: Step 2, calculate the base timeline used to unify the timelines of all dynamic products and static products, including: (1) Calculate the maximum data acquisition frequency HZ_max = max({HZ_j} j=1,...,m ), (2) Calculate the start time and end time range of the reference time axis [t_start, t_end], (3) Calculate the reference time axis [t1,t n ]; Step 3, perform preliminary time alignment, including: Calculate the time axis of all dynamic products and static products and the time axis of each point tt_j_i n ], take t_i as the reference time of time tt_j_i, take the time series data data_j_i at time tt_j_i as the time series data at time t_i, and record the time series data and its corresponding time axis data set after preliminary alignment as {data_new_j} and {t_new_j} respectively. j,i }, where {data_new_j}∈data_j_i}, {t j,i The data in} can be repeated; Step 4, determine whether each product is a dynamic product, if the product is a dynamic product, go to step 5, otherwise go to step 6; Step 5: Process the dynamic product in the following steps, including: (1) For {t j,i } corresponds to multiple data in the time series data set at the same time, the first complete data is taken as the data at the current reference time. At this time, the time points in {data_new_j} are unique. (2) According to {data_new_j} and {t j,i }, the filtering algorithm is used to remove the wild points. The time series data set and the corresponding time axis set after removing the wild points are {data_new2_j}, {t j,m }, (3) At this time, [t1,t n ] does not belong to {t j,m } i The time series data at the moment is missing and is filled in using interpolation methods or transformer model calculations; Step 6, performing the following steps on the static product, including: (1) When there are multiple data_j_i at the same time in {data_new_j}, take the mean of data_j_i (non-empty) as the data at the current reference time. At this time, the time points in {data_new_j} are unique. (2) Based on {data_new_j}, calculate the mean square error of the position data (x1, x2, ...) over a period of time and remove outliers accordingly. (3) At this time, [t1,t n ] does not belong to {t j,m } i If the time series data of a moment is missing, it is filled with the mean of the time series data corresponding to the moments before and after it. Step 7, looping steps 1 to 6, unifying the time series data of all dynamic products or static products on the same time axis.
3. The method for time alignment of test data of multi-source time series according to claim 2, characterized in that: In step 2, the reference time axis [t1,t n ]include: t i+1 =t i +interval,i=1,...,n-1, mod(t1,interval)=0, where mod is a modulus operation. {t i } i=1,...,n ∈[t_start,t_end], n is the number of time points on the base time axis.
4. The method for time alignment of test data of multi-source time series according to claim 3, characterized in that: The step 3 also includes: (1) Calculate the minimum interval [t j_i_start ,t j_i_end ],in: t j_i_start =tt_j_i-mod(tt_j_i,interval), t j_i_end =tt_j_i+(interval-mod(tt_j_i,interval)), (2) Calculate the point t_i closest to tt_j_i, where: t j_i_start_error =abs(tt_j_i-t j_i_start ), t j_i_end_error =abs(tt_j_i-t j_i_end ), t_i=min(t j_i_start_error ,t j_i_end_error ), (3) Record the timing data and timeline after preliminary alignment: The time series data data_j_i (data_j_i∈{data_new_j}) corresponding to the time tt_j_i is used as the time series data at the time t_i after matching, and has the following corresponding relationship: The corresponding time on the reference time axis is t_i, and the corresponding time series data is data_j_i. The corresponding time series data set obtained after matching is recorded as {data_new_j}, and the time set on the reference time axis is recorded as {t j,i }.
5. The method for time alignment of test data of multi-source time series according to claim 4, characterized in that: Eliminating outliers in step 5 also includes: t j,i time, take [t j,i-3 ,t j,i+3 ]The position data within the time period is used to calculate the mean square error: in: for [t j,i-3 ,t j,i+3 ] is x in the time period m The mean of the directions, If the mean square error is greater than a certain threshold, it is considered that there is an outlier at that moment, and all data at that moment are removed. The time series data set and the corresponding time axis set after removing the outliers are recorded as {data_new2_j}, {t j,m }.
6. The method for time alignment of test data of multi-source time series according to any one of claims 1 to 5, characterized in that: Among the product attributes: The dynamic product is a product whose position changes over time during the test. The static product is a product whose position remains unchanged over time during the test.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program can be executed by a processor to implement the steps of a method for time alignment of experimental data of multiple source time series as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Water surface boat and buoy time sequence position data matching method
CN114416818A
Multi-source time sequence data cleaning and time sequence matching method
CN115640284A
Fuzzy time alignment method of multi-source heterogeneous data for equipment state analysis
CN117370415A
Cited By
Spatial load material scientific experiment data processing method and system and related equipment
CN121071413A