A multi-source time series data alignment method, device, equipment and medium
Patent Information
- Application Number
- CN202610877960.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]然而,强制时间同步可能导致数据丢失或计算中断,影响系统稳定性;而直接丢弃异常数据则容易造成分析结果失真,降低业务决策的准确性
[0023]本发明实施例的技术方案,通过获取目标业务在历史时间段对应的基准数据源及待对齐数据源对应的历史采集数据,历史采集数据包括至少一个采集数据组,采集数据组包括采集值及采集值对应的采集时间;根据基准数据源与待对齐数据源对应的历史采集数据,确定目标动态窗口;针对基准数据源对应的历史采集数据中各采集数据组,根据采集数据组及目标动态窗口,在待对齐数据源对应的历史采集数据中筛选,得到采集数据组对应的至少一个候选数据组;针对各候选数据组,根据候选数据组、采集数据组及目标动态窗口,确定候选数据组对应的数据组权重;根据各候选数据组对应的数据组权重中的权重最小值对应的候选数据组,确定为采集数据组对应的目标数据组,并将采集数据组与目标数据组时间对齐,通过动态时间窗口、非对称最近邻匹配及分级指标补偿技术,有效解决多源异构时序数据时间错位问题,提升数据处理稳定性与业务决策可靠性。
Smart Images

Figure CN122838784A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of financial technology and data processing technology, and in particular to a method, apparatus, device and medium for aligning multi-source time-series data. Background Technology
[0002] With the rapid development of technology, the number of business systems is gradually increasing, and different business systems can process different business data. The received business data from different business systems can be aggregated and analyzed to generate decision-making information. Furthermore, different business systems transmit business data at different frequencies; to ensure the accuracy of the generated decision-making information, time alignment processing of the business data from different business systems is necessary.
[0003] Currently, multi-source time-series data alignment is achieved by forcing time synchronization or directly discarding abnormal timestamp data.
[0004] However, forced time synchronization may lead to data loss or computational interruption, affecting system stability; while directly discarding abnormal data can easily cause distortion of analysis results and reduce the accuracy of business decisions. Summary of the Invention
[0005] This invention provides a method, apparatus, device, and medium for aligning multi-source time-series data to improve the accuracy of multi-source time-series data alignment.
[0006] In a first aspect, embodiments of the present invention provide a method for aligning multi-source time-series data, the method comprising:
[0007] Obtain the baseline data source and the historical data collection data corresponding to the data source to be aligned for the target business in the historical time period. The historical data collection data includes at least one data collection group, and the data collection group includes the collection value and the collection time corresponding to the collection value.
[0008] The target dynamic window is determined based on the historical data collected from the benchmark data source and the data source to be aligned.
[0009] For each data group in the historical data corresponding to the benchmark data source, based on the data group and the target dynamic window, filter the historical data corresponding to the data source to be aligned to obtain at least one candidate data group.
[0010] For each candidate data group, the weight of the corresponding data group is determined based on the candidate data group, the collected data group, and the target dynamic window.
[0011] The candidate data group corresponding to the minimum weight in the weight of each candidate data group is determined as the target data group, and the time of the collected data group and the target data group are aligned.
[0012] Secondly, embodiments of the present invention also provide a multi-source time-series data alignment device, the device comprising:
[0013] The data acquisition module is used to acquire the historical data of the benchmark data source and the data source to be aligned for the target business in the historical time period. The historical data includes at least one data set, and the data set includes the acquired value and the acquisition time corresponding to the acquired value.
[0014] The window determination module is used to determine the target dynamic window based on the historical data collected by the benchmark data source and the data source to be aligned.
[0015] The data filtering module is used to filter the historical data of the data source to be aligned based on the data group and the target dynamic window, and obtain at least one candidate data group corresponding to the data group.
[0016] The weight determination module is used to determine the weight of each candidate data group based on the candidate data group, the collected data group, and the target dynamic window.
[0017] The target group determination module is used to determine the target data group corresponding to the collected data group based on the candidate data group with the minimum weight among the data group weights of each candidate data group, and to align the collected data group with the target data group in time.
[0018] Thirdly, embodiments of the present invention also provide a multi-source time-series data alignment device, the multi-source time-series data alignment device comprising:
[0019] At least one processor; and
[0020] A memory that is communicatively connected to at least one processor; wherein,
[0021] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to execute the multi-source timing data alignment method of any embodiment of the present invention.
[0022] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement the multi-source timing data alignment method of any embodiment of the present invention.
[0023] The technical solution of this invention involves acquiring historical data from a benchmark data source and a data source to be aligned for a target business within a historical time period. The historical data includes at least one data set, each containing a collected value and its corresponding collection time. A target dynamic window is determined based on the historical data from the benchmark and data sources to be aligned. For each data set in the historical data from the benchmark data source, at least one candidate data set is obtained by filtering the historical data from the data source to be aligned, based on the data set and the target dynamic window. For each candidate data set, a data set weight is determined based on the candidate data set, the data set, and the target dynamic window. The candidate data set corresponding to the candidate data set with the minimum weight among the data set weights is identified as the target data set, and the data set and target data set are time-aligned. Through dynamic time windows, asymmetric nearest neighbor matching, and hierarchical index compensation techniques, this solution effectively addresses the time misalignment problem of multi-source heterogeneous time-series data, improving data processing stability and business decision reliability.
[0024] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a flowchart of a multi-source time-series data alignment method provided by an embodiment of the present invention;
[0027] Figure 2 This is a flowchart of a multi-source time-series data alignment method provided by an embodiment of the present invention;
[0028] Figure 3 This is a structural diagram of a multi-source timing data alignment device according to an embodiment of the present invention;
[0029] Figure 4 This is a schematic diagram of the structure of a multi-source time-series data alignment device provided in an embodiment of the present invention. Detailed Implementation
[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0032] The acquisition, storage, and application of historical data involved in the technical solutions of this invention comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0033] Example 1
[0034] Figure 1 This is a flowchart illustrating a multi-source timing data alignment method according to Embodiment 1 of the present invention. This embodiment of the invention is applicable to multi-source timing data alignment. The method can be executed by a multi-source timing data alignment device, which can be implemented in hardware and / or software.
[0035] See Figure 1 The multi-source time-series data alignment method shown includes:
[0036] S101. Obtain the historical data collected by the benchmark data source and the data source to be aligned for the target business in the historical time period. The historical data collected includes at least one data collection group, and the data collection group includes the collected value and the collection time corresponding to the collected value.
[0037] The target business can be any business requiring alignment of multi-source time-series data. The historical time period can refer to a pre-defined, continuous historical time interval used for sample statistics. The benchmark data source can be a data source among multi-source heterogeneous data streams that serves as an alignment reference standard, acting as the baseline for time-series alignment. The collection time of this data source is used as a reference to complete the matching and alignment of other data. The data source to be aligned can be a data source relative to the benchmark data source that requires time-series matching and alignment. Historical collected data can be the full volume of time-series data continuously collected by the benchmark data source and the data source to be aligned within the aforementioned historical time period according to their respective sampling rules, and which has been stored on the database. A collected data group can be the smallest structured unit of time-series data, encapsulating a single time-series record into an independent data group; a single historical collected data record can be split into one or more collected data groups. The collected value can be the business indicator value collected by the data source at the corresponding collection time. The collection time can be the timestamp of the data source completing data collection and generating the corresponding collected value, used to mark the temporal position of data generation and serving as the basis for achieving time-dimensional alignment.
[0038] Specifically, based on the current operational business scenario, the target business is determined. A fixed-length historical time period corresponding to the target business is obtained. Based on the target business, a baseline data source and a data source to be aligned are defined, and the time-series databases corresponding to each type of data source are located; the baseline data source serves as the alignment reference, and the data source to be aligned is the object to be matched. For both data sources, all raw time-series records continuously generated by the data source within the historical time period according to its own sampling frequency are obtained. Each raw time-series record is parsed and split, extracting the collection time (timestamp) and collection value (business indicator value) from the record, binding the two together, and encapsulating them into an independent collection data group. All raw time-series records are traversed, and the above encapsulation operation is repeated, ultimately forming a structured dataset corresponding to the data source, consisting of at least one collection data group, thus determining the historical collection data corresponding to the data source.
[0039] S102. Determine the target dynamic window based on the historical data collected by the benchmark data source and the data source to be aligned.
[0040] The target dynamic window can be a dynamic tolerance time interval, representing the maximum arrival time difference between two data sources that the system can accept; the window does not have a fixed value and can be adaptively adjusted according to data latency fluctuations and business priorities.
[0041] Specifically, from the historical data collected from the benchmark data source and the data source to be aligned, multiple sets of data to be compared are formed according to business relationships. For each set of data to be compared, the collection time of the data from the benchmark data source and the data source to be aligned is read, and the arrival time difference of the two sets of data to the data processing end is calculated, resulting in multiple sets of time difference sample values. Based on all data arrival time difference samples, the average time difference and the standard deviation of the time difference are calculated using mathematical statistics. The average time difference reflects the normal transmission latency level of the two types of data sources; the standard deviation of the time difference characterizes the fluctuation range of the data arrival time difference, and the window expansion range is positively correlated with the standard deviation of the time difference. Using the average time difference as the benchmark, the benchmark interval is extended by combining the standard deviation of the time difference to construct an initial dynamic window. This window determines the time tolerance range only based on the objective laws of data latency. The data priority weights (i.e., business priority weight factors) corresponding to the two types of data sources are retrieved: for data sources corresponding to high-priority businesses, the time tolerance range is expanded; for data sources corresponding to ordinary-priority businesses, the time tolerance range is maintained or narrowed. The initial dynamic window's range is corrected and updated using priority weights, ultimately outputting a target dynamic window that adapts to the business level and data latency characteristics.
[0042] S103. For each data group in the historical data collected from the benchmark data source, filter the historical data collected from the data source to be aligned based on the data group and the target dynamic window to obtain at least one candidate data group.
[0043] Among them, the candidate data group can be the data group collected from the data source to be aligned, whose time difference with the data group collected from the current benchmark data source falls within the target dynamic window range. The data group collected from a single benchmark data source can correspond to one or more candidate data groups.
[0044] Specifically, each data group from the historical data collection of the benchmark data source is read sequentially, and each group is used as the current comparison benchmark. The filtering logic is executed on only one data group from the benchmark data source at a time, until all data groups have been traversed. For the data group from the currently being processed benchmark data source, its collection time is parsed and extracted as the benchmark time for time comparison. All data groups from the data sources to be aligned are retrieved, and the collection time corresponding to each data group from each data source to be aligned is extracted. The time difference between this collection time and the benchmark time is calculated. The calculated time difference is compared with the range of the target dynamic window: if the time difference is within the tolerance range of the target dynamic window, the data group from the data source to be aligned is determined to meet the time constraint; if the time difference exceeds the tolerance range of the target dynamic window, the data group from the data source to be aligned is determined to be invalid data. All data groups from the data sources to be aligned that meet the time constraint are summarized to obtain at least one candidate data group corresponding to the current benchmark data group. After completing the filtering of a single benchmark data group, the above process is repeated for the next benchmark data group.
[0045] S104. For each candidate data group, determine the data group weight corresponding to the candidate data group based on the candidate data group, the collected data group, and the target dynamic window.
[0046] Among them, the data group weight can be a quantitative indicator that represents the spatiotemporal matching degree between the candidate data group and the benchmark data group. The larger the weight value, the lower the matching degree between the two data groups, and the smaller the value, the higher the matching degree.
[0047] Specifically, the current candidate data group and the baseline data group are analyzed separately, extracting their corresponding collected values and times. Simultaneously, the maximum tolerable time difference corresponding to the determined target dynamic window and the maximum historical value range of the full data are read. The collected values of the candidate data group and the baseline data group are compared, and the absolute value of their numerical difference is calculated. This difference is then normalized using the maximum historical value range to obtain a first similarity value representing the degree of numerical difference. The absolute value of the time difference between the candidate data group's collection time and the baseline data group's collection time is calculated, and this difference is normalized using the maximum tolerable time difference of the target dynamic window to obtain a second similarity value representing the degree of time difference. The type of the currently running target service is identified, and based on the service attributes, a first weight corresponding to the first similarity value and a second weight corresponding to the second similarity value are configured: for services emphasizing time accuracy, the second weight is increased; for services emphasizing numerical accuracy, the first weight is increased. The sum of the two weights is a fixed value. The first similarity value is multiplied by the first weight, and the second similarity value is multiplied by the second weight. The sum of these products is then calculated to obtain the data group weight corresponding to the candidate data group. Following the above logic, calculations are performed sequentially for each candidate data group to complete the weight calculation for all candidate data groups.
[0048] S105. Based on the candidate data group corresponding to the minimum weight in the weight of each candidate data group, determine the target data group corresponding to the collected data group, and align the collected data group with the target data group in time.
[0049] The minimum weight value can be the smallest weight value among all candidate data groups corresponding to the same baseline data group, corresponding to the candidate with the best matching effect. The target data group can be the candidate data group pointed to by the minimum weight value, which is the data group to be aligned that has the highest matching degree with the current baseline data group selected from all candidates. Time alignment can be based on the time sequence position of the baseline data group, and the target data group and the baseline data group can be associated and bound in the time dimension, so that the two sets of business data form a one-to-one time sequence relationship, completing the alignment processing of multi-source heterogeneous time sequence data.
[0050] Specifically, for all candidate data groups corresponding to the current benchmark data group, the calculated weights of each candidate data group are aggregated to form a set of weight values. The set of weight values is traversed, all weights are compared, and the minimum weight value is selected; simultaneously, the candidate data group uniquely corresponding to this weight value is located. The candidate data group corresponding to the minimum weight value is designated as the target data group paired with the current benchmark data group. An association mapping relationship is established between the benchmark data group and the target data group. Using the acquisition time of the benchmark data group as a time-series reference, the target data group is time-bound with the benchmark data group, forming a pairing relationship in both business logic and time dimensions, completing the time alignment of a single data group. The process is repeated for the next data group acquired from the benchmark data source until all data groups within the benchmark data source have been paired and time-aligned.
[0051] The technical solution of this invention involves acquiring historical data from a benchmark data source and a data source to be aligned for a target business within a historical time period. The historical data includes at least one data set, each containing a collected value and its corresponding collection time. A target dynamic window is determined based on the historical data from the benchmark and data sources to be aligned. For each data set in the historical data from the benchmark data source, at least one candidate data set is obtained by filtering the historical data from the data source to be aligned, based on the data set and the target dynamic window. For each candidate data set, a data set weight is determined based on the candidate data set, the data set, and the target dynamic window. The candidate data set corresponding to the candidate data set with the minimum weight among the data set weights is identified as the target data set, and the data set and target data set are time-aligned. Through dynamic time windows, asymmetric nearest neighbor matching, and hierarchical index compensation techniques, this solution effectively addresses the time misalignment problem of multi-source heterogeneous time-series data, improving data processing stability and business decision reliability.
[0052] Example 2
[0053] Figure 2 This is a flowchart illustrating a multi-source time-series data alignment method according to Embodiment 2 of the present invention. Based on the above embodiments, this embodiment optimizes and improves the multi-source time-series data alignment operation.
[0054] Furthermore, the step of "determining the target dynamic window based on the historical data collected from the benchmark data source and the data source to be aligned" is refined into "determining at least one set of data to be compared and the time difference of arrival of each set of data based on the historical data collected from the benchmark data source and the data source to be aligned; calculating the average time difference and standard deviation of the time difference based on the time difference of arrival of each set of data; determining the initial dynamic window based on the average time difference and standard deviation of the time difference; updating the initial dynamic window based on the data priority weights of the benchmark data source and the data source to be aligned to obtain the target dynamic window," in order to improve the operation of multi-source time series data alignment.
[0055] It should be noted that for parts not described in detail in the embodiments of the present invention, please refer to the descriptions in other embodiments.
[0056] See Figure 2 The multi-source time-series data alignment method shown includes:
[0057] S201. Obtain the historical data collected by the benchmark data source and the data source to be aligned for the target business in the historical time period. The historical data collected includes at least one data collection group, and the data collection group includes the collected value and the collection time corresponding to the collected value.
[0058] S202. Based on the historical data collected by the benchmark data source and the data source to be aligned, determine at least one set of data to be compared and the time difference of data arrival for each set of data to be compared.
[0059] The data arrival time difference can be the time difference between the data collected on the reference side and the data collected on the side to be aligned in the same set of data to be compared, when the data processing system receives the data, and is used to characterize the transmission delay difference between the two types of business data.
[0060] Specifically, all historical data collected from the baseline data source and the data source to be aligned are read separately. The collection time and value for each data set are retrieved, forming two complete time-series datasets. Based on the logical relationships of the target business and the order in which the data was generated, the two datasets are paired: baseline data sets representing the same business behavior and those with corresponding business relationships are paired with data sets to be aligned to form a set of comparison data. This pairing process is repeated until at least one or more independent sets of comparison data are generated. For each completed set of comparison data, the collection time of the baseline data set and the collection time of the data set to be aligned are extracted. Using the baseline collection time as a reference, the difference between the two collection times is calculated to obtain the arrival time difference of the data for the current comparison data set. This calculation is performed sequentially on all comparison data sets to obtain a complete set of time difference samples.
[0061] S203. Calculate the average time difference and standard deviation of the time difference based on the arrival time difference of the data corresponding to each data group to be compared.
[0062] Specifically, the arrival time differences of all data groups to be compared are summarized to form a complete time difference sample dataset. The arrival time differences of all data in the sample dataset are summed, and then divided by the total number of samples to calculate the arithmetic mean, which is the average time difference. The statistic obtained by calculating the arithmetic mean of all data arrival time difference samples characterizes the normal transmission latency level of the two data sources over a historical period. Using the average time difference as a benchmark, the deviation of each data arrival time difference from the average time difference is calculated sequentially; the squares of all deviations are then averaged, and the square root of this average is taken to obtain the standard deviation of the time difference, which is used to quantify the fluctuation and dispersion of the data arrival time difference; the larger the standard deviation, the more significant the fluctuation in data latency. The calculated average time difference and standard deviation of the time difference are stored as input parameters for subsequently constructing the initial dynamic window.
[0063] S204. Determine the initial dynamic window based on the average time difference and the standard deviation of the time difference.
[0064] Specifically, the system reads the average time difference and standard deviation of the time difference from the historical time-series data statistics output. The average time difference reflects the overall transmission latency level of the benchmark data source and the data source to be aligned under normal conditions, while the standard deviation quantifies the random fluctuation range of data arrival latency. The window expansion range is set to be positively correlated with the standard deviation of the time difference: when factors such as network fluctuations or increased system load cause greater fluctuations in data arrival latency and an increase in the standard deviation of the time difference, the window interval is automatically expanded; when the data transmission status is stable, latency fluctuations decrease, and the standard deviation of the time difference decreases, the window interval is automatically narrowed. This achieves adaptive adjustment of the window size according to the data latency status, effectively adapting to the dynamic changes in data latency under complex operating conditions. In the interval calculation stage, the average time difference is set as the central reference benchmark of the time window, and the benchmark value is bidirectionally extended based on the standard deviation of the time difference. The lower bound of the initial dynamic window is determined based on the difference between the average time difference and the standard deviation of the time difference; the upper bound of the initial dynamic window is determined based on the sum of the average time difference and the standard deviation of the time difference. This time interval is defined as the initial dynamic window.
[0065] S205. Update the initial dynamic window according to the data priority weights corresponding to the baseline data source and the data source to be aligned, and obtain the target dynamic window.
[0066] Specifically, the data priority weights corresponding to the baseline data source and the data source to be aligned are read to obtain the calculated initial dynamic window and its corresponding lower and upper time bounds. The business types of the two data sources are analyzed, and high-priority and ordinary-priority businesses are distinguished based on pre-configured rules to clarify the adjustment direction of the window correction. The data priority weights are applied to the upper and lower boundaries of the initial dynamic window: if the data source corresponds to a high-priority business, the data priority weights are used to amplify the range of the initial window, widening the overall time tolerance range; if the data source corresponds to an ordinary-priority business, the initial window range remains unchanged, or the corresponding weights are used to slightly narrow the time tolerance range. After weight correction, new lower and upper time bounds are obtained, and a new continuous time interval is formed by the new upper and lower boundaries. This interval is defined as the target dynamic window.
[0067] S206. For each data group in the historical data collected from the benchmark data source, filter the historical data collected from the data source to be aligned based on the data group and the target dynamic window to obtain at least one candidate data group.
[0068] S207. For each candidate data group, determine the data group weight corresponding to the candidate data group based on the candidate data group, the collected data group, and the target dynamic window.
[0069] S208. Based on the candidate data group corresponding to the minimum weight in the weight of each candidate data group, determine the target data group corresponding to the data collection group, and align the data collection group and the target data group in time.
[0070] This invention, through its embodiments, determines at least one set of data groups to be compared and the corresponding time differences of data arrival for each set based on historical data collected from a benchmark data source and a data source to be aligned. It then calculates the average time difference and standard deviation of the time difference based on these differences. An initial dynamic window is determined based on the average time difference and standard deviation. The initial dynamic window is updated according to the data priority weights of the benchmark and data sources to be aligned, resulting in a target dynamic window. This generates an initial dynamic window that can adapt to changes in latency, solving the problem that fixed time windows cannot adapt to dynamic transmission latency. Furthermore, by combining the business priority weights of the data sources with the window's adjustment, it differentiates the time tolerance standards for core and ordinary business operations, effectively preventing the loss of core data and interruption of business calculations.
[0071] Optionally, the weight of the data group corresponding to the candidate data group is determined based on the candidate data group, the collected data group, and the target dynamic window, including: determining the first similarity value based on the collected values corresponding to the candidate data group and the collected data group; determining the second similarity value based on the collection time and the target dynamic window corresponding to the candidate data group and the collected data group; determining the first weight corresponding to the first similarity value and the second weight corresponding to the second similarity value based on the business type corresponding to the target business; and determining the weight of the data group corresponding to the candidate data group by weighted summation of the first similarity value, the first weight, the second similarity value, and the second weight.
[0072] The first similarity value is a quantified value calculated based on the collected values of the candidate data group and the benchmark data group. It characterizes the degree of difference between the two data groups in the numerical dimension; a larger value indicates a lower numerical matching degree, and a smaller value indicates a higher numerical matching degree. The second similarity value is a quantified value calculated based on the collection time of the candidate data group and the benchmark data group, combined with the target dynamic window. It characterizes the degree of difference between the two data groups in the time dimension; a larger value indicates a lower time matching degree, and a smaller value indicates a higher time matching degree. The first weight can be a weighting coefficient of the first similarity value, used to characterize the proportion of the numerical dimension in the overall matching degree calculation. The second weight can be a weighting coefficient of the second similarity value, used to characterize the proportion of the time dimension in the overall matching degree calculation.
[0073] Specifically, the collected values corresponding to the candidate data group and the baseline collected data group are extracted separately, and the absolute value of the difference between the two sets of collected values is calculated. Normalization is performed by combining the numerical fluctuation range of all historical collected values to obtain the first similarity value representing the numerical difference. The collection time corresponding to the candidate data group and the baseline collected data group is extracted separately, and the absolute value of the difference between the two sets of collection time is calculated. The maximum tolerable time difference corresponding to the generated target dynamic window is retrieved, and the time difference is normalized using this value as a benchmark to obtain the second similarity value representing the time difference. The type of the currently running target service is identified, and the preset weight allocation rules are invoked: if it is a service that prioritizes time accuracy, a larger second weight and a smaller first weight are configured; if it is a service that prioritizes numerical accuracy, a larger first weight and a smaller second weight are configured. According to the above rules, the first weight and second weight under the current scenario are determined. The first similarity value is multiplied by the first weight to obtain the numerical dimension weighting term; the second similarity value is multiplied by the second weight to obtain the time dimension weighting term; the two weighting results are summed, and the summed value is the data group weight corresponding to the current candidate data group. Following the above process, calculations are performed one by one on each candidate data group corresponding to the current benchmark data group to complete the weighting of all candidate data groups.
[0074] The first similarity value is determined based on the collected values corresponding to the candidate data group and the collected data group; the second similarity value is determined based on the collection time and target dynamic window corresponding to the candidate data group and the collected data group; the first weight corresponding to the first similarity value and the second weight corresponding to the second similarity value are determined based on the business type corresponding to the target business; the weight of the data group corresponding to the candidate data group is determined by weighted summation of the first similarity value, the first weight, the second similarity value and the second weight. The similarity value is calculated in different dimensions, while taking into account time synchronization and business numerical consistency, to achieve spatiotemporal dual-constraint matching and improve the accuracy of time series data alignment.
[0075] Optionally, after obtaining the historical data collected from the benchmark data source and the data source to be aligned for the target business in the historical time period, the historical data collected includes at least one data collection group, which includes the collected value and the collection time corresponding to the collected value. The process further includes: performing data missing detection on the historical data collected from each data source in the data source set to determine the data missing type for each data source. The data source set includes the benchmark data source and the data source to be aligned. For each data source, based on the data missing type, a data compensation rule is determined, and the missing data in the historical data collected from the data source is updated according to the data compensation rule to obtain the compensation data corresponding to the missing data. The effective data range corresponding to the data source is obtained. Based on the effective data range and the compensation data, an offset value is determined. The compensation data is updated based on the offset value.
[0076] Data missing detection can be a process of traversing and verifying historically collected data item by item and time-series position to identify missing data points and statistically analyze the distribution characteristics of missing data. Data missing types can be categorized based on the temporal distribution of missing data, such as single-point missing types and continuous missing types, to match differentiated compensation strategies. Data compensation rules can be pre-defined data completion logic for different data missing types, with independent compensation rules corresponding to different missing types. Missing data can be data in historically collected data where there is no valid collected value at the corresponding time-series position, or where the data content is empty. Data update can be the operation of filling missing data points and correcting the original dataset using compliant values according to the compensation rules. Compensated data can be the numerical result used by the compensation rules to fill missing data points. The effective data range can be the normal numerical fluctuation range obtained from the statistical analysis of historical collected values over a long period of time from the data source, defining the reasonable upper and lower limits of the data source's business indicators as the basis for verifying the reasonableness of the compensated data. The offset value can be the deviation between the compensated data and the effective data range of the data source.
[0077] Specifically, the process iterates through the baseline data source and the data source to be aligned within the data source set, performing a time-series scan and integrity check on the historical data collected by each data source. Missing data at all time-series locations is identified, and the distribution characteristics of the missing data at each time point are statistically analyzed to distinguish between single-point and continuous missing data, thus determining the data missing type for each data source. For each data source, based on the identified data missing type, the pre-defined corresponding data compensation rules are retrieved. Following the execution logic in the rules, missing data points in the historical data are numerically filled and updated, and compensation data is calculated and generated to fill the gaps, completing the initial completion of the original missing data. Statistical analysis is performed on the full volume of normal historical data collected from a single data source to extract the maximum and minimum values of the business indicator over a long period, defining the effective data range for that data source indicator. The initially generated compensation data is compared with the effective data range of the corresponding data source, and the difference between the compensation data and the boundary of the effective range is calculated to obtain an offset value representing the degree of deviation. If the compensation data falls within the effective data range, the offset value is within a reasonable range; if the compensation data exceeds the effective range, the offset value exceeds the limit and is determined to be abnormal compensation data. The compensation data is corrected based on the calculated offset value: when the offset value is within a reasonable range, the current compensation data remains unchanged; when the offset value exceeds the limit or the compensation data exceeds the normal range, the compensation data is adjusted to fall back within the valid data range. This process is repeated for all data sources in the data source set until all historical data collected from all data sources has been completed and corrected for compliance, resulting in the output of a time-series dataset.
[0078] By performing data missing detection on historical data collected from each data source in the data source set (including a baseline data source and data sources to be aligned), the data missing type for each data source is determined. For each data source, data compensation rules are determined based on the data missing type, and the missing data in the historical data collected from that data source is updated according to these rules to obtain the compensated data. The effective data range for each data source is then obtained. Based on the effective data range and the compensated data, an offset value is determined. The compensated data is then updated based on the offset value. Different data compensation measures are implemented for different data missing types, refining the data compensation steps and improving the accuracy of time-series data acquisition.
[0079] Optionally, perform data missing detection on the historical data collected from each data source in the data source set to determine the data missing type for each data source. This includes: identifying the historical data collected from each data source in the data source set to determine at least one missing data point for each data source; for each data source, obtaining the missing time point corresponding to each missing data point; performing continuity detection based on the missing time points corresponding to each missing data point to determine at least one detection data group; obtaining the number of time points corresponding to each detection data group; for each detection data group, comparing the number of time points corresponding to the detection data group with a preset threshold to obtain a threshold comparison result; when the threshold comparison result shows that the number of time points corresponding to the detection data group is greater than or equal to the preset threshold, determining the data missing type of the missing data in the detection data group as a continuous missing type; when the threshold comparison result shows that the number of time points corresponding to the detection data group is less than the preset threshold, determining the data missing type of the missing data in the detection data group as a single point missing type.
[0080] Specifically, the entire historical data collection of each data source within the data source set is traversed, and the validity of the collected values in each data group is verified one by one. Data groups with empty or invalid values are identified and marked as missing data, and at least one or more missing data points are determined in each data source. For each identified missing data point, its corresponding missing time point is extracted, and the specific time of each data gap in the time series is recorded to establish a one-to-one correspondence between missing data and missing time points. All missing time points are sorted in chronological order, and a time series continuity test is performed: consecutive missing time points that are adjacent to each other and have no valid data intervals are grouped into the same detection data group; missing time points with valid data intervals are divided into separate groups or assigned to other groups, ultimately forming one or more independent detection data groups. For each partitioned detection data group, the number of missing time points contained in the group is counted to obtain the statistical value of missing points in each group. A pre-configured preset threshold is obtained, and the number of time points corresponding to a single detection data group is compared with the threshold to generate the corresponding threshold comparison result. If the threshold comparison result shows that the number of time points in the detected data group is greater than or equal to the preset threshold, then all missing data in that group are uniformly classified as consecutive missing data. If the threshold comparison result shows that the number of time points in the detected data group is less than the preset threshold, then all missing data in that group are uniformly classified as single-point missing data. After completing the type determination for all detected data groups under the current data source, switch to the next data source and repeat the above process until all data sources in the data source set have been processed.
[0081] By identifying historical data collected from each data source in the data source set, at least one missing data point is determined for each data source. For each data source, the missing time points corresponding to the missing data are obtained. Continuity detection is performed based on the missing time points corresponding to each missing data point to determine at least one detection data group. The number of time points corresponding to each detection data group is obtained. For each detection data group, the number of time points corresponding to the detection data group is compared with a preset threshold to obtain a threshold comparison result. When the threshold comparison result shows that the number of time points corresponding to the detection data group is greater than or equal to the preset threshold, the data missing type of the corresponding missing data in the detection data group is determined to be a continuous missing type. When the threshold comparison result shows that the number of time points corresponding to the detection data group is less than the preset threshold, the data missing type of the corresponding missing data in the detection data group is determined to be a single-point missing type. This method accurately classifies data groups with different time-series characteristics, improving the accuracy of data missing type determination.
[0082] Optionally, based on the data missing type corresponding to the data source, a data compensation rule is determined, and the missing data in the historical data collected from the data source is updated according to the data compensation rule to obtain the compensation data corresponding to the missing data. This includes: obtaining the historical data collected from the data source within the historical time period; for each detection data group, when the data missing type corresponding to the detection data group is a continuous missing type, identifying the historical data collected from the data source according to the pre-trained time series data prediction model, and predicting the compensation data corresponding to the missing time point of each missing data in the detection data group; when the data missing type corresponding to the detection data group is a single missing type, determining the data change type based on the historical data collected from the data source; and generating the compensation data corresponding to the missing time point of the missing data in the detection data group according to the data compensation rule corresponding to the data change type.
[0083] Specifically, the process iterates through each detection data set, executing compensation logic in two branches based on the determined data missing type: If the current detection data set is of the continuous missing type, the historical data collected from this data source is input into a pre-trained time-series data prediction model. The model learns the time-series change trends, numerical fluctuation patterns, and periodic characteristics of the historical data, performs numerical prediction for each missing time point within the detection data set, and outputs the compensation data corresponding to each missing position. If the current detection data set is of the single-point missing type, the effective collected data adjacent to the missing point is extracted, the numerical trend of the local time series is analyzed, and the data change type of the current data is determined. The data compensation rule matching this data change type is retrieved, and compensation data corresponding to the missing time point is generated. The compensation data corresponding to each missing time point is filled into the empty positions of the original historical collected data, completing the completion of a single detection data set; the above process is repeated until all detection data sets under the current data source are processed. The process then switches to the next data source in the data source set and repeats the above compensation process until both the baseline data source and the data source to be aligned have completed data completion.
[0084] By acquiring historical data from the data source within a historical time period, and for each detection data group, when the data missing type of the detection data group is continuous, the system identifies the historical data from the data source based on a pre-trained time series data prediction model, and predicts the compensation data corresponding to the missing time points of each missing data in the detection data group; when the data missing type of the detection data group is single-point missing, the system determines the data change type based on the historical data from the data source; and generates compensation data corresponding to the missing time points of the missing data in the detection data group according to the data compensation rules corresponding to the data change type. Differentiated compensation strategies are adopted for continuous and single-point missing scenarios, combining intelligent time series prediction with lightweight compensation rules to ensure the quality of data completion while optimizing the system's computing power allocation and improving the integrity of the original data.
[0085] Optionally, based on the data compensation rules corresponding to the data change type, compensation data corresponding to the missing time point of the missing data in the detection data group is generated, including: when the data change type is the first change type, the preceding data of the missing data in the detection data group is determined as the compensation data corresponding to the missing data; when the data change type is the second change type, the compensation data corresponding to the missing time point of the missing data is calculated by interpolation based on at least one adjacent data of the missing data in the detection data group.
[0086] The first type of change can be a constant time-series data pattern in a single-point-of-missing scenario, meaning that the effective collected values before and after the missing point remain relatively stable without significant fluctuations, and the values remain constant within a short time interval. Preceding data can be the collected values corresponding to the data group that is temporally adjacent to and valid before the current missing data point in the time series, serving as the direct data source for equivalent replacement complements. The second type of change can be a linear / continuous gradual change pattern in time-series data in a single-point-of-missing scenario, meaning that the effective collected values around the missing point show an upward or downward trend over time. Adjacent data can be one or more effective collected data points distributed at temporal positions before and after the missing point.
[0087] Specifically, for detection data sets identified as single-point missing, multiple valid data points before and after the missing time point are extracted. The numerical fluctuations and trends of the local time series are analyzed to determine whether the current data belongs to the first or second change type. If it is determined to be the first change type, i.e., the local data values are generally constant without a clear gradual trend, the immediately preceding valid data of the missing data is extracted, and the collected value of the preceding data is directly used as the compensation data corresponding to the current missing data to complete the missing position filling. If it is determined to be the second change type, i.e., the local data values show a smooth gradual trend over time, at least one adjacent valid data point before and after the missing time point is extracted. Combining the collection time and collection value corresponding to each adjacent data point, a preset interpolation algorithm is used to perform interpolation calculations to obtain the compensation data corresponding to the missing time point. The above determination and completion operations are sequentially performed on all detection data sets belonging to the single-point missing type in the current data source until all single-point missing points are filled with data, and then the same process is performed on other data sources.
[0088] When the data change type is the first type, the preceding data of the missing data in the detection data group is determined as the compensation data corresponding to the missing data. When the data change type is the second type, the compensation data corresponding to the missing time point of the missing data is calculated by interpolation based on at least one adjacent data of the missing data in the detection data group. Different completion schemes are matched for the trends of single-point data of constant and gradual types to ensure smooth transition of values in the missing interval, avoid numerical jumps, and improve the authenticity and rationality of the completed data.
[0089] Example 3
[0090] Figure 3 This is a schematic diagram of a multi-source timing data alignment device provided in Embodiment 3 of the present invention. This embodiment of the invention is applicable to multi-source timing data alignment, and the device can execute a multi-source timing data alignment method. The device can be implemented in hardware and / or software.
[0091] See Figure 3The multi-source time-series data alignment device shown includes: a data acquisition module 301, a window determination module 302, a data filtering module 303, a weight determination module 304, and a target group determination module 305, wherein...
[0092] Data acquisition module 301 is used to acquire the historical acquisition data corresponding to the benchmark data source and the data source to be aligned for the target business in the historical time period. The historical acquisition data includes at least one acquisition data group, and the acquisition data group includes the acquisition value and the acquisition time corresponding to the acquisition value.
[0093] The window determination module 302 is used to determine the target dynamic window based on the historical data collected by the benchmark data source and the data source to be aligned.
[0094] The data filtering module 303 is used to filter the historical data of the data source to be aligned based on the data group and the target dynamic window, and obtain at least one candidate data group corresponding to the data group.
[0095] The weight determination module 304 is used to determine the weight of the data group corresponding to each candidate data group based on the candidate data group, the collected data group and the target dynamic window.
[0096] The target group determination module 305 is used to determine the target data group corresponding to the collected data group based on the candidate data group corresponding to the minimum weight in the weight of each candidate data group, and to align the collected data group with the target data group in time.
[0097] The technical solution of this invention involves acquiring historical data from a benchmark data source and a data source to be aligned for a target business within a historical time period. The historical data includes at least one data set, each containing a collected value and its corresponding collection time. A target dynamic window is determined based on the historical data from the benchmark and data sources to be aligned. For each data set in the historical data from the benchmark data source, at least one candidate data set is obtained by filtering the historical data from the data source to be aligned, based on the data set and the target dynamic window. For each candidate data set, a data set weight is determined based on the candidate data set, the data set, and the target dynamic window. The candidate data set corresponding to the candidate data set with the minimum weight among the data set weights is identified as the target data set, and the data set and target data set are time-aligned. Through dynamic time windows, asymmetric nearest neighbor matching, and hierarchical index compensation techniques, this solution effectively addresses the time misalignment problem of multi-source heterogeneous time-series data, improving data processing stability and business decision reliability.
[0098] Optionally, the window determination module 302 is specifically used for:
[0099] Based on the historical data collected by the benchmark data source and the data source to be aligned, determine at least one set of data to be compared and the time difference of data arrival for each set of data to be compared.
[0100] The mean time difference and standard deviation of the time difference are calculated based on the time difference of arrival of the data corresponding to each data group to be compared.
[0101] The initial dynamic window is determined based on the average time difference and the standard deviation of the time difference;
[0102] Based on the data priority weights corresponding to the baseline data source and the data source to be aligned, the initial dynamic window is updated to obtain the target dynamic window.
[0103] Optionally, the weight determination module 304 is specifically used for:
[0104] The first similarity value is determined based on the collected values corresponding to the candidate data group and the collected data group;
[0105] The second similarity value is determined based on the collection time and target dynamic window corresponding to the candidate data group and the collected data group;
[0106] Based on the business type corresponding to the target business, determine the first weight corresponding to the first similarity value and the second weight corresponding to the second similarity value;
[0107] The weights of the candidate data groups are determined by weighted summation of the first similarity value, the first weight, the second similarity value, and the second weight.
[0108] Optionally, the multi-source time-series data alignment device also includes:
[0109] The data detection module is used to obtain the historical data collected by the benchmark data source and the data source to be aligned for the target business in the historical time period. The historical data collected includes at least one data collection group, which includes the collected value and the collection time corresponding to the collected value. Then, it performs data missing detection on the historical data collected by each data source in the data source set to determine the data missing type of each data source. The data source set includes the benchmark data source and the data source to be aligned.
[0110] The rule determination module is used to determine the data compensation rules for each data source based on the data missing type corresponding to the data source, and update the missing data in the historical data collected by the data source according to the data compensation rules to obtain the compensation data corresponding to the missing data.
[0111] The range determination module is used to obtain the valid data range corresponding to the data source;
[0112] The offset value determination module is used to determine the offset value based on the valid data range and compensation data;
[0113] The data update module is used to update the compensation data based on the offset value.
[0114] Optional, a data detection module, specifically used for:
[0115] Identify the historical data collected by each data source in the data source set and determine at least one missing data for each data source;
[0116] For each data source, obtain the missing time points corresponding to each missing data point in the data source;
[0117] Based on the missing time points corresponding to each missing data, perform continuous detection to determine at least one detection data group;
[0118] Obtain the number of time points corresponding to each detection data group;
[0119] For each set of detection data, the number of time points corresponding to the set of detection data is compared with a preset threshold to obtain the threshold comparison result;
[0120] When the threshold comparison result is that the number of time points corresponding to the detection data group is greater than or equal to the preset threshold, the data missing type of the corresponding missing data in the detection data group is determined to be the continuous missing type.
[0121] When the threshold comparison result shows that the number of time points corresponding to the detection data group is less than the preset threshold, the data missing type of the corresponding missing data in the detection data group is determined to be a single point missing type.
[0122] Optional, the rule determination module includes:
[0123] The data acquisition unit is used to acquire historical data collected from the data source within a historical time period.
[0124] The data identification unit is used to identify the historical data collected from the data source based on a pre-trained time series data prediction model when the data missing type of the detection data group is continuous missing. It then predicts the compensation data corresponding to the missing time point of each missing data in the detection data group.
[0125] The type determination unit is used to determine the data change type based on the historical data collected from the data source when the data missing type corresponding to the detection data group is a single point missing type.
[0126] The data compensation unit is used to generate compensation data corresponding to the missing time points of the missing data in the detection data group according to the data compensation rules corresponding to the data change type.
[0127] Optional, a data compensation unit, specifically used for:
[0128] When the data change type is the first change type, the preceding data of the missing data in the detection data group is determined as the compensation data corresponding to the missing data;
[0129] When the data change type is the second change type, the compensation data corresponding to the missing time point of the missing data is calculated by interpolation based on at least one adjacent data of the missing data in the detection data group.
[0130] The multi-source time-series data alignment device provided in the embodiments of the present invention can execute the multi-source time-series data alignment method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the multi-source time-series data alignment method.
[0131] Example 4
[0132] Figure 4 A schematic diagram of the structure of a multi-source timing data alignment device 400 that can be used to implement embodiments of the present invention is shown.
[0133] like Figure 4 As shown, the multi-source timing data alignment device 400 includes at least one processor 401 and a memory, such as a read-only memory (ROM) 402 and a random access memory (RAM) 403, communicatively connected to the at least one processor 401. The memory stores computer programs executable by the at least one processor. The processor 401 can perform various appropriate actions and processes based on the computer program stored in the ROM 402 or loaded from storage unit 408 into the RAM 403. The RAM 403 can also store various programs and data required for the operation of the multi-source timing data alignment device 400. The processor 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0134] Multiple components in the multi-source timing data alignment device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a disk, optical disk, etc.; and a communication unit 409, such as a network interface card, modem, wireless transceiver, etc. The communication unit 409 allows the multi-source timing data alignment device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0135] Processor 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 401 performs the various methods and processes described above, such as multi-source timing data alignment methods.
[0136] In some embodiments, the multi-source timing data alignment method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed onto the multi-source timing data alignment device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by processor 401, one or more steps of the multi-source timing data alignment method described above may be performed. Alternatively, in other embodiments, processor 401 may be configured to perform the multi-source timing data alignment method by any other suitable means (e.g., by means of firmware).
[0137] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0138] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0139] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0140] To provide user interaction, the systems and techniques described herein can be implemented on a multi-source timing data alignment device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the multi-source timing data alignment device. Other types of devices can also be used to provide user interaction; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0141] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0142] A computing system can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability.
[0143] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0144] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for aligning multi-source time-series data, characterized in that, The method includes: Obtain the historical data collected from the baseline data source and the data source to be aligned for the target business in a historical time period. The historical data collected includes at least one data collection group, and the data collection group includes the collected value and the collection time corresponding to the collected value. The target dynamic window is determined based on the historical data collected by the benchmark data source and the data source to be aligned. For each of the historical data collection groups corresponding to the benchmark data source, at least one candidate data group corresponding to the data collection group is obtained by filtering the historical data collection data corresponding to the data source to be aligned, based on the data collection group and the target dynamic window. For each candidate data group, the data group weight corresponding to the candidate data group is determined based on the candidate data group, the collected data group, and the target dynamic window. The candidate data group corresponding to the minimum weight among the weights of each candidate data group is determined as the target data group corresponding to the collected data group, and the collected data group and the target data group are time-aligned.
2. The method according to claim 1, characterized in that, The step of determining the target dynamic window based on the historical data collected by the benchmark data source and the data source to be aligned includes: Based on the historical data collected by the benchmark data source and the data source to be aligned, at least one set of data to be compared and the data arrival time difference corresponding to each set of data to be compared are determined. The mean time difference and standard deviation of the time difference are calculated based on the time difference of arrival of the data corresponding to each of the data groups to be compared. The initial dynamic window is determined based on the average time difference and the standard deviation of the time difference; The initial dynamic window is updated based on the data priority weights corresponding to the baseline data source and the data source to be aligned, to obtain the target dynamic window.
3. The method according to claim 1, characterized in that, The step of determining the data group weight corresponding to the candidate data group based on the candidate data group, the collected data group, and the target dynamic window includes: A first similarity value is determined based on the candidate data group and the collected values corresponding to the collected data group; The second similarity value is determined based on the candidate data group, the collection time corresponding to the collected data group, and the target dynamic window; Based on the business type corresponding to the target business, determine the first weight corresponding to the first similarity value and the second weight corresponding to the second similarity value; The weight of the data group corresponding to the candidate data group is determined by weighted summation of the first similarity value, the first weight, the second similarity value, and the second weight.
4. The method according to claim 1, characterized in that, After obtaining the historical data of the baseline data source and the data source to be aligned corresponding to the target service in the historical time period, wherein the historical data includes at least one data collection group, and the data collection group includes the collection value and the collection time corresponding to the collection value, the process further includes: Data missing detection is performed on the historical data collected from each data source in the data source set to determine the data missing type for each data source. The data source set includes a baseline data source and a data source to be aligned. For each data source, based on the data missing type corresponding to the data source, determine the data compensation rules, and update the missing data in the historical data collected by the data source according to the data compensation rules to obtain the compensation data corresponding to the missing data. Obtain the valid data range corresponding to the data source; The offset value is determined based on the effective data range and the compensation data; The compensation data is updated based on the offset value.
5. The method according to claim 4, characterized in that, The process of performing data missing detection on the historical data collected from each data source in the data source set, and determining the data missing type for each data source, includes: Identify the historical data collected by each data source in the data source set and determine at least one missing data for each data source; For each data source, obtain the missing time points corresponding to the missing data in each data source; Based on the missing time points corresponding to each of the missing data, a continuous detection is performed to determine at least one group of detection data; Obtain the number of time points corresponding to each of the aforementioned detection data groups; For each of the aforementioned detection data groups, the number of time points corresponding to the detection data group is compared with a preset threshold to obtain a threshold comparison result; When the threshold comparison result is that the number of time points corresponding to the detection data group is greater than or equal to the preset threshold, the data missing type of the missing data in the detection data group is determined to be the continuous missing type. When the threshold comparison result shows that the number of time points corresponding to the detection data group is less than the preset threshold, the data missing type of the missing data in the detection data group is determined to be a single point missing type.
6. The method according to claim 5, characterized in that, The process involves determining data compensation rules based on the data loss type corresponding to the data source, and updating the missing data in the historical data collected from the data source according to the data compensation rules to obtain the compensation data corresponding to the missing data, including: Retrieve historical data collected from the data source within a historical time period; For each detection data set, when the data missing type corresponding to the detection data set is continuous missing type, the historical data collected by the data source is identified based on the pre-trained time series data prediction model, and the compensation data corresponding to the missing time point of each missing data in the detection data set is predicted. When the data missing type corresponding to the detection data group is a single point missing type, the data change type is determined based on the historical data collected from the data source. Based on the data compensation rules corresponding to the data change type, compensation data is generated for the missing time points of the missing data in the detection data group.
7. The method according to claim 6, characterized in that, The step of generating compensation data corresponding to the missing time points of missing data in the detection data group according to the data compensation rules corresponding to the data change type includes: When the data change type is the first change type, the preceding data of the missing data in the detection data group is determined as the compensation data corresponding to the missing data; When the data change type is the second change type, the compensation data corresponding to the missing time point of the missing data is calculated by interpolation based on at least one adjacent data of the missing data in the detection data group.
8. A multi-source time-series data alignment device, characterized in that, The device includes: The data acquisition module is used to acquire historical data of the benchmark data source and the data source to be aligned for the target business in a historical time period. The historical data includes at least one data collection group, and the data collection group includes the collection value and the collection time corresponding to the collection value. The window determination module is used to determine the target dynamic window based on the historical data collected by the benchmark data source and the data source to be aligned. The data filtering module is used to filter the historical data corresponding to the data source to be aligned, based on the data group and the target dynamic window, to obtain at least one candidate data group corresponding to the data group. The weight determination module is used to determine the data group weight corresponding to each candidate data group based on the candidate data group, the collected data group, and the target dynamic window. The target group determination module is used to determine the candidate data group corresponding to the collected data group based on the candidate data group with the minimum weight among the weights of the data groups corresponding to each candidate data group, and to align the collected data group with the target data group in time.
9. A multi-source time-series data alignment device, characterized in that, The multi-source time-series data alignment device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the multi-source timing data alignment method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the multi-source timing data alignment method of any one of claims 1-7.