A method and system for continuous aggregation of time series data
Through the continuous aggregation time series data analysis method, a hierarchical sampling rate list and a meta-aggregation hierarchical sampling rate calculation task list are constructed, which solves the problem of insufficient processing speed of traditional data analysis methods and realizes efficient and real-time time series data analysis.
Patent Information
- Application Number
- CN202211437979.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-16
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-11-16
AI Technical Summary
In large systems, due to insufficient processing speed, traditional data analysis methods lack real-time and interactive query process of timing monitoring data, which is difficult to meet the timeliness needs of users.
The time series data analysis method of continuous aggregation is adopted to obtain the preset sampling rate list and the index sampling accuracy list, and a meta-aggregation hierarchical sampling rate calculation task list is constructed by obtaining the preset sampling rate list and the meta-algorithm of the preset aggregation algorithm. Then, a continuous aggregation strategy is constructed, a continuous aggregation timing task is generated, aggregation data is obtained, and a query analysis task is constructed through query conditions to obtain the complete query analysis results.
Through continuous incremental aggregation calculation of hierarchical sampling rates, the amount of data to be analyzed is significantly reduced, the speed of data query analysis is improved, and the accuracy and timeliness of data query analysis are ensured.
Smart Images

Figure CN116089489B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis, and in particular, to a method and system for continuous aggregation of time series data analysis. Background Art
[0002] With the rapid popularization and development of the Internet of Things technology, a large amount of time series monitoring data will be generated. Not only the data acquisition scale is increasing continuously, but also higher requirements are put forward for data analysis technology. For large systems, the amount of original data is relatively huge. Affected by the data scale and limited by the memory capacity, it is difficult to meet the timeliness requirements of users for querying, analyzing and calculating big data.
[0003] Traditional data analysis methods are based on the analysis and query of all original data. In such big data volume scenarios, the original data pulled is extremely huge, occupying a high amount of computing resources and taking a long time.
[0004] For ad-hoc queries and statistical data analysis tasks, it is required to complete the data query and give the query results within a given time. However, due to insufficient processing speed, the real-time performance and interactivity of the traditional data analysis methods are insufficient during the query process.
[0005] The above content is only used to assist in understanding the technical solution of the present invention, and does not represent an admission that the above content is prior art. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides a method for continuous aggregation of time series data analysis, including:
[0007] S1: Obtain a preset sampling rate list and an index sampling accuracy list of time series monitoring data, and construct a hierarchical sampling rate list through the preset sampling rate list and the index sampling accuracy list;
[0008] S2: Construct a meta-aggregation hierarchical sampling rate calculation task list through the hierarchical sampling rate list and the meta-algorithm of the preset aggregation algorithm;
[0009] S3: Construct a continuous aggregation strategy, generate a continuous aggregation timing task through the continuous aggregation strategy and the meta-aggregation hierarchical sampling rate calculation task list, and obtain aggregated data through the continuous aggregation timing task;
[0010] S4: Construct a query analysis task through query conditions, and execute the query analysis task on the aggregated data to obtain a complete query analysis result.
[0011] Preferably, step S1 is specifically:
[0012] S11: Group and construct an index sampling accuracy list according to the time interval during data collection of the index, and configure the time interval for index sampling according to actual business needs to group and construct a preset sampling rate list;
[0013] S12: Calculate the greatest common accuracy T of the sampling accuracies in the index sampling accuracy list, and retain the preset sampling rates greater than or equal to T in the preset sampling rate list to obtain a hierarchical sampling rate list.
[0014] Preferably, step S2 is specifically as follows:
[0015] S21: Construct a preset aggregation algorithm according to the aggregation element algorithm, and generate an aggregation element algorithm task classification i, where i represents the number of the aggregation element algorithm task classification;
[0016] S22: For each aggregation element algorithm task classification, construct calculation tasks in sequence according to each hierarchical sampling rate in the hierarchical sampling rate list to obtain a meta-aggregation hierarchical sampling rate calculation task list.
[0017] Preferably, step S3 is specifically as follows:
[0018] S31: Construct a continuous aggregation strategy according to different continuous aggregation calculation requirements, select corresponding calculation tasks from the meta-aggregation hierarchical sampling rate calculation task list through the continuous aggregation strategy, and construct continuous aggregation timing tasks through the selected calculation tasks;
[0019] S32: Execute the continuous aggregation timing tasks, and use the data of different meta-algorithms and different sampling rates output after execution as aggregation data.
[0020] Preferably, step S4 is specifically as follows:
[0021] S41: Use the maximum hierarchical sampling rate in the hierarchical sampling rate list that is less than the query accuracy as the best hierarchical sampling rate for the query analysis task;
[0022] S42: Find the aggregation element algorithm for the query analysis task according to the query conditions;
[0023] S43: Obtain the time range, and divide the time range through a preset time threshold to obtain a query time period sequence for the query analysis task;
[0024] S44: Construct a query analysis task through the best hierarchical sampling rate, aggregation element algorithm, and query time period sequence of the query analysis task;
[0025] S45: Execute the query analysis task on the aggregation data to obtain a complete query analysis result.
[0026] A time series data analysis system for continuous aggregation, including:
[0027] A hierarchical sampling rate list construction module, configured to obtain a preset sampling rate list and an index sampling accuracy list of time series monitoring data, and construct a hierarchical sampling rate list through the preset sampling rate list and the index sampling accuracy list;
[0028] A task list calculation module, configured to construct a meta-aggregation hierarchical sampling rate calculation task list through the hierarchical sampling rate list and a meta-algorithm of a preset aggregation algorithm;
[0029] An aggregated data acquisition module, configured to construct a continuous aggregation strategy, generate a continuous aggregation timing task through the continuous aggregation strategy and the meta-aggregation hierarchical sampling rate calculation task list, and obtain aggregated data through the continuous aggregation timing task;
[0030] An analysis result output module, configured to construct a query analysis task through query conditions, execute the query analysis task on the aggregated data, and obtain a complete query analysis result.
[0031] The present invention has the following beneficial effects:
[0032] The present invention adopts a method of continuous incremental aggregation calculation of hierarchical sampling rates. During the continuous acquisition of time series monitoring data, continuous incremental calculation is performed according to the preset to generate a hierarchical sampling rate list; in subsequent query analysis, the best hierarchical sampling rate is obtained according to the aggregation algorithm and sampling accuracy of the query conditions, and time series data analysis is performed in combination with the data in the latest time period; time series data analysis is performed by matching the best hierarchical sampling rate and the continuous aggregation strategy, so that the amount of data to be analyzed is reduced by multiples compared with the original data, greatly improving the speed of data query analysis and ensuring the accuracy and timeliness of data query analysis. Description of the Drawings
[0033] Figure 1 is a flowchart of the method of the embodiment of the present invention;
[0034] Figure 2 is a schematic diagram of a meta-aggregation hierarchical sampling rate calculation task list;
[0035] The implementation, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiment
[0036] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0037] Referring to Figure 1 , the present invention provides a method for time series data analysis with continuous aggregation, including:
[0038] S1: Obtain a preset sampling rate list and an index sampling accuracy list of time series monitoring data, and construct a hierarchical sampling rate list through the preset sampling rate list and the index sampling accuracy list;
[0039] S2: Construct a meta-aggregation hierarchical sampling rate calculation task list through a hierarchical sampling rate list and a meta-algorithm of a preset aggregation algorithm;
[0040] S3: Construct a continuous aggregation strategy, generate a continuous aggregation timing task through the continuous aggregation strategy and the meta-aggregation hierarchical sampling rate calculation task list, and obtain aggregated data through the continuous aggregation timing task;
[0041] S4: Construct a query analysis task through query conditions, execute the query analysis task on the aggregated data, and obtain a complete query analysis result.
[0042] Further, step S1 is specifically as follows:
[0043] S11: Group and construct an index sampling accuracy list according to the time interval of the index during data collection, and group and construct a preset sampling rate list according to the time interval configured for index sampling according to actual business needs;
[0044] Specifically, the index sampling accuracies of the index sampling accuracy list grouping are such as 1s, 10s, 1m, 10m, etc.;
[0045] The preset sampling rates of the preset sampling rate list grouping are such as 1m, 10m, 1h, 8h, 1d, etc.;
[0046] Among them, s is seconds, m is minutes, h is hours, and d is days; each time interval corresponds to a group and the time series monitoring data of the corresponding collection time interval is put in;
[0047] S12: Calculate and obtain the greatest common accuracy T of the sampling accuracies in the index sampling accuracy list, and retain the preset sampling rates greater than or equal to T in the preset sampling rate list to obtain a hierarchical sampling rate list;
[0048] Specifically, for example, if the sampling accuracies of the index sampling accuracy list are 2m and 10m, and the preset sampling rates of the preset sampling rate list are 1m, 10m, 1h, 8h, and 1d, then the greatest common accuracy T is 2m, and the hierarchical sampling rates of the hierarchical sampling rate list are 10m, 1h, 8h, 1d.
[0049] Further, step S2 is specifically as follows:
[0050] S21: Construct a preset aggregation algorithm according to the aggregation meta-algorithm, generate an aggregation meta-algorithm task classification i, where i represents the number of the aggregation meta-algorithm task classification;
[0051] Specifically, the aggregation meta-algorithm classifications include: maximum (max), minimum (min), sum (sum), count (ct.), average (avg), range (R), standard deviation (S), variance (S 2 ) etc.;
[0052] For the time series monitoring data sequence x1, x2, ……, x n , there is the following calculation formula:
[0053] max = max(x1, x2,... x n )
[0054] min = min(x1, x2,... x n )
[0055] sum = x1 + x2 +... + x n
[0056] ct. = n
[0057] avg = sum / ct.
[0058] R = max - min
[0059]
[0060]
[0061] The aggregation element algorithm is the basic aggregation algorithm used by the aggregation calculation method. For example, if the average value aggregation calculation needs to use the summation aggregation algorithm and the counting aggregation algorithm, then the summation aggregation algorithm and the counting aggregation algorithm are the meta-algorithms of the average value aggregation algorithm;
[0062] S22: For each aggregation element algorithm task classification, calculate tasks are sequentially constructed according to each level sampling rate in the level sampling rate list to obtain a meta-aggregation level sampling rate calculation task list.
[0063] Specifically, the meta-aggregation level sampling rate calculation task list is as Figure 2 shown. The level sampling rates in the level sampling rate list are: 10m, 1h, 8h, 1d, and the aggregation element algorithm classifications are maximum (max), minimum (min), counting (count);
[0064] Then the generated calculation tasks are: max_10m, max_1h, max_8h, max_1d; min_10m, min_1h, min_8h, min_1d; count_10m, count_1h, count_8h, count_1d.
[0065] Furthermore, step S3 is specifically as follows:
[0066] S31: Construct a continuous aggregation strategy according to different continuous aggregation calculation requirements. Select corresponding calculation tasks from the task list of sampling rate calculation at the meta-aggregation level through the continuous aggregation strategy, and construct a continuous aggregation timing task through the selected calculation tasks;
[0067] Specifically, continuous aggregation is a process of performing aggregation analysis on newly added data when data is continuously stored in the database by utilizing the continuous aggregation feature of the TimescaleDB database;
[0068] The continuous aggregation strategy is to create different TimescaleDB continuous aggregation timing tasks according to different continuous aggregation calculation requirements;
[0069] S32: Execute the continuous aggregation timing task, and use the data with different meta-algorithms and different sampling rates output after execution as the aggregated data.
[0070] Further, step S4 is specifically as follows:
[0071] S41: Take the hierarchical sampling rate in the hierarchical sampling rate list that is less than the maximum hierarchical sampling rate in the query accuracy as the optimal hierarchical sampling rate for the query analysis task;
[0072] Specifically, obtain the optimal hierarchical sampling rate according to the query accuracy in the query analysis conditions;
[0073] S42: Find the aggregated meta-algorithm for the query analysis task according to the query conditions;
[0074] Specifically, find the corresponding meta-algorithm according to the aggregation calculation in the query analysis conditions;
[0075] S43: Obtain the time range, and divide the time range through a preset time threshold to obtain the query time period sequence of the query analysis task;
[0076] Specifically, the time threshold is the time point preset by the system for dividing the time range, and the entire time range can be divided into a time period for querying aggregated data and a time period for querying raw data;
[0077] S44: Construct a query analysis task through the optimal hierarchical sampling rate, aggregated meta-algorithm, and query time period sequence of the query analysis task;
[0078] S45: Execute the query analysis task on the aggregated data to obtain the complete query analysis result.
[0079] An embodiment of the time series data analysis method for continuous aggregation according to the present invention is as follows:
[0080] The query indicator is the liquid level, the site number is 01, the indicator sampling precisions in the indicator sampling precision list are 2m and 10m, and the preset sampling rates in the preset sampling rate list are 1m, 10m, 1h, 8h, and 1d. Then, the hierarchical sampling rates in the calculated hierarchical sampling rate list are 10m, 1h, 8h, and 1d.
[0081] Assume the current time is 2022-10-13 14:00:00, and the query analysis conditions are as follows:
[0082] Query precision: 4h;
[0083] Aggregation calculation: maximum (max);
[0084] Time range: 2022-10-01 00:00:00 to 2022-10-31 23:59:59;
[0085] Aggregation meta-algorithm classification: maximum (max), minimum (min), count (count);
[0086] Hierarchical sampling rate: 10m, 1h, 8h, 1d;
[0087] Time threshold: 00:00:00 of the current day.
[0088] The query results are as follows:
[0089] The optimal hierarchical sampling rate is: 1h;
[0090] The aggregation meta-algorithm is: max;
[0091] The query data includes:
[0092] Query aggregated data from 2022-10-01 00:00:00 to 2022-10-12 23:59:59,
[0093] Query original data from 2022-10-13 00:00:00 to 2022-10-31 23:59:59.
[0094]
[0095]
[0096] The present invention provides a continuous aggregation time series data analysis system, including:
[0097] A hierarchical sampling rate list construction module, configured to obtain a preset sampling rate list and an indicator sampling precision list of time series monitoring data, and construct a hierarchical sampling rate list through the preset sampling rate list and the indicator sampling precision list;
[0098] A task list calculation module, configured to construct a meta-aggregation hierarchical sampling rate calculation task list through a hierarchical sampling rate list and a meta-algorithm of a preset aggregation algorithm;
[0099] An aggregated data acquisition module, configured to construct a continuous aggregation strategy, generate a continuous aggregation timing task through the continuous aggregation strategy and the meta-aggregation hierarchical sampling rate calculation task list, and obtain aggregated data through the continuous aggregation timing task;
[0100] An analysis result output module, configured to construct a query analysis task through query conditions, execute the query analysis task on the aggregated data, and obtain a complete query analysis result.
[0101] It should be noted that in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or system. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or system including that element.
[0102] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments. Among the several apparatus unit claims listing several apparatuses, several of these apparatuses may be embodied by the same hardware item. The use of the words first, second, and third, etc. does not indicate any order and these words can be interpreted as identifiers.
[0103] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for analyzing time series data of continuous aggregation, characterized in that, Including: S1: Obtain the preset sampling rate list and index sampling accuracy list of the time series monitoring data, and construct a hierarchical sampling rate list through the preset sampling rate list and index sampling accuracy list; S2: Construct a meta-aggregation hierarchical sampling rate calculation task list through the hierarchical sampling rate list and the meta-algorithm of the preset aggregation algorithm; S3: Construct a continuous aggregation strategy, generate a continuous aggregation timing task through the continuous aggregation strategy and the meta-aggregation hierarchical sampling rate calculation task list, and obtain aggregated data through the continuous aggregation timing task; S4: Construct a query analysis task through the query conditions, execute the query analysis task on the aggregated data, and obtain a complete query analysis result; Among them, step S1 is specifically: S11: Group and construct the index sampling accuracy list according to the time interval of the index during data collection, and group and construct the preset sampling rate list according to the time interval configured for index sampling according to actual business needs; S12: Calculate the greatest common accuracy T of the sampling accuracies in the index sampling accuracy list, and retain the preset sampling rates greater than or equal to T in the preset sampling rate list to obtain the hierarchical sampling rate list; Step S4 is specifically: S41: Take the hierarchical sampling rate less than the maximum hierarchical sampling rate in the query accuracy in the hierarchical sampling rate list as the optimal hierarchical sampling rate of the query analysis task; S42: Find the aggregation meta-algorithm of the query analysis task according to the query conditions; S43: Obtain the time range, and split the time range through the preset time threshold to obtain the query time period sequence of the query analysis task; S44: Construct a query analysis task through the optimal hierarchical sampling rate, aggregation meta-algorithm and query time period sequence of the query analysis task; S45: Execute the query analysis task on the aggregated data to obtain a complete query analysis result.
2. The method for analyzing time series data of continuous polymerization according to claim 1, wherein Step S2 is specifically: S21: Construct a preset aggregation algorithm according to the aggregation meta-algorithm, and generate the aggregation meta-algorithm task classification i, where i represents the number of the aggregation meta-algorithm task classification; S22: For each aggregation meta-algorithm task classification, construct calculation tasks in turn according to each hierarchical sampling rate in the hierarchical sampling rate list to obtain the meta-aggregation hierarchical sampling rate calculation task list.
3. The method for analyzing time series data of continuous polymerization according to claim 1, wherein Step S3 is specifically: S31: Construct a continuous aggregation strategy according to different continuous aggregation calculation requirements, select the corresponding calculation tasks in the meta-aggregation hierarchical sampling rate calculation task list through the continuous aggregation strategy, and construct a continuous aggregation timing task through the selected calculation tasks; S32: Execute the continuous aggregation timing task, and use the data of different meta-algorithms and different sampling rates output after execution as the aggregated data.
4. A time series data analysis system for continuous aggregation, characterized in that, It is used to implement the continuous aggregation time series data analysis method described in any one of claims 1 to 3.
Citation Information
Patent Citations
Methods and systems for detection in industrial internet of things data collection environment with large data sets
CN110073301A
Rapid mass time-series data processing method based on aggregated edge and time-series aggregated edge
WO2021134318A1