Method for dynamically determining data scanning and calculation in time sequence database

By dynamically calculating parallelism and resource allocation, and using real-time statistical information to optimize data scanning and calculation of timing databases, the limitations of timing databases in dynamic processing are solved, and the performance of processing large-scale data queries and system stability are significantly improved.

CN119988427APending Publication Date: 2025-05-13上海沄熹科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510061794.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Time series databases have limitations in dynamically determining data scanning and computing, lack special optimization of data characteristics, cannot fully utilize the inherent properties of time series data, and real-time statistics and parallel processing strategies cannot cope with dynamic changes in system load and available resources.

Method used

By collecting real-time statistical information, the average data volume of the entity is calculated, and the query data volume is dynamically estimated based on the data compression ratio, data distribution and the size of the query window. According to the complexity of the query, the system's available resources, the query time range, the number of entities and the amount of entity data, the optimal parallelism is dynamically calculated, and the system resource usage, query execution time and load status are monitored in real time, and the parallelism and resource allocation strategy are dynamically adjusted using the feedback mechanism.

Benefits of technology

It significantly improves the performance when handling large-scale data queries, enhances the adaptability to complex business scenarios and variable system loads, reduces the impact of data skew problems, improves the overall efficiency and execution speed of queries, and improves the stability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988427A_ABST
    Figure CN119988427A_ABST
Patent Text Reader

Abstract

The invention discloses a method for dynamically determining data scanning and calculation in a time sequence database, and relates to the technical field of time sequence database processing. Comprising the steps that 1, real-time statistical information is collected, and the statistical information comprises entity data volume distribution, a query time range and maximum and minimum timestamp information in a database; 2, calculating the average data volume of the entities based on the collected statistical information, and dynamically estimating the queried data volume in combination with a data compression ratio, data distribution and the size of a query window; 3, calculating the optimal parallelism degree according to the query complexity, the available resources of the system, the query time range, the quantity of entities and the quantity of entity data, and 4, allocating data processing tasks to calculation nodes or threads according to the calculated parallelism degree, and executing parallel data processing; and 5, monitoring the use condition of system resources in real time, querying execution time and a load state, and dynamically adjusting the degree of parallelism and a resource allocation strategy by utilizing a feedback mechanism according to a monitoring result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention discloses a method for dynamically determining data scanning and calculation in a time series database, and relates to the technical field of time series database processing. Background Art

[0002] Parallel processing methods are already quite mature in traditional relational databases. The parallel processing of time series databases is significantly different from that of traditional relational databases. Therefore, time series databases still have some significant limitations in dynamically determining data scanning and calculation. The main issues involved include:

[0003] Lack of data specificity considerations: The data in time series databases have a high degree of time correlation and pattern predictability, such as periodic fluctuations, data surges due to emergencies, etc. Existing technologies often adopt generalized data management and computing strategies, lack of specialized optimization for these characteristics, and cannot fully utilize the inherent properties of time series data to optimize data processing.

[0004] Limitations of real-time statistics and parallel processing strategies: The degree of parallelism is usually determined by system parameters or statistics. However, these methods rely on non-real-time statistical information data and cannot cope with dynamic changes in system load and available resources. For example, some systems set a static degree of parallelism based on historical data volume and query mode before query execution, or rely on fixed system parameters for resource allocation. However, when the amount of data increases dramatically, the query complexity increases, or the system concurrent load changes, static parallel processing strategies may lead to unbalanced resource allocation, reduce system efficiency and increase response time. Summary of the invention

[0005] In view of the problems of the prior art, the present invention provides a method for dynamically determining data scanning and calculation in a time series database. The method is suitable for scenarios that require efficient processing of large-scale time series data and optimization of query performance, and has important application value in data-intensive and query-frequent industrial, financial and Internet of Things applications.

[0006] The specific scheme proposed by the present invention is:

[0007] The present invention also provides a method for dynamically determining data scanning and calculation in a time series database, comprising:

[0008] Step 1: Collect real-time statistics, including entity data volume distribution, query time range, and maximum and minimum timestamp information in the database;

[0009] Step 2: Calculate the average data volume of the entity based on the collected statistical information, and dynamically estimate the query data volume based on the data compression ratio, data distribution, and query window size;

[0010] Step 3: Based on the complexity of the query, system available resources, query time range, number of entities, and amount of entity data, use the following formula:

[0011]

[0012] Calculate the optimal degree of parallelism, where entityCount represents the number of entities that meet the conditions, entityDataVolume represents the average amount of data for each entity, queryTimeSpan represents the time range of the timestamp field involved in the query, groupCount represents the number of groups after pre-grouping, maxData Timestamp / minData Timestam represents the earliest / latest timestamp recorded in the database, and constant represents the total amount of data processed by each parallel task, which is a constant coefficient.

[0013] Step 4: According to the calculated degree of parallelism, the data processing tasks are assigned to computing nodes or threads to perform parallel data processing;

[0014] Step 5: Monitor the usage of system resources, query execution time, and load status in real time, and use the feedback mechanism to dynamically adjust the parallelism and resource allocation strategy based on the monitoring results.

[0015] Furthermore, in step 1 of the method for dynamically determining data scanning and calculation in a time series database, the collected real-time statistical information is analyzed, including:

[0016] Analyze the horizontal and vertical distribution of entity data volume, including collecting entity data distribution in different time intervals to evaluate the total amount of data between different entities, and also to evaluate the uniformity of data in each entity within the time range, to avoid unreasonable parallelism distribution caused by horizontal or vertical data tilt.

[0017] Analyze the span of the query time range, including the time range involved in the query, and also judge whether the query is concentrated on certain high-frequency data points by the distribution density of the sampled data.

[0018] Analyze the overall data situation of the database, use the maximum and minimum timestamps to calculate the time span, and dynamically adjust the parallelism for cold data and hot data in combination with the data life cycle management strategy.

[0019] Furthermore, step 2 of the method for dynamically determining data scanning and calculation in a time series database specifically includes:

[0020] The average data volume of the entity is obtained by analyzing the ratio of the number of rows to the number of key static data records, wherein the data volume estimation in the dynamic window is combined: the ratio is dynamically adjusted according to the size of the time window;

[0021] And consider the data compression ratio: infer the data compression ratio of the entity based on statistical information, thereby determining the data decompression cost and the actual data volume.

[0022] Furthermore, in step 3 of the method for dynamically determining data scanning and calculation in a time series database, optimal parallelism adjustment is also performed, including:

[0023] Correcting data skew: If the amount of data of some entities is significantly greater than the average value in the statistical information, the parallelism of these entities is adjusted in a weighted manner, and the data skew correction factor skewFactor is added to the original formula. The formula is:

[0024] Degree=Degree×(1+skewFactor)

[0025] The skewFactor is calculated by analyzing the deviation of each entity’s data distribution from the average data volume;

[0026] Perform resource restrictions and load forecasting: Combine historical load and load forecasting mechanisms to reserve certain system resources to cope with load fluctuations, dynamically adjust the upper and lower limits of availableThreads, and avoid high load peaks in a short period of time that may lead to system resource exhaustion.

[0027] Correct the data column width and type: Queries involving arrays or objects will take extra decoding and calculation time. In this case, the weighted parallelism coefficient columnComplexityFactor is calculated based on the column type and data volume. The formula for adjusting the parallelism is as follows:

[0028] Degree=Degree×columnComplexityFactor

[0029] The columnComplexityFactor is dynamically calculated by analyzing the processing cost of the data column types involved in the query.

[0030] Furthermore, step 5 of the method for dynamically determining data scanning and calculation in a time series database specifically includes:

[0031] Monitor the usage of system resources: Monitor resource utilization, which includes CPU, memory, and I / O usage. Dynamically adjust the values ​​of availableThreads and constant according to the fluctuation of resource utilization.

[0032] Query execution time: Based on the query execution time at different parallelism levels, a feedback mechanism is used to dynamically adjust the parallelism of subsequent queries.

[0033] Managing fault tolerance and recovery: When failures of some compute nodes in a parallel task are detected, the parallelism is adjusted dynamically and the unfinished tasks are reallocated to other healthy nodes.

[0034] The present invention also provides a device for dynamically determining data scanning and calculation in a time series database, comprising a collection module, a calculation module, an execution module and a monitoring and adjustment module.

[0035] The collection module collects real-time statistical information, including entity data volume distribution, query time range, and maximum and minimum timestamp information in the database;

[0036] The calculation module calculates the average data volume of the entity based on the collected statistical information, and dynamically estimates the query data volume based on the data compression ratio, data distribution, and the size of the query window;

[0037] The calculation module uses the following formula based on the complexity of the query, the available system resources, the query time range, the number of entities, and the amount of entity data:

[0038]

[0039] Calculate the optimal degree of parallelism, where entityCount represents the number of entities that meet the conditions, entityDataVolume represents the average amount of data for each entity, queryTimeSpan represents the time range of the timestamp field involved in the query, groupCount represents the number of groups after pre-grouping, maxData Timestamp / minData Timestam represents the earliest / latest timestamp recorded in the database, and constant represents the total amount of data processed by each parallel task, which is a constant coefficient.

[0040] The execution module assigns data processing tasks to computing nodes or threads according to the calculated parallelism and performs parallel data processing;

[0041] The monitoring and adjustment module monitors the usage of system resources, query execution time and load status in real time, and dynamically adjusts the parallelism and resource allocation strategy based on the monitoring results using the feedback mechanism.

[0042] Furthermore, the collection module of the device for dynamically determining data scanning and calculation in a time series database also analyzes the collected real-time statistical information, including:

[0043] Analyze the horizontal and vertical distribution of entity data volume, including collecting entity data distribution in different time intervals to evaluate the total amount of data between different entities, and also to evaluate the uniformity of data in each entity within the time range, to avoid unreasonable parallelism distribution caused by horizontal or vertical data tilt.

[0044] Analyze the span of the query time range, including the time range involved in the query, and also judge whether the query is concentrated on certain high-frequency data points by the distribution density of the sampled data.

[0045] Analyze the overall data situation of the database, use the maximum and minimum timestamps to calculate the time span, and dynamically adjust the parallelism for cold data and hot data in combination with the data life cycle management strategy.

[0046] Further, the computing module of the device for dynamically determining data scanning and computing in a time series database obtains the average data volume of the entity by analyzing the ratio of the number of rows to the number of key static data records, wherein the ratio is dynamically adjusted according to the size of the time window in combination with the estimation of the data volume in the dynamic window;

[0047] And consider the data compression ratio: infer the data compression ratio of the entity based on statistical information, thereby determining the data decompression cost and the actual data volume.

[0048] Furthermore, the device computing module for dynamically determining data scanning and computing in a time series database performs optimal parallelism adjustment, including:

[0049] Correcting data skew: If the amount of data of some entities is significantly greater than the average value in the statistical information, the parallelism of these entities is adjusted in a weighted manner, and the data skew correction factor skewFactor is added to the original formula. The formula is:

[0050] Degree=Degree×(1+skewFactor)

[0051] The skewFactor is calculated by analyzing the deviation of each entity’s data distribution from the average data volume;

[0052] Perform resource restrictions and load forecasting: Combine historical load and load forecasting mechanisms to reserve certain system resources to cope with load fluctuations, dynamically adjust the upper and lower limits of availableThreads, and avoid high load peaks in a short period of time that may lead to system resource exhaustion.

[0053] Correct the data column width and type: Queries involving arrays or objects will take extra decoding and calculation time. In this case, the weighted parallelism coefficient columnComplexityFactor is calculated based on the column type and data volume. The formula for adjusting the parallelism is as follows:

[0054] Degree=Degree×columnComplexityFactor

[0055] The columnComplexityFactor is dynamically calculated by analyzing the processing cost of the data column types involved in the query.

[0056] Furthermore, the monitoring and adjustment module of the device for dynamically determining data scanning and calculation in a time series database monitors the usage of system resources: monitors resource utilization, which includes the usage of CPU, memory and I / O, and dynamically adjusts the values ​​of availableThreads and constant according to the fluctuation of resource utilization,

[0057] Query execution time: Based on the query execution time at different parallelism levels, a feedback mechanism is used to dynamically adjust the parallelism of subsequent queries.

[0058] Managing fault tolerance and recovery: When failures of some compute nodes in a parallel task are detected, the parallelism is adjusted dynamically and the unfinished tasks are reallocated to other healthy nodes.

[0059] The benefits of the present invention are:

[0060] The present invention not only significantly improves the performance of processing large-scale data queries by dynamically adjusting the data scanning and calculation parallelism in the time series database, but also provides powerful adaptability for coping with complex business scenarios and changing system loads.

[0061] First, based on real-time statistics (RTS) and optimizer strategies, it can intelligently allocate parallel tasks according to query complexity, data distribution, and the availability of system resources, avoiding resource waste and performance bottlenecks caused by static allocation in traditional methods. This dynamic adjustment mechanism greatly reduces the impact of data skew problems, thereby improving the overall efficiency and execution speed of queries.

[0062] Secondly, by real-time monitoring of resource usage, the allocation of computing resources can be flexibly adjusted according to load changes to ensure optimal resource utilization under high-load environments. Compared with traditional fixed parallelism solutions, this invention greatly improves the efficiency of CPU, memory, I / O and other resources, and can significantly shorten query response time. In particular, when processing a large number of concurrent queries or complex queries, the system can maintain stable high-performance output.

[0063] In addition, the present invention also has strong fault tolerance. When a computing node failure, resource bottleneck or task execution abnormality is detected, the parallelism can be automatically adjusted and tasks can be dynamically reallocated to avoid system crashes or query failures. This adaptive fault tolerance mechanism not only improves the stability and reliability of the system, but also greatly reduces maintenance costs and improves user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 It is a schematic flow chart of the method of the present invention. DETAILED DESCRIPTION

[0065] Data parallelism refers to the extent to which a task in a database operation is divided into multiple subtasks and these subtasks can be executed simultaneously on different processing units.

[0066] The data in a time series database has some notable characteristics: first, the amount of data is huge and its total amount grows rapidly over time, but the time series data rarely changes; second, the growth of time series measurement data has certain rules, especially in an environment with a fixed collection frequency; third, the static data in the time series database (such as attribute data or label information) changes less, and the main data changes come from the measurement data in the time series, and the measurement data can usually be horizontally separated from entities that can be distinguished by static data, that is, a group of static data defines a logical or physical entity, and the entity has corresponding time series data collected and stored, and similarly other entities also have corresponding time series data collected and stored, and the data between entities are relatively independent; third, the frequency of query and use of the measurement data of each entity varies over time. The older the data in the time series database, the less frequently it is used. Therefore, in terms of storage efficiency, it will naturally be compressed and stored.

[0067] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.

[0068] Example 1

[0069] The present invention also provides a method for dynamically determining data scanning and calculation in a time series database, comprising:

[0070] Step 1: Collect real-time statistics, including entity data volume distribution, query time range, and maximum and minimum timestamp information in the database.

[0071] In step 1, the collected real-time statistical information is also analyzed, including:

[0072] Analyze the horizontal and vertical distribution of entity data volume, including collecting entity data distribution in different time intervals to evaluate the total amount of data between different entities, and also to evaluate the uniformity of data in each entity within the time range, to avoid unreasonable parallelism distribution caused by horizontal or vertical data skew. There may be horizontal data skew between different entities, and there may be vertical data skew in different time periods in a single entity.

[0073] Analyze the span of the query time range, including the time range involved in the query, and also determine whether the query is concentrated on certain high-frequency data points by the distribution density of the sampled data. For example, if the query has more data access in a specific time period, it may be necessary to increase the parallelism in that time period to balance the load.

[0074] Analyze the overall data situation of the database, using the maximum and minimum timestamps to calculate the time span, and dynamically adjust the parallelism for cold data and hot data in combination with the data lifecycle management strategy. For example, in a database architecture with cold data and hot data tiers, data at different levels have huge differences in storage media and access costs. The system should dynamically adjust the parallelism strategy based on the storage tier where the data is located. Hot data may require a higher degree of parallelism, while cold data can be appropriately reduced to save resources.

[0075] Step 2: Calculate the average data volume of the entity based on the collected statistical information, and dynamically estimate the query data volume based on the data compression ratio, data distribution, and query window size.

[0076] These may further include:

[0077] The average data volume of the entity is obtained by analyzing the ratio of the number of rows to the number of key static data records, which is combined with the estimation of the data volume in the dynamic window: the ratio is dynamically adjusted according to the size of the time window. For example, a larger query time window may contain more historical data, and this data may be in different storage media (such as disk or memory), affecting the speed of data access and parallel allocation.

[0078] And consider the data compression ratio: infer the data compression ratio of the entity based on statistical information, thereby determining the data decompression cost and the actual data volume. A higher compression ratio may mean lower I / O overhead, allowing more parallel tasks to be assigned to such entities.

[0079] Step 3: Based on the complexity of the query, system available resources, query time range, number of entities, and amount of entity data, use the following formula:

[0080]

[0081] Calculate the optimal degree of parallelism, where entityCount represents the number of entities that meet the conditions, the total number of entities involved in the database query that meet the specific conditions. Entities are defined based on one or more key static attribute data columns, which together identify unique entities in the database.

[0082] entityDataVolume represents the average data volume of each entity.

[0083] queryTimeSpan indicates the time range of the timestamp field involved in the query, that is, from the earliest to the latest timestamp. This parameter determines the size of the query time window, which directly affects the amount of data and query complexity.

[0084] groupCount indicates the number of groups after pre-grouping. By organizing data with the same key value together, parallel computing and data aggregation can be performed more efficiently, especially when processing complex aggregate queries, which can significantly improve performance and response speed.

[0085] maxData Timestamp / minData Timestam represents the earliest / latest timestamp recorded in the database. These two parameters define the time span of the data in the database.

[0086] constant represents the total amount of data processed by each parallel task, which is a constant coefficient, such as 1000000. This constant helps the system maintain the scale of parallel processing tasks and may need to be adjusted in different systems and data environments to maintain optimal performance.

[0087] In step 3, the optimal parallelism adjustment is also performed, including:

[0088] Correcting data skew: If the amount of data of some entities is significantly greater than the average value in the statistical information, the parallelism of these entities is adjusted in a weighted manner, and the data skew correction factor skewFactor is added to the original formula. The formula is:

[0089] Degree = Degree × (1 + skewFactor) SkewFactor is calculated by analyzing the deviation of the data distribution of each entity from the average data volume;

[0090] Perform resource restrictions and load prediction: Combine historical load and load prediction mechanisms to reserve certain system resources to cope with load fluctuations, dynamically adjust the upper and lower limits of availableThreads, and avoid high load peaks in a short period of time that may lead to system resource exhaustion. In addition, dynamic allocation of I / O resources can also be considered, such as disk read and write rates and network bandwidth, and give priority to frequently accessed data in high-parallelism scenarios.

[0091] Correct the data column width and type: Queries involving complex data types, such as arrays or objects, will take extra decoding and calculation time. In this case, the weighted parallelism coefficient columnComplexityFactor is calculated based on the column type and data volume. The formula for adjusting the parallelism is as follows:

[0092] Degree=Degree×columnComplexityFactor

[0093] The columnComplexityFactor is calculated dynamically by analyzing the processing cost of the data column types involved in the query, such as floating point numbers, strings, or JSON objects.

[0094] Step 4: Based on the calculated degree of parallelism, data processing tasks are assigned to computing nodes or threads to perform parallel data processing; the straw model of parallel processing mechanism can be used to maximize resource utilization and response speed. This step ensures that each query achieves the best execution efficiency based on available resources and data characteristics.

[0095] Step 5: Monitor the usage of system resources, query execution time, and load status in real time, and use the feedback mechanism to dynamically adjust the parallelism and resource allocation strategy based on the monitoring results.

[0096] Wherein step 5 may further specifically include:

[0097] Monitor the usage of system resources: Monitor resource utilization, which includes CPU, memory, and I / O usage. Dynamically adjust the values ​​of availableThreads and constant according to the fluctuation of resource utilization.

[0098] Query execution time: Based on the query execution time at different parallelism levels, a feedback mechanism is used to dynamically adjust the parallelism of subsequent queries. These subsequent queries are not limited to queries with similar structures, but are optimized and adjusted based on factors such as the objects accessed, the computing mode, and the available resources of the system. Even if the queries are written differently, as long as the resource consumption patterns or data access paths involved in their execution are similar, the parallelism of these queries can be more accurately tuned using previous execution information. For example, for queries with long execution times, the complexity of parallel tasks can be reduced by narrowing the query window or adding data preprocessing steps.

[0099] Managing fault tolerance and recovery: When failures of some compute nodes in a parallel task are detected, the parallelism is adjusted dynamically and the unfinished tasks are reallocated to other healthy nodes.

[0100] Example 2

[0101] The present invention also provides a device for dynamically determining data scanning and calculation in a time series database, comprising a collection module, a calculation module, an execution module and a monitoring and adjustment module.

[0102] The collection module collects real-time statistical information, including entity data volume distribution, query time range, and maximum and minimum timestamp information in the database;

[0103] The calculation module calculates the average data volume of the entity based on the collected statistical information, and dynamically estimates the query data volume based on the data compression ratio, data distribution, and the size of the query window;

[0104] The calculation module uses the following formula based on the complexity of the query, the available system resources, the query time range, the number of entities, and the amount of entity data:

[0105]

[0106] Calculate the optimal degree of parallelism, where entityCount represents the number of entities that meet the conditions, entityDataVolume represents the average amount of data for each entity, queryTimeSpan represents the time range of the timestamp field involved in the query, groupCount represents the number of groups after pre-grouping, maxData Timestamp / minData Timestam represents the earliest / latest timestamp recorded in the database, and constant represents the total amount of data processed by each parallel task, which is a constant coefficient.

[0107] The execution module assigns data processing tasks to computing nodes or threads according to the calculated parallelism and performs parallel data processing;

[0108] The monitoring and adjustment module monitors the usage of system resources, query execution time and load status in real time, and dynamically adjusts the parallelism and resource allocation strategy based on the monitoring results using the feedback mechanism.

[0109] As the information interaction and execution process between the modules of the above-mentioned device are based on the same concept as the embodiment of the method of the present invention, the specific contents can be found in the description of the embodiment of the method of the present invention and will not be repeated here.

[0110] Similarly, the device of the present invention not only significantly improves the performance of processing large-scale data queries by dynamically adjusting the data scanning and calculation parallelism in the time series database, but also provides powerful adaptability for coping with complex business scenarios and changing system loads.

[0111] First, based on real-time statistics (RTS) and optimizer strategies, it can intelligently allocate parallel tasks according to query complexity, data distribution, and the availability of system resources, avoiding resource waste and performance bottlenecks caused by static allocation in traditional methods. This dynamic adjustment mechanism greatly reduces the impact of data skew problems, thereby improving the overall efficiency and execution speed of queries.

[0112] Secondly, by real-time monitoring of resource usage, the allocation of computing resources can be flexibly adjusted according to load changes to ensure optimal resource utilization under high-load environments. Compared with traditional fixed parallelism solutions, this invention greatly improves the efficiency of CPU, memory, I / O and other resources, and can significantly shorten query response time. In particular, when processing a large number of concurrent queries or complex queries, the system can maintain stable high-performance output.

[0113] In addition, the present invention also has strong fault tolerance. When a computing node failure, resource bottleneck or task execution abnormality is detected, the parallelism can be automatically adjusted and tasks can be dynamically reallocated to avoid system crashes or query failures. This adaptive fault tolerance mechanism not only improves the stability and reliability of the system, but also greatly reduces maintenance costs and improves user experience.

[0114] It should be noted that not all steps and modules in the above-mentioned processes and device structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted as needed. The system structure described in the above-mentioned embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or some components in multiple independent devices may be implemented together.

[0115] The above-described embodiments are only preferred embodiments for fully illustrating the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or changes made by those skilled in the art based on the present invention are within the protection scope of the present invention. The protection scope of the present invention shall be subject to the claims.

Claims

1. A method for dynamically determining data scanning and calculation in a time series database, characterized by: include: Step 1: Collect real-time statistics, including entity data volume distribution, query time range, and maximum and minimum timestamp information in the database; Step 2: Calculate the average data volume of the entity based on the collected statistical information, and dynamically estimate the query data volume based on the data compression ratio, data distribution, and query window size; Step 3: Based on the complexity of the query, system available resources, query time range, number of entities, and amount of entity data, use the following formula: Calculate the optimal degree of parallelism, where entityCount represents the number of entities that meet the conditions, entityDataVolume represents the average amount of data for each entity, queryTimeSpan represents the time range of the timestamp field involved in the query, groupCount represents the number of groups after pre-grouping, maxData Timestamp / minData Timestam represents the earliest / latest timestamp recorded in the database, and constant represents the total amount of data processed by each parallel task, which is a constant coefficient. Step 4: According to the calculated degree of parallelism, the data processing tasks are assigned to computing nodes or threads to perform parallel data processing; Step 5: Monitor the usage of system resources, query execution time, and load status in real time, and use the feedback mechanism to dynamically adjust the parallelism and resource allocation strategy based on the monitoring results.

2. The method for dynamically determining data scanning and calculation in a time series database according to claim 1, characterized in that In step 1, real-time statistics collected are also analyzed, including: Analyze the horizontal and vertical distribution of entity data volume, including collecting entity data distribution in different time intervals to evaluate the total amount of data between different entities, and also to evaluate the uniformity of data in each entity within the time range, to avoid unreasonable parallelism distribution caused by horizontal or vertical data tilt. Analyze the span of the query time range, including the time range involved in the query, and also judge whether the query is concentrated on certain high-frequency data points by the distribution density of the sampled data. Analyze the overall data situation of the database, use the maximum and minimum timestamps to calculate the time span, and dynamically adjust the parallelism for cold data and hot data in combination with the data life cycle management strategy.

3. The method for dynamically determining data scanning and calculation in a time series database according to claim 1, characterized in that Step 2 specifically includes: The average data volume of the entity is obtained by analyzing the ratio of the number of rows to the number of key static data records, wherein the data volume estimation in the dynamic window is combined: the ratio is dynamically adjusted according to the size of the time window; And consider the data compression ratio: infer the data compression ratio of the entity based on statistical information, thereby determining the data decompression cost and the actual data volume.

4. The method for dynamically determining data scanning and calculation in a time series database according to claim 1, characterized in that In step 3, the optimal parallelism adjustment is also performed, including: Correcting data skew: If the amount of data of some entities is significantly greater than the average value in the statistical information, the parallelism of these entities is adjusted in a weighted manner, and the data skew correction factor skewFactor is added to the original formula. The formula is: Degree=Degree×(1+skewFactor) The skewFactor is calculated by analyzing the deviation of each entity’s data distribution from the average data volume; Perform resource restrictions and load forecasting: Combine historical load and load forecasting mechanisms to reserve certain system resources to cope with load fluctuations, dynamically adjust the upper and lower limits of availableThreads, and avoid high load peaks in a short period of time that may lead to system resource exhaustion. Correct the data column width and type: Queries involving arrays or objects will take extra decoding and calculation time. In this case, the weighted parallelism coefficient columnComplexityFactor is calculated based on the column type and data volume. The formula for adjusting the parallelism is as follows: Degree=Degree×columnComplexityFactor The columnComplexityFactor is dynamically calculated by analyzing the processing cost of the data column types involved in the query.

5. The method for dynamically determining data scanning and calculation in a time series database according to claim 1, characterized in that Step 5 specifically includes: Monitor the usage of system resources: Monitor resource utilization, which includes CPU, memory, and I / O usage. Dynamically adjust the values ​​of availableThreads and constant according to the fluctuation of resource utilization. Query execution time: Based on the query execution time at different parallelism levels, a feedback mechanism is used to dynamically adjust the parallelism of subsequent queries. Managing fault tolerance and recovery: When failures of some compute nodes in a parallel task are detected, the parallelism is adjusted dynamically and the unfinished tasks are reallocated to other healthy nodes.

6. A device for dynamically determining data scanning and calculation in a time series database, characterized in that It includes collection module, calculation module, execution module and monitoring and adjustment module. The collection module collects real-time statistical information, including entity data volume distribution, query time range, and maximum and minimum timestamp information in the database; The calculation module calculates the average data volume of the entity based on the collected statistical information, and dynamically estimates the query data volume based on the data compression ratio, data distribution, and the size of the query window; The calculation module uses the following formula based on the complexity of the query, the available system resources, the query time range, the number of entities, and the amount of entity data: Calculate the optimal degree of parallelism, where entityCount represents the number of entities that meet the conditions, entityDataVolume represents the average amount of data for each entity, queryTimeSpan represents the time range of the timestamp field involved in the query, groupCount represents the number of groups after pre-grouping, maxData Timestamp / minData Timestam represents the earliest / latest timestamp recorded in the database, and constant represents the total amount of data processed by each parallel task, which is a constant coefficient. The execution module assigns data processing tasks to computing nodes or threads according to the calculated parallelism and performs parallel data processing; The monitoring and adjustment module monitors the usage of system resources, query execution time and load status in real time, and dynamically adjusts the parallelism and resource allocation strategy based on the monitoring results using the feedback mechanism.

7. The device for dynamically determining data scanning and calculation in a time series database according to claim 6, characterized in that The collection module also analyzes the collected real-time statistics, including: Analyze the horizontal and vertical distribution of entity data volume, including collecting entity data distribution in different time intervals to evaluate the total amount of data between different entities, and also to evaluate the uniformity of data in each entity within the time range, to avoid unreasonable parallelism distribution caused by horizontal or vertical data tilt. Analyze the span of the query time range, including the time range involved in the query, and also judge whether the query is concentrated on certain high-frequency data points by the distribution density of the sampled data. Analyze the overall data situation of the database, use the maximum and minimum timestamps to calculate the time span, and dynamically adjust the parallelism for cold data and hot data in combination with the data life cycle management strategy.

8. The device for dynamically determining data scanning and calculation in a time series database according to claim 6, characterized in that The calculation module obtains the average data volume of the entity by analyzing the ratio of the number of rows to the number of key static data records, wherein the data volume estimation in the dynamic window is combined: the ratio is dynamically adjusted according to the size of the time window; And consider the data compression ratio: infer the data compression ratio of the entity based on statistical information, thereby determining the data decompression cost and the actual data volume.

9. The device for dynamically determining data scanning and calculation in a time series database according to claim 6, characterized in that The computing module performs optimal parallelism adjustment, including: Correcting data skew: If the amount of data of some entities is significantly greater than the average value in the statistical information, the parallelism of these entities is adjusted in a weighted manner, and the data skew correction factor skewFactor is added to the original formula. The formula is: Degree=Degree×(1+skewFactor) The skewFactor is calculated by analyzing the deviation of each entity’s data distribution from the average data volume; Perform resource restrictions and load forecasting: Combine historical load and load forecasting mechanisms to reserve certain system resources to cope with load fluctuations, dynamically adjust the upper and lower limits of availableThreads, and avoid high load peaks in a short period of time that may lead to system resource exhaustion. Correct the data column width and type: Queries involving arrays or objects will take extra decoding and calculation time. In this case, the weighted parallelism coefficient columnComplexityFactor is calculated based on the column type and data volume. The formula for adjusting the parallelism is as follows: Degree=Degree×columnComplexityFactor The columnComplexityFactor is dynamically calculated by analyzing the processing cost of the data column types involved in the query.

10. The device for dynamically determining data scanning and calculation in a time series database according to claim 5, characterized in that The monitoring and adjustment module monitors the usage of system resources: monitors resource utilization, which includes CPU, memory, and I / O usage. According to the fluctuation of resource utilization, the values ​​of availableThreads and constant are dynamically adjusted. Query execution time: Based on the query execution time at different parallelism levels, a feedback mechanism is used to dynamically adjust the parallelism of subsequent queries. Managing fault tolerance and recovery: When failures of some compute nodes in a parallel task are detected, the parallelism is adjusted dynamically and the unfinished tasks are reallocated to other healthy nodes.