Database pre-downsampling method and device, electronic equipment and storage medium

By merging pre-downsampling tasks in the computing nodes, reducing the number of data pulls and network traffic, the problem of memory pressure of storage nodes in the timing database is solved, and data query speed and database cluster availability are improved.

CN120277124APending Publication Date: 2025-07-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410017723.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-05
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The large amount of data stored in the timing database leads to slow data query speed, and the pre-down sampling process occupies memory of the storage node, affecting the smooth execution of the data query process, and reducing the availability of the timing database cluster.

Method used

Perform pre-down sampling tasks in the computing node, merge multiple pre-down sampling tasks with the sharded tasks associated with the same storage shard into a merge task, and perform pre-down sampling processing through the computing node pulling data from the storage shard, reducing the number of data pulls and network traffic consumption.

Benefits of technology

In the case of consuming less network traffic, it saves memory nodes, alleviates memory pressure, ensures the smooth execution of pre-down sampling and data query processes, and improves the availability of time-series database clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277124A_ABST
    Figure CN120277124A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a database pre-downsampling method and device, electronic equipment and a storage medium, which can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving and the like. The method comprises the steps that after time series data are written into all storage fragments of a storage node, a plurality of preset pre-downsampling tasks are obtained, and each pre-downsampling task is used for executing corresponding pre-downsampling processing on the time series data; merging a plurality of fragment tasks associated with the same storage fragment in the plurality of pre-downsampling tasks to obtain each merged task; and for each merging task, instructing the computing node to pull the written target data from the storage fragment associated with the merging task, executing corresponding pre-downsampling processing on the target data according to the merging task, and writing the obtained pre-downsampling data into the storage node. According to the invention, the memory of the storage node can be saved, and the availability of the time sequence database cluster is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a database pre-downsampling method, apparatus, electronic device, and storage medium. Background Art

[0002] In recent years, the application of time series databases (hereinafter referred to as time series databases) has become increasingly widespread. A time series database is used to store and process data with time tags, and this data with time tags changes over time, also known as time series data (hereinafter referred to as time series data).

[0003] Currently, the amount of data stored in a time series database is large, resulting in a slow data query speed. To improve the data query speed, when writing time series data into the time series database, the written time series data can be pre-downsampled according to a set rule, and the pre-downsampled data is stored. Pre-downsampling is a pre-computation method that can store the written data after reducing its precision; for example, if the written time series data is data collected every hour within a day, and assuming that the set rule for pre-downsampling is to calculate the average value every hour, then after pre-downsampling, 24 average values are obtained, that is, the data precision is reduced. In this way, when querying data, the pre-downsampled data can be queried according to the downsampling precision specified in the query condition to reduce the amount of data that needs to be calculated during the query, thereby improving the data query speed.

[0004] In the related art, for a time series database cluster including storage nodes, after writing time series data into a storage node, the storage node performs pre-downsampling on the written time series data. This storage node is used to store and query data. Since the pre-downsampling process requires a certain amount of memory, it increases the memory pressure on the storage node, which may affect the smooth execution of the data query process, thereby reducing the availability of the time series database cluster. Summary of the Invention

[0005] Embodiments of this application provide a database pre-downsampling method, apparatus, electronic device, and storage medium, which can save the memory of storage nodes when performing the pre-downsampling task, ensure the smooth execution of the data query process, and improve the availability of the time series database cluster.

[0006] On the one hand, a database pre-downsampling method provided by an embodiment of this application is applied to a scheduling node in a time series database cluster. The time series database cluster further includes at least one storage node and at least one computing node. The method includes:

[0007] After writing the time-series data to at least one storage shard of the storage node, obtain a plurality of preset pre-downsampling tasks; wherein each pre-downsampling task is used to perform corresponding pre-downsampling processing on the time-series data, and at least one shard task included in each pre-downsampling task is respectively associated with a corresponding storage shard;

[0008] Merge multiple shard tasks associated with the same storage shard among at least one shard task of each of the plurality of pre-downsampling tasks to obtain at least one merge task;

[0009] For each of the at least one merge task, perform the following operations respectively: instruct a computing node to pull target data from the written time-series data in the storage shard associated with one merge task, and perform corresponding pre-downsampling processing on the target data according to the one merge task, and write the obtained pre-downsampled data to the storage node.

[0010] On the one hand, a database pre-downsampling method provided by an embodiment of the present application is applied to a computing node of a time-series database cluster, and the method includes:

[0011] Obtain at least one merge task, and each merge task is determined by the following method: after writing the time-series data to at least one storage shard of the storage node of the time-series database cluster, merge multiple shard tasks associated with the corresponding storage shard among a plurality of preset pre-downsampling tasks; wherein each pre-downsampling task is used to perform corresponding pre-downsampling processing on the time-series data, and at least one shard task included in each pre-downsampling task is respectively associated with a corresponding storage shard;

[0012] For each of the at least one merge task, perform the following operations respectively: pull target data from the written time-series data in the storage shard associated with one merge task, and perform corresponding pre-downsampling processing on the target data according to the one merge task, and write the obtained pre-downsampled data to the storage node.

[0013] On the one hand, a database pre-downsampling device provided by an embodiment of the present application is applied to a scheduling node in a time-series database cluster, and the time-series database cluster further includes at least one storage node and at least one computing node, and the device includes:

[0014] A task acquisition unit, configured to obtain a plurality of preset pre-downsampling tasks after writing the time-series data to at least one storage shard of the storage node; wherein each pre-downsampling task is used to perform corresponding pre-downsampling processing on the time-series data, and at least one shard task included in each pre-downsampling task is respectively associated with a corresponding storage shard;

[0015] A merging unit, configured to merge multiple shard tasks associated with the same storage shard among at least one shard task of each of the multiple pre-downsampling tasks, to obtain at least one merged task;

[0016] A processing unit, configured to respectively perform the following operations for the at least one merged task: instruct a computing node to pull target data from the written time-series data in a storage shard associated with a merged task, and perform corresponding pre-downsampling processing on the target data according to the merged task, and write the obtained pre-downsampled data into a storage node.

[0017] On the one hand, a database pre-downsampling device provided by an embodiment of the present application is applied to a computing node of a time-series database cluster. The device includes:

[0018] A task acquisition unit, configured to acquire at least one merged task. Each merged task is determined by the following method: after writing time-series data into at least one storage shard of a storage node of the time-series database cluster, merging multiple shard tasks associated with the corresponding storage shard among a preset multiple pre-downsampling tasks; wherein each pre-downsampling task is configured to perform corresponding pre-downsampling processing on the time-series data, and at least one shard task included in each pre-downsampling task is respectively associated with a corresponding storage shard;

[0019] A processing unit, configured to respectively perform the following operations for the at least one merged task: pull target data from the written time-series data in a storage shard associated with a merged task, and perform corresponding pre-downsampling processing on the target data according to the merged task, and write the obtained pre-downsampled data into a storage node.

[0020] An electronic device provided by an embodiment of the present application includes a processor and a memory. Wherein, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of any one of the above database pre-downsampling methods.

[0021] An embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the computer program is configured to cause the electronic device to execute the steps of any one of the above database pre-downsampling methods.

[0022] An embodiment of the present application provides a computer program product, the computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the steps of any one of the above database pre-downsampling methods.

[0023] The present application has at least the following beneficial effects:

[0024] An embodiment of the present application provides a database pre-downsampling method, device, electronic device and storage medium. After writing time series data into at least one storage shard of a storage node, corresponding pre-downsampling processing needs to be performed on the written time series data according to multiple pre-downsampling tasks; before performing the pre-downsampling processing, first merge multiple shard tasks associated with the same storage shard in the multiple pre-downsampling tasks to obtain at least one merged task; for each merged task, instruct a computing node to pull target data in the written time series data from the storage shard associated with the merged task, and perform corresponding pre-downsampling processing on the target data according to the merged task. In this way, in the embodiment of the present application, the pre-downsampling process is performed by the computing node, so that the pre-downsampling process does not need to occupy the memory of the storage node, alleviating the memory pressure of the storage node.

[0025] At the same time, the embodiment of the present application takes into account that when performing pre-downsampling processing in a computing node, the computing node needs to pull the written time series data from the storage node, which will consume network traffic. In order to reduce the consumed network traffic as much as possible, multiple shard tasks associated with the same storage shard are merged. In this way, it is not necessary to pull data from the corresponding storage shard for each shard task, but to pull data from the corresponding storage shard for each merged task, greatly reducing the number of data pulls, and thus greatly reducing the consumed network traffic. Therefore, the embodiment of the present application can save the storage memory of the storage node while consuming less network traffic, ensure that both the pre-downsampling process and the data query process can be executed smoothly, thereby improving the availability of the time series database cluster.

[0026] In addition, in the process of processing multiple pre-downsampling tasks in the embodiment of the present application, new pre-downsampling tasks can also be introduced, so that new pre-downsampling tasks can be infinitely extended.

[0027] Other features and advantages of the present application will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0029] Figure 1 It is a schematic diagram of an application scenario of a database pre-downsampling method in an embodiment of the present application;

[0030] Figure 2 It is a flowchart of a database pre-downsampling method in an embodiment of the present application;

[0031] Figure 3 It is a schematic diagram of the relationship between a pre-downsampling task and a storage shard in an embodiment of the present application;

[0032] Figure 4 It is a schematic diagram of the merging of sharding tasks of a pre-downsampling task in an embodiment of the present application;

[0033] Figure 5 It is a schematic diagram of the distribution of merging tasks of computing nodes in an embodiment of the present application;

[0034] Figure 6 It is a schematic diagram of the execution status of a pre-downsampling task in an embodiment of the present application;

[0035] Figure 7 It is a schematic diagram of expanding a new merging task in an embodiment of the present application;

[0036] Figure 8 It is a flowchart of another database pre-downsampling method in an embodiment of the present application;

[0037] Figure 9 It is a logical schematic diagram of a database pre-downsampling method in an embodiment of the present application;

[0038] Figure 10 It is a schematic diagram of data pulling of computing nodes in an embodiment of the present application;

[0039] Figure 11 It is a schematic diagram of expanding a new pre-downsampling task in an embodiment of the present application;

[0040] Figure 12 It is a schematic diagram of the composition structure of a database pre-downsampling device in an embodiment of the present application;

[0041] Figure 13 It is a schematic diagram of the composition structure of another database pre-downsampling device in an embodiment of the present application;

[0042] Figure 14It is a schematic diagram of a hardware component structure of an electronic device applying an embodiment of the present application;

[0043] Figure 15 It is a schematic diagram of a hardware component structure of another electronic device applying an embodiment of the present application. Specific implementation manners

[0044] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the technical solutions of the present application, rather than all of the embodiments. Based on the embodiments recorded in this application document, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the technical solutions of the present application.

[0045] Some concepts involved in the embodiments of the present application will be introduced below.

[0046] Time series data: That is, time - series data, which refers to data with time tags and can be stored in a time - series database.

[0047] Pre - downsampling: A pre - calculation method. When data is written, the data is stored after reducing its precision according to the configured pre - downsampling rule. When querying, the data closest to the pre - downsampling precision is automatically queried according to the downsampling precision specified in the query condition, so as to reduce the amount of data that needs to be calculated in real - time query and reduce access latency.

[0048] The term "exemplary" used hereinafter means "serving as an example, an embodiment or an illustration". Any embodiment described as "exemplary" does not have to be construed as superior to or better than other embodiments.

[0049] The terms "first" and "second" in the text are only used for descriptive purposes and cannot be construed as explicitly or implicitly indicating relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.

[0050] Cloud technology is a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide - area network or a local - area network to realize the calculation, storage, processing, and sharing of data.

[0051] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used as needed, and is flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the highly developed and applied Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires the support of a powerful system background, which can only be achieved through cloud computing. The database pre-downsampling method in the embodiments of this application can be implemented through cloud computing.

[0052] A database, in short, can be regarded as an electronic filing cabinet - a place to store electronic files. Users can perform operations such as adding, querying, updating, and deleting data in the files. The so-called "database" is a data set stored together in a certain way, can be shared by multiple users, has the smallest possible redundancy, and is independent of application programs. The time series database involved in the embodiments of this application belongs to a type of database.

[0053] A database management system (DBMS) is a computer software system designed to manage a database and generally has basic functions such as storage, interception, security guarantee, and backup. The database management system can be classified according to the database model it supports, such as relational, Extensible Markup Language (XML); or according to the type of computer it supports, such as server clusters, mobile phones; or according to the query language it uses, such as Structured Query Language (SQL), XML Query (XQuery); or according to the key points of performance impulse, such as maximum scale, highest running speed; or other classification methods. No matter which classification method is used, some DBMS can cross categories. For example, they can support multiple query languages at the same time. The time series database cluster in the embodiments of this application can be managed through a database management system.

[0054] The design concept of the embodiments of this application is briefly introduced below.

[0055] At present, the large amount of data stored in the time series database makes the data query speed slow. To improve the data query speed, when writing data into the time series database, the written data can be pre-downsampled according to a set rule, and the pre-downsampled data can be stored. In this way, when querying data, the pre-downsampled data can be queried according to the downsampling accuracy specified in the query condition, so as to reduce the amount of data that needs to be calculated during the query, thereby improving the data query speed. In the related art, for a time series database cluster including storage nodes, the time series database cluster can be a storage-computation separation architecture or a non-separation architecture of storage and computation. After writing the time series data into the storage node, the storage node performs pre-downsampling on the written time series data. The storage node is used to store and query data. Since the pre-downsampling process requires a certain amount of memory, the memory pressure on the storage node is increased, which may affect the smooth execution of the data query process, thereby reducing the availability of the time series database.

[0056] In view of this, the present application is directed to a time series database cluster with a storage-computation separation architecture. The time series database cluster includes separate storage nodes and computing nodes. It is considered to execute the pre-downsampling task in the computing node to save the memory of the storage node and reduce the memory pressure on the storage node. However, this requires the computing node to pull the written time series data from the storage node, which consumes network traffic.

[0057] In order to minimize the consumed network traffic, after writing the time series data into at least one storage shard of the storage node, a preset number of pre-downsampling tasks are obtained. Among the at least one shard task of each of these pre-downsampling tasks, the multiple shard tasks associated with the same storage shard are merged to obtain at least one merged task. For each merged task, a computing node is instructed to pull the target data in the written time series data from the storage shard associated with the merged task, and perform corresponding pre-downsampling processing on the target data according to the merged task. In this way, it is not necessary to pull data from the corresponding storage shard for each shard task, but to pull data from the corresponding storage shard for each merged task, greatly reducing the number of data pulls, thereby reducing the consumed network traffic. Therefore, the embodiments of the present application can save the storage memory of the storage node, relieve the memory pressure on the storage node, ensure that both the pre-downsampling process and the data query process can be smoothly executed, and thus improve the availability of the time series database cluster under the condition of consuming less network traffic.

[0058] The following describes the preferred embodiments of the present application with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0059] AsFigure 1 As shown in the figure, it is a schematic diagram of an application scenario of an embodiment of the present application. The application scenario diagram includes a terminal device 110, a scheduling node 120, a storage node 130, and a computing node 140. Among them, the scheduling node 120, the storage node 130, and the computing node 140 can form a time series database cluster, and the time series database cluster can adopt an architecture with separated storage and computing.

[0060] In the embodiments of the present application, the scheduling node 120, the storage node 130, and the computing node 140 can all be deployed on a server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart voice interaction device, a smart speaker, a smart watch, a smart home appliance, a vehicle terminal, an aircraft, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this.

[0061] It should be noted that the database pre-downsampling method in each embodiment of the present application can be executed by the scheduling node 120 or the computing node 140. The following takes the scheduling node 120 as an example for illustration.

[0062] In some embodiments, the user can, through the terminal device 110, preset multiple pre-downsampling tasks for the time series data written to the storage node of the time series database cluster. These pre-downsampling tasks have the same time sampling window, and each pre-downsampling task is used to perform corresponding pre-downsampling processing on the written time series data; the terminal device 110 can upload these pre-downsampling tasks to the time series database cluster. After the terminal device 110 uploads the collected time series data to the time series database cluster, the scheduling node 120 of the time series database cluster writes the time series data into at least one storage shard of the storage node, and then obtains multiple pre-downsampling tasks preset for the storage node, and then based on the data pre-downsampling method of the embodiment of the present application, performs corresponding pre-downsampling processing on the written time series data.

[0063] Specifically, to save the memory of the storage node, the scheduling node 120 may instruct at least one computing node to perform corresponding pre-downsampling processing on the written time-series data. Moreover, to reduce the network traffic for the computing node to pull data from the storage node, among at least one shard task of each of the multiple pre-downsampling tasks, multiple shard tasks associated with the same storage shard are merged to obtain at least one merged task. For each merged task, a computing node is instructed to pull the target data from the written time-series data in the storage shard associated with the merged task, and perform corresponding pre-downsampling processing on the target data according to the merged task.

[0064] It should be noted that Figure 1 The illustration is only an example. In fact, the numbers of the terminal device 110, the scheduling node 120, the storage node 130, and the computing node 140 are not limited and are not specifically defined in the embodiments of the present application.

[0065] In addition, the embodiments of the present application can be applied to various scenarios, including not only database scenarios, but also, without limitation, scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving.

[0066] Next, in combination with the above-described application scenarios, the database pre-downsampling method provided by the exemplary embodiment of the present application will be described with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard.

[0067] Refer to Figure 2 As shown, it is a flowchart of the implementation of a database pre-downsampling method provided by an embodiment of the present application. Taking the scheduling node of the time-series database cluster as the execution subject as an example, the time-series database system includes at least one storage node and at least one computing node. The specific implementation process of the method includes the following S21-S23:

[0068] S21: After writing the time-series data into at least one storage shard of the storage node, obtain a preset multiple pre-downsampling tasks; wherein, each pre-downsampling task is used to perform corresponding pre-downsampling processing on the time-series data, and at least one shard task included in each pre-downsampling task is respectively associated with a corresponding storage shard.

[0069] Among them, after obtaining the time-series data sent by the terminal device, the scheduling node may write the time-series data into the corresponding storage node, and the storage node includes database tables, and each database table includes multiple storage shards. It should be noted that the scheduling node may write the time-series data into at least one storage shard of one storage node, or may also write the time-series data into multiple storage shards of multiple storage nodes, which is not limited in this regard.

[0070] Users can preset multiple pre-downsampling tasks for the written time-series data. These pre-downsampling tasks target the same time-series data and have the same time sampling window. Multiple pre-downsampling tasks can be configured through different query statements.

[0071] Exemplarily, the query statement is SQL, and the following two pre-downsampling tasks are preset:

[0072] 1. select min(field1), max(field1) from m1 group by(12m), *

[0073] 2. select median(field1), mean(field1) from m1 group by(12m), *

[0074] Among them, m1 represents the written time-series data, group by(12m) represents that the time sampling window is 12 minutes, min(field1) and max(field1) respectively represent taking the minimum value and the maximum value of field1 (representing a certain field, such as the data in a certain column) in m1, and median(field1) and mean(field1) respectively represent taking the median and the average value of field1 in m1.

[0075] The first pre-downsampling task in the above example is to calculate the minimum value and the maximum value for the time-series data with a 12-minute time sampling window, and the second pre-downsampling task is to calculate the median and the average value for the time-series data with a 12-minute time sampling window.

[0076] After writing the time-series data into at least one storage shard of the storage node, for each pre-downsampling task, it is necessary to perform the corresponding pre-downsampling processing on the data written in at least one storage shard. That is to say, each pre-downsampling task can be divided into at least one shard task, and each shard task is used to perform pre-downsampling processing on the data written in a storage shard.

[0077] It can be understood that different shard tasks of the same pre-downsampling task are associated with different storage shards, and the shard tasks of different pre-downsampling tasks can be associated with the same storage shard.

[0078] Exemplarily, such as Figure 3As shown in the figure, assume that time series data is written into storage shard 1, storage shard 2, and storage shard 3 of a storage node, and the user has pre-set pre-downsampling task 1 and pre-downsampling task 2 for this storage node. Then, pre-downsampling task 1 includes three shard tasks respectively associated with the above three storage shards: shard task 11, shard task 12, and shard task 13. Pre-downsampling task 2 includes three shard tasks respectively associated with the above three storage shards: shard task 21, shard task 22, and shard task 23. That is to say, shard task 11 and shard task 21 are both associated with storage shard 1, shard task 12 and shard task 22 are both associated with storage shard 2, and shard task 13 and shard task 23 are both associated with storage shard 3.

[0079] S22: Merge multiple shard tasks of each of multiple pre-downsampling tasks that are associated with the same storage shard to obtain at least one merged task.

[0080] Continuing with the above example, as Figure 4 shown in the figure, for pre-downsampling task 1 and pre-downsampling task 2, merge shard task 11 and shard task 21 associated with storage shard 1 into merged task 1. This merged task 1 is used to perform corresponding pre-downsampling processing on the target data written into storage shard 1 (including the processing processes of shard task 11 and shard task 21). Merge shard task 21 and shard task 22 associated with storage shard 2 into merged task 2. This merged task 2 is used to perform corresponding pre-downsampling processing on the target data written into storage shard 2 (including the processing processes of shard task 21 and shard task 22). Merge shard task 13 and shard task 23 associated with storage shard 3 into merged task 3. This merged task 3 is used to perform corresponding pre-downsampling processing on the target data written into storage shard 3 (including the processing processes of shard task 13 and shard task 23).

[0081] S23: For each of the at least one merged task, perform the following operations respectively: instruct a computing node to pull the target data from the time series data written into the storage shard associated with a merged task, perform corresponding pre-downsampling processing on the target data according to a merged task, and write the obtained pre-downsampled data into the storage node.

[0082] Among them, when there are multiple merged tasks and multiple computing nodes, the multiple merged tasks can be evenly divided into multiple computing nodes, specifically according to the load conditions of each computing node. The division method is not limited here. As Figure 5As shown, assume there are 3 merging tasks: Merging Task 1, Merging Task 2, and Merging Task 3, and the time series database cluster includes 2 computing nodes: Computing Node 1 and Computing Node 2. Optionally, Merging Task 1 and Merging Task 2 can be assigned to Computing Node 1, and Merging Task 3 can be assigned to Computing Node 2.

[0083] In this way, the scheduling node can instruct Computing Node 1 to respectively pull the target data written from Storage Shard 1 and Storage Shard 2 for Merging Task 1 and Merging Task 2, and perform corresponding pre-downsampling processing on the target data pulled from Storage Shard 1 according to Merging Task 1, and perform corresponding pre-downsampling processing on the target data pulled from Storage Shard 2 according to Merging Task 2; the scheduling node also instructs Computing Node 2 to respectively pull the target data written from Storage Shard 3 and Storage Shard 4 for Merging Task 3 and Merging Task 4, and perform corresponding pre-downsampling processing on the target data pulled from Storage Shard 3 according to Merging Task 3, and perform corresponding pre-downsampling processing on the target data pulled from Storage Shard 4 according to Merging Task 4.

[0084] Specifically, when writing time series data to each storage shard of the storage node, the time series data is first written to the write-ahead log corresponding to each storage shard. Therefore, the computing node can pull the write-ahead log from the storage shard, and the write-ahead log contains the target data written.

[0085] In the embodiments of the present application, the pre-downsampling task is executed in the computing node to save the memory of the storage node and reduce the memory pressure on the storage node; in order to minimize the network traffic for the computing node to pull data from the storage node, among at least one shard task of each pre-downsampling task, multiple shard tasks associated with the same storage shard are merged to obtain at least one merging task; for each merging task, a computing node is instructed to pull the target data in the time series data written from the storage shard associated with the merging task, and perform corresponding pre-downsampling processing on the target data according to the merging task. In this way, instead of pulling data from the corresponding storage shard once for each shard task, data is pulled from the corresponding storage shard once for each merging task, greatly reducing the number of data pulls, thereby reducing the consumed network traffic.

[0086] For example, time series data is written into 4 storage shards, and there are 4 pre-downsampling tasks. Each pre-downsampling task contains shard tasks for the 4 storage shards respectively, that is, there are a total of 16 shard tasks. If data is pulled once for each shard task, 16 data pulls are required; while every 4 shard tasks associated with the same storage shard among the 16 shard tasks are combined into one combined task, obtaining 4 combined tasks, only 4 data pulls are needed. It can be understood that the more pre-downsampling tasks there are and the more storage shards there are, the more data pull times can be saved, that is, the more network traffic can be saved.

[0087] Therefore, the embodiments of the present application can save the memory of the storage node and reduce the load of the processor while consuming less network traffic, relieve the memory pressure and load pressure of the storage node, ensure that both the pre-downsampling process and the data query process can be executed smoothly, and thus improve the availability of the time series database cluster.

[0088] In some embodiments, since the time sampling windows corresponding to multiple pre-downsampling tasks are the same, therefore, in the above S23, performing the corresponding pre-downsampling process on the target data according to one combined task may specifically include the following steps A1 - A2:

[0089] A1. During the process of sampling the target data according to the time sampling window, every time sampling data within one time sampling window is obtained, perform the corresponding pre-downsampling process on the sampling data according to one combined task to obtain the corresponding pre-downsampled sub-data.

[0090] Specifically, for the target data, divide the time sampling window on the time axis, and perform pre-downsampling on the sampling data within each time sampling window. For example, if the target data contains data collected every hour within a day and the time sampling window can be one hour, then every time the sampling data within one hour is obtained, find the maximum value, minimum value, median value, and average value of the sampling data, that is, perform pre-downsampling.

[0091] A2. Whenever the pre-downsampled sub-data corresponding to N time sampling windows is obtained, write the N pre-downsampled sub-data into the storage node; where N is an integer greater than or equal to 1.

[0092] Among them, when N is 1, for each pre-downsampled sub-data obtained in a time sampling window, the pre-downsampled sub-data is written into the corresponding storage node. In this way, the real-time writing of the pre-downsampled data can be ensured, meeting the user's requirements for writing real-time performance. When N is an integer greater than 1, the specific value can be set as needed. In this way, out-of-order data within N time sampling windows can be tolerated because the time-series data written by the user may contain out-of-order data. When it is determined that there is out-of-order data within N time sampling windows, since the pre-downsampled sub-data of these N time sampling windows is saved, the out-of-order data can be reprocessed accordingly.

[0093] In the embodiments of the present application, in order to balance the real-time writing of the pre-downsampled data and the problem of out-of-order data in the sampled data, each computing node allows storing the pre-downsampled sub-data of N time sampling windows. When strong real-time writing performance of the pre-downsampled data is required, N can be set to 1, that is, after the sampled data within each time sampling window completes the pre-downsampling process, the obtained pre-downsampled sub-data is written into the storage node, giving the strongest real-time performance; when there are out-of-order data points in the sampled data, N can be set to an integer greater than 1 to tolerate the out-of-order data within N time sampling windows.

[0094] In an optional implementation manner, when N is an integer greater than 1, after obtaining the sampled data in the latest time sampling window, the scheduling node can specifically perform the following steps A11 - A12:

[0095] A11. If there are out-of-order data points in the sampled data in the latest time sampling window, and the time tags carried by the out-of-order data points are within any one of the previous N - 1 time sampling windows, insert the out-of-order data points into the sampled data within any one of the time sampling windows, and delete the out-of-order data points from the latest sampled data.

[0096] Exemplarily, assume that N is 3, and the latest time sampling window is from 11 o'clock to 12 o'clock. When the sampled data from 11 o'clock to 12 o'clock is obtained and it is found that there is an out-of-order data point with a time tag of 9:30 among them. At this time, the pre-downsampled sub-data of the two time sampling windows from 9 o'clock to 10 o'clock and from 10 o'clock to 11 o'clock is saved. Therefore, the out-of-order data point of 9:30 can be inserted into the sampled data of the time sampling window from 9 o'clock to 10 o'clock, and the out-of-order data point of 9:30 can be deleted from the sampled data from 11 o'clock to 12 o'clock.

[0097] A12. According to a merging task, perform the corresponding pre-downsampling process on the sampled data after inserting the out-of-order data points again, and perform the corresponding pre-downsampling process on the latest sampled data after deleting the out-of-order data points.

[0098] In the embodiments of the present application, for the out-of-order data points within the most recent time sampling window, the historical time sampling window to which the out-of-order data points belong is determined. If the pre-downsampled sub-data of the historical time sampling window has not been written to the storage node, the out-of-order data points can be inserted into the historical time sampling window, so as to re-perform pre-downsampling processing on the sampling data after inserting the out-of-order data points, and perform pre-downsampling processing on the most recent sampling data after deleting the out-of-order data points. In this way, the inaccuracy of the pre-downsampling result caused by the out-of-order data points can be avoided, and the accuracy of the pre-downsampling result can be improved.

[0099] In addition, when the pre-downsampled sub-data of the historical time sampling window to which the out-of-order data points belong has been written to the storage node, the out-of-order data points can be deleted from the most recent sampling data, and pre-downsampling processing can be performed on the most recent sampling data after deleting the out-of-order data points. At the same time, the execution failure information of the out-of-order data points can be recorded.

[0100] It should be noted that the embodiments of the present application can also check whether there are out-of-order data points in the target data pulled from the storage shard. If there are out-of-order data points, the out-of-order data points are inserted into the correct positions. After that, the adjusted target data is sampled according to the time sampling window, and pre-downsampling processing is performed. In this way, the inaccuracy of the pre-downsampling result caused by the out-of-order data points can also be avoided.

[0101] In some embodiments, during the process of performing corresponding pre-downsampling processing on the target data according to at least one merging task respectively, the execution status of each merging task can be obtained periodically, and then the execution status of each pre-downsampling task can be determined. The scheduling node can specifically perform the following steps B1-B2:

[0102] B1. For at least one merging task, perform the following operations respectively: obtain the execution status of a merging task from the corresponding computing node every set period, and the execution status includes the execution status of each of the multiple shard tasks included in a merging task.

[0103] Among them, the set period can be set as needed and is not limited thereto. The execution status of each shard task includes, but is not limited to, the number of data points successfully executed, the data points with execution failures, the memory occupied by the shard task, etc.

[0104] Considering that exceptions may occur during the execution of each merging task, the execution status of each merging task is obtained periodically to timely detect abnormal data (such as data with execution failures) during the pre-downsampling process. It can be understood that each merging task includes multiple shard tasks, so the execution status of each merging task includes the execution status of multiple shard tasks.

[0105] B2. For multiple pre-downsampling tasks, perform the following operations respectively: At each set period, merge the execution statuses of at least one shard task included in a pre-downsampling task to obtain the execution status of a pre-downsampling task.

[0106] Among the multiple shard tasks included in each merge task, there is at least one shard task included in each pre-downsampling task. Therefore, after obtaining the execution statuses of the multiple shard tasks included in each merge task each time, merge the execution statuses of at least one shard task of each pre-downsampling task, that is, obtain the execution status of each pre-downsampling task.

[0107] Exemplarily, as Figure 6 shown, obtain the execution statuses of merge task 1, merge task 2, and merge task 3. Merge task 1 includes the execution statuses of shard task 11 and shard task 21 respectively, merge task 2 includes the execution statuses of shard task 12 and shard task 22 respectively, and merge task 3 includes the execution statuses of shard task 13 and shard task 23 respectively. Merge the execution statuses of shard task 11, shard task 12, and shard task 13 respectively to obtain the execution status of pre-downsampling task 1, and merge the execution statuses of shard task 21, shard task 22, and shard task 23 respectively to obtain the execution status of pre-downsampling task 2.

[0108] In the embodiments of the present application, the scheduling node can periodically obtain the execution statuses of the merge tasks on each computing node, parse the execution statuses of multiple shard tasks from the execution statuses of the merge tasks, and for each pre-downsampling task, merge the execution statuses of at least one shard task included therein, so as to obtain the execution status of each pre-downsampling task. Through the execution status, abnormal data in the pre-downsampling process can be found in a timely manner, so as to re-perform pre-downsampling processing on the abnormal data subsequently.

[0109] In some embodiments, in order to facilitate the user to timely understand the execution statuses of multiple pre-downsampling tasks, the user is allowed to query the execution status of each pre-downsampling task. Specifically, the scheduling node can provide a status acquisition interface, such as a Remote Procedure Call (RPC) interface. The query end can call this status acquisition interface to query the execution status of any pre-downsampling task.

[0110] When the scheduling node receives a status query request for any pre-downsampling task sent by the query end, it can obtain the latest execution status of any pre-downsampling task and send the latest execution status of any pre-downsampling task to the corresponding query end.

[0111] In addition, each computing node can also directly access the scheduling node to obtain the latest execution status of any pre-downsampling task queried by the query end, and send the latest execution status of any pre-downsampling task to the corresponding query end.

[0112] In the embodiments of the present application, the user queries the execution status of the pre-downsampling task to timely discover abnormal data in the pre-downsampling process, so that a re-execution request can be initiated for the abnormal data, avoiding inaccurate pre-downsampling results caused by the abnormal data and improving the accuracy of the pre-downsampling results.

[0113] In some embodiments, after the scheduling node sends the latest execution status of any pre-downsampling task to the corresponding query end, it can also receive a supplementary task of any pre-downsampling task sent by the query end; wherein, the supplementary task is sent when there is abnormal pre-downsampling data in the latest execution status.

[0114] Furthermore, the scheduling node instructs the corresponding computing node to re-execute the corresponding pre-downsampling process on the first original data corresponding to the abnormal pre-downsampling data according to the supplementary task, and the first original data is included in the time series data.

[0115] Among them, when the first original data corresponding to the abnormal pre-downsampling data is still saved in the computing node, the corresponding pre-downsampling process can be directly re-executed on the first original data; when the first original data is not saved in the computing node, the first original data can be pulled from the corresponding storage slice of the storage node, and the corresponding pre-downsampling process is re-executed on the first original data.

[0116] In the embodiments of the present application, through the execution status of the pre-downsampling task, the user can discover abnormal data in the pre-downsampling process. For example, when the number of the above N time sampling windows configured by the user is small, resulting in any downsampling task losing out-of-order data points and making the pre-downsampling result inaccurate, at this time, a supplementary task of the pre-downsampling task can be initiated for the out-of-order data points that failed to execute, so as to re-execute the pre-downsampling process on the sampling data within the time sampling window where the out-of-order data points are located, ensuring the accuracy of the pre-downsampling result.

[0117] In some embodiments, when a computing node experiences abnormal power-off, data loss may occur. At this time, the scheduling node can obtain the execution status of the merging task in the computing node to determine the lost data, and specifically can perform the following steps C1 - C3:

[0118] C1. When an abnormal condition occurs in a computing node, obtain the execution status of each of at least one merging task in the computing node.

[0119] Among them, an abnormal condition of a computing node includes but is not limited to: power failure, bugs, downtime, etc. At this time, the data in the memory of the computing node may be lost.

[0120] The scheduling node can obtain the execution status of at least one merging task in the computing node saved before the abnormal condition of the computing node occurs. The execution status of each merging task includes the execution status of multiple shard tasks. The execution status of each shard task, for example, includes: the number of data points successfully executed, the data points with execution failures, the memory occupied by the shard task, etc.

[0121] C2. According to the execution status of any merging task, when it is determined that the pre-downsampled data of any merging task is lost, obtain the second original data corresponding to the lost pre-downsampled data from the storage node. The second original data is included in the time series data.

[0122] Specifically, according to the execution status of any merging task, the execution status of multiple shard tasks in the merging task can be obtained. According to the execution status of these shard tasks, it can be determined whether each shard task has lost pre-downsampled data. For the lost pre-downsampled data, the corresponding second original data can be re-pulled from the storage node.

[0123] C3. After the computing node is restored, instruct the computing node to re-execute the corresponding pre-downsampling process on the second original data according to any merging task, and write the obtained pre-downsampled data into the storage node.

[0124] Optionally, instruct the restored computing node to perform the corresponding pre-downsampling process on the second original data according to the shard task that has lost pre-downsampled data in any merging task. In addition, when the computing node with abnormal conditions cannot be restored for a long time, other computing nodes can also be instructed to perform the corresponding pre-downsampling process on the second original data.

[0125] In the embodiments of the present application, when an abnormal condition occurs in a computing node, it can be determined whether each merging task has lost pre-downsampled data according to the execution status of each merging task in the computing node. When it is determined that any merging task has lost pre-downsampled data, the corresponding second original data can be re-pulled from the storage node so that after the computing node is restored, the corresponding pre-downsampling process can be re-executed on the second original data, thereby avoiding the situation of data loss in the pre-downsampling result.

[0126] In some embodiments, during the execution of multiple pre-downsampling tasks, the user may create a new pre-downsampling task in the storage node. When the scheduling node obtains the new pre-downsampling task, the following operations are respectively performed on at least one new shard task included in the new pre-downsampling task:

[0127] When there is a target merging task corresponding to the storage shard associated with a new sharding task, if it is determined that the memory used to execute a new sharding task is not greater than the remaining memory of the computing node where the target merging task is located, then the new sharding task is assigned to the target merging task.

[0128] Among them, the new pre-downsampling task is used to perform corresponding pre-downsampling processing on the target data in each storage shard written to the storage node. For example, a new sharding task in the new pre-downsampling task is associated with storage shard 1. For the created merging task 1 (i.e., the target merging task) in storage shard 1, at this time, when it is determined that the memory used to execute this new sharding task is not greater than the remaining memory of the computing node where merging task 1 is located, this new sharding task can be assigned to merging task 1, instructing the computing node where merging task 1 is located to execute this new sharding task based on the pulled target data.

[0129] In this embodiment, for any new sharding task in the new pre-downsampling task, when it is determined that there is a target merging task for the storage shard associated with this new sharding task, in principle, this new sharding task can be assigned to the target merging task to save the network traffic consumed by pulling data. However, since the memory used to execute this new sharding task is uncertain, and after this new sharding task is assigned to the target merging task, the computing node executing the target merging task may run out of memory. For example, the computing node has a total of 16G of memory, and the target merging task already includes two sharding tasks, and each sharding task requires 7G of memory. If the new sharding task is also executed on this computing node, it will exceed the memory limit of this computing node. Therefore, when it is determined that the memory used to execute the new sharding task is less than the remaining memory of the computing node where the target merging task is located, then this new sharding task is assigned to the target merging task.

[0130] If it is determined that the memory used to execute the new sharding task is greater than the remaining memory of the computing node where the target merging task is located, this new sharding task can be assigned to other computing nodes for execution.

[0131] In the embodiment of the present application, when a new pre-downsampling task is obtained, the new sharding tasks included in the new pre-downsampling task can be assigned to the target merging task to save the network traffic consumed by pulling data, and to ensure that the memory of the computing node where the target merging task is located does not exceed the upper limit. Therefore, the embodiment of the present application can expand the new pre-downsampling task during the sharding task merging process.

[0132] In an alternative embodiment, to determine the memory used for executing a new sharding task, the scheduling node may instruct other computing nodes other than the computing node where the target merging task is located to execute a new sharding task; when it is determined that the memory occupied by a new sharding task in other computing nodes is not greater than the remaining memory of the computing node where the target merging task is located, the new sharding task is transferred from other computing nodes to the computing node where the target merging task is located.

[0133] Wherein, when executing a new sharding task in other computing nodes, the existing data in the other computing nodes may be used to execute the new sharding task to determine the memory used for executing the new sharding task. Optionally, the other computing nodes may be computing nodes with relatively sufficient memory, and specifically, other computing nodes may be selected according to the actual situation.

[0134] Exemplarily, as Figure 7 shown, assume that merging task 1 is running in computing node 1, and this merging task 1 is used to perform pre-downsampling processing on the target data pulled from storage shard 1. When the user creates a new pre-downsampling task, the new pre-downsampling task includes a new sharding task corresponding to storage shard 1. At this time, the new sharding task is first run in computing node 2. When it is determined that the sum of the occupied memory t1 of merging task 1 and the occupied memory t2 of the new sharding task is not greater than the memory T of computing node 1, the new sharding task can be divided into merging task 1 to obtain a new merging task 1. Thus, it is not necessary to pull data from storage shard 1 again for the new sharding task, but rather to execute the new sharding task on the data that computing node 1 has already pulled from storage shard 1.

[0135] In addition, when it is determined that the sum of the occupied memory t1 of merging task 1 and the occupied memory t2 of the new sharding task is greater than the memory T of computing node 1, the new sharding task can continue to be executed on computing node 2.

[0136] In the above embodiments of the present application, when it is not determined whether the occupied memory of the new sharding task is not greater than the remaining memory of the computing node where the target merging task is located, the new sharding task may be first run on other computing nodes to determine the occupied memory of the new sharding task. In this way, when the remaining memory of the computing node where the target merging task is located is insufficient to execute the new sharding task, the smooth execution of the target merging task is avoided.

[0137] Based on the same inventive concept, the embodiments of the present application further provide a database pre-downsampling method, which is applied to a computing node of a time series database cluster, as Figure 8 shown, this database pre-downsampling method may include the following steps S81 - S82:

[0138] S81: Obtain at least one merging task, where each merging task is determined as follows: after writing time-series data into at least one storage shard of a storage node in a time-series database cluster, merge multiple shard tasks associated with the corresponding storage shard among multiple preset pre-downsampling tasks.

[0139] Among them, each pre-downsampling task is used to perform corresponding pre-downsampling processing on time-series data, and at least one shard task included in each pre-downsampling task is respectively associated with a corresponding storage shard;

[0140] S82: For at least one merging task, respectively perform the following operations: Pull the target data in the written time-series data from the storage shard associated with a merging task, and perform corresponding pre-downsampling processing on the target data according to a merging task, and write the obtained pre-downsampled data into the storage node.

[0141] In some embodiments, the time sampling windows corresponding to multiple pre-downsampling tasks are the same; in the above S82, performing corresponding pre-downsampling processing on the target data according to a merging task specifically includes the following steps:

[0142] During the process of sampling the target data according to the time sampling window, each time sampling data within a time sampling window is obtained, perform corresponding pre-downsampling processing on the sampling data according to a merging task to obtain corresponding pre-downsampled sub-data;

[0143] Whenever the pre-downsampled sub-data corresponding to N time sampling windows are obtained, write the N pre-downsampled sub-data into the storage node; where N is an integer greater than or equal to 1.

[0144] Optionally, when N is an integer greater than 1, after obtaining the sampling data within the latest time sampling window, the following steps are further included:

[0145] If there are out-of-order data points in the sampling data within the latest time sampling window, and the time tag carried by the out-of-order data points is within any time sampling window among the previous N - 1 time sampling windows, insert the out-of-order data points into the sampling data within any time sampling window and delete the out-of-order data points from the latest sampling data;

[0146] According to a merging task, perform corresponding pre-downsampling processing on the sampling data after inserting the out-of-order data points again, and perform corresponding pre-downsampling processing on the latest sampling data after deleting the out-of-order data points.

[0147] Specifically, for the specific implementation processes of S81 - S82 above, refer to the specific implementation processes of S21 - S23 in the above embodiments of this application, which will not be elaborated here. In addition, in addition to the above S81 - S82, the computing node can also perform other operations. For details, refer to the above embodiments of this application, which will not be elaborated here.

[0148] The following combines Figure 9 to exemplarily introduce the overall logic of the data pre - downsampling method in the embodiments of this application.

[0149] As Figure 9 shown, the user writes time - series data into 3 storage shards of the storage node. Specifically, 12 * 60 data are written into storage shard 1, storage shard 2, and storage shard 3 respectively. Assuming that 12 * 60 data represent 60 data collected per hour in 12 hours, there are preset pre - downsampling task 1 (calculating the maximum and minimum values once per hour) and pre - downsampling task 2 (calculating the median and average values once per hour) for this storage node; after merging the 3 shard tasks of these two pre - downsampling tasks respectively, merged tasks 1, 2, and 3 for the 3 storage shards are obtained. Merged task 1 and merged task 2 are assigned to computing node 1, and merged task 3 is assigned to computing node 2; computing node 1 pulls data from storage shard 1 and storage shard 2 respectively. For the 12 * 60 data pulled from storage shard 1, the maximum value, minimum value, median value, and average value are calculated once per hour, and for the 12 * 60 data pulled from storage shard 2, the maximum value, minimum value, median value, and average value are calculated once per hour; computing node 2 pulls data from storage shard 3 and calculates the maximum value, minimum value, median value, and average value once per hour for the 12 * 60 data pulled.

[0150] The data pre - downsampling method in the embodiments of this application can be applied to any time - series database cluster with separated storage and computing, and is specifically used for the pre - downsampling scenario of the time - series database cluster.

[0151] The following introduces the implementation principle of the data pre - downsampling method in the embodiments of this application.

[0152] The data pre - downsampling method in the embodiments of this application can be implemented through the following aspects:

[0153] I. Task merging model

[0154] Assume that the user creates two pre - downsampling tasks: the source database (which can be understood as the storage node) is the same, the groupby time is the same, and other configurations are the same, only the SQL is different:

[0155] select min(field1),max(field1)from m1 group by(12m),*

[0156] select median(field1),mean(field1)from m1 group by(12m),*

[0157] When the data of m1 is distributed among 4 storage shards (partitions), a pre-downsampling task is divided into 4 shard tasks (PartitionTask). Among the 4 shard tasks of each of these two pre-downsampling tasks, there are multiple shard tasks that belong to the same storage shard. At this time, multiple shard tasks that belong to the same storage shard can be merged to obtain 4 merge tasks (MergeTask).

[0158] Suppose there are 2 computing nodes in the time series database cluster. At this time, there are 8 shard tasks, which can be merged into 4 merge tasks. These 4 merge tasks will be evenly distributed on these two computing nodes. As Figure 10 shown, the merge tasks associated with storage shard 1 and the merge tasks associated with storage shard 2 are assigned to computing node 1, and the merge tasks associated with storage shard 3 and the merge tasks associated with storage shard 4 are assigned to computing node 2. Computing node 1 needs to pull the write-ahead logs from storage shards 1 and 2, and computing node 2 needs to pull the write-ahead logs from storage shards 3 and 4.

[0159] II. Time Sampling Window Model

[0160] To ensure the real-time writing of pre-downsampled data and the balance of out-of-order data, each computing node allows storing the pre-downsampled sub-data of N time sampling windows. When the user requires strong real-time performance, N can be set to 1, that is, the sampled data of one time sampling window is immediately written to the storage node after the pre-downsampling process is completed, giving the strongest real-time performance; when there is out-of-order data in the time series data, N can be set to a value greater than 1 to tolerate the out-of-order data within N time sampling windows.

[0161] III. Merge Status Interface

[0162] The scheduling node can periodically obtain the execution status of the merge tasks on each computing node, parse the execution status of each shard task from the execution status of the merge tasks, and merge the execution status of the multiple shard tasks included in each pre-downsampling task to obtain the execution status of each pre-downsampling task.

[0163] It is understandable that the scheduling node can maintain the execution status of all pre-downsampling tasks in memory and provide an RPC interface. Each computing node can directly access the scheduling node to obtain the execution status of the pre-downsampling tasks that need to be obtained. The user can obtain the execution status of any pre-downsampling task based on the RPC interface. In this way, the abnormal status of the pre-downsampling task can be known in a timely manner. For example, when the number of the above N time sampling windows is small, resulting in the loss of out-of-order data points in the pre-downsampling task and inaccurate pre-downsampling results.

[0164] IV. Disaster recovery capability

[0165] The computing node is stateless. In the pre-downsampling process, data is stored in the memory of the computing node and will only be written into the storage node after the sampling data within the time sampling window has completed the pre-downsampling process. When an abnormal situation occurs in the computing node, such as power failure, bug, or downtime, it may cause the loss of memory data, which will lead to a breakpoint in the pre-downsampling task. At this time, the scheduler node can determine the lost pre-downsampling data and actively call SQL to obtain the original data of the lost pre-downsampling data from the storage node, and then re-execute the corresponding pre-downsampling process on the original data.

[0166] V. Scalability

[0167] When the user creates a new pre-downsampling task, in principle, a new sharding task should be added to each merging task in the first part above (add a new sharding task). If the new sharding task is added to the corresponding merging task, the original merging task needs to be stopped first and a new merging task needs to be created. There are two problems in this process: First, it will trigger a large amount of SQL to supplement the lost data in memory; Second, it is uncertain whether the memory size occupied by the new sharding task will exceed the upper limit of the computing node's memory after being merged into the new merging task.

[0168] To address the above situation, as Figure 11 shown, the embodiment of the present application introduces a sub-task. A merging task contains multiple sub-tasks, and a sub-task corresponds to multiple sharding tasks. When there are already two sharding tasks in an existing merging task and a new sharding task is added, the merging task contains two sub-tasks. One sub-task corresponds to two sharding tasks t1 and t2, and the other sub-task corresponds to a new sharding task t3. These two sub-tasks run on computing node 1 and computing node 2 respectively. When the scheduling node finds that the memory occupied by these three sharding tasks is not greater than the memory of computing node 1, the new sharding task is merged with the above merging task to obtain a new merging task to reduce the pull traffic of the write-ahead log, otherwise the merging is not performed.

[0169] Generally, there are tens of thousands of pre-downsampling tasks in an online time-series data cluster, and the number of shard tasks included in each pre-downsampling task ranges from 8 to 108. After adopting the solution of the embodiments of the present application, the memory consumption of all pre-downsampling tasks is concentrated on the computing nodes. The computing nodes are stateless and can be infinitely scaled, thus solving the memory problem, high load problem, and scalability problem of the storage nodes caused by pre-downsampling.

[0170] In practical applications, different database tables often adopt the same downsampling task. That is to say, these pre-downsampling tasks only have different SQLs. In this case, the merging of a large number of pre-downsampling tasks can reduce the WAL pulling traffic to 1% of the original.

[0171] Based on the same inventive concept as the above method embodiments of the present application, an embodiment of the present application also provides a database pre-downsampling device. The principle of the device to solve problems is similar to that of the method in the above embodiments. Therefore, the implementation of the device can refer to the implementation of the above method, and the repeated parts will not be described again.

[0172] Refer to Figure 12 As shown, a database pre-downsampling device 1200 provided by an embodiment of the present application is applied to a scheduling node in a time-series database cluster. The time-series database cluster further includes at least one storage node and at least one computing node. The device includes:

[0173] A task acquisition unit 1201, configured to obtain a plurality of preset pre-downsampling tasks after writing time-series data into at least one storage shard of a storage node; wherein, each pre-downsampling task is used to perform corresponding pre-downsampling processing on the time-series data, and at least one shard task included in each pre-downsampling task is respectively associated with a corresponding storage shard;

[0174] A merging unit 1202, configured to merge a plurality of shard tasks associated with the same storage shard among at least one shard task of each of the plurality of pre-downsampling tasks, to obtain at least one merged task;

[0175] A processing unit 1203, configured to perform the following operations respectively for at least one merged task: instruct a computing node to pull target data from the time-series data written into the storage shard associated with a merged task, and perform corresponding pre-downsampling processing on the target data according to a merged task, and write the obtained pre-downsampled data into the storage node.

[0176] In the embodiments of the present application, the pre-downsampling task is executed in a computing node to save the memory of the storage node and reduce the memory pressure on the storage node. In order to minimize the network traffic for the computing node to pull data from the storage node, among at least one shard task of each of multiple pre-downsampling tasks, multiple shard tasks associated with the same storage shard are merged to obtain at least one merged task. For each merged task, a computing node is instructed to pull target data from the storage shard associated with the merged task from the written time series data, and perform corresponding pre-downsampling processing on the target data according to the merged task. In this way, instead of pulling data from the corresponding storage shard once for each shard task, data is pulled from the corresponding storage shard once for each merged task, greatly reducing the number of data pulls, and thus reducing the consumed network traffic.

[0177] Therefore, the embodiments of the present application can save the memory of the storage node and reduce the load of the processor with less consumed network traffic, relieve the memory pressure and load pressure of the storage node, ensure that both the pre-downsampling process and the data query process can be executed smoothly, and thus improve the availability of the time series database cluster.

[0178] Optionally, the time sampling windows corresponding to each of the multiple pre-downsampling tasks are the same;

[0179] When performing corresponding pre-downsampling processing on the target data according to a merged task, the processing unit 1203 is specifically configured to:

[0180] During the process of sampling the target data according to the time sampling window, for each piece of sampled data obtained within a time sampling window, corresponding pre-downsampling processing is performed on the sampled data according to a merged task to obtain corresponding pre-downsampled sub-data;

[0181] Whenever the pre-downsampled sub-data corresponding to N time sampling windows are obtained, the N pieces of pre-downsampled sub-data are written into the storage node; where N is an integer greater than or equal to 1.

[0182] Optionally, when N is an integer greater than 1, after obtaining the sampled data within the latest time sampling window, the processing unit 1203 is further configured to:

[0183] If there are out-of-order data points in the sampled data within the latest time sampling window, and the time tags carried by the out-of-order data points are within any one of the first N - 1 time sampling windows, insert the out-of-order data points into the sampled data within any one of the time sampling windows, and delete the out-of-order data points from the latest sampled data;

[0184] According to a merging task, re - perform the corresponding pre - downsampling process on the sampled data after inserting out - of - order data points, and perform the corresponding pre - downsampling process on the latest sampled data after deleting out - of - order data points.

[0185] Optionally, during the process of respectively performing the corresponding pre - downsampling process on the target data according to at least one merging task, the processing unit 1203 is further configured to:

[0186] For each of the at least one merging task, respectively perform the following operations: obtain the execution status of a merging task from the corresponding computing node every set period, where the execution status of a merging task includes: the execution status of each of the multiple shard tasks in a merging task;

[0187] For each of the multiple pre - downsampling tasks, respectively perform the following operations: merge the execution status of each of the at least one shard task included in a pre - downsampling task every set period to obtain the execution status of a pre - downsampling task.

[0188] Optionally, the apparatus further includes:

[0189] A request receiving unit, configured to obtain the latest execution status of any pre - downsampling task when receiving a status query request of any pre - downsampling task sent by the query end;

[0190] A status sending unit, configured to send the latest execution status of any pre - downsampling task to the corresponding query end.

[0191] Optionally, after sending the latest execution status of any pre - downsampling task to the corresponding query end, the processing unit 1203 is further configured to:

[0192] Receive a supplementary task of any pre - downsampling task sent by the query end; where the supplementary task is sent when there is abnormal pre - downsampled data in the latest execution status;

[0193] Instruct the corresponding computing node to re - perform the corresponding pre - downsampling process on the first original data corresponding to the abnormal pre - downsampled data according to the supplementary task, where the first original data is included in the time - series data.

[0194] Optionally, the processing unit 1203 is further configured to:

[0195] When an abnormal condition occurs in a computing node, obtain the execution status of each of the at least one merging task in the computing node;

[0196] When it is determined according to the execution status of any merging task that there is missing pre - downsampled data in any merging task, obtain the second original data corresponding to the missing pre - downsampled data from the storage node, where the second original data is included in the time - series data;

[0197] After the computing node recovers, it instructs the computing node to re - execute the corresponding pre - downsampling process on the second original data according to any merging task, and write the obtained pre - downsampled data into the storage node.

[0198] Optionally, the processing unit 1203 is further configured to:

[0199] When a new pre - downsampling task is obtained, the following operations are respectively performed for at least one shard task included in the new pre - downsampling task:

[0200] When the storage shard associated with a new shard task has a target merging task, if it is determined that the memory used to execute a new shard task is not greater than the remaining memory of the computing node where the target merging task is located, then divide a new shard task into the target merging task.

[0201] Optionally, when it is determined that the memory used to execute a new shard task is not greater than the remaining memory of the computing node where the target merging task is located, and then divide a new shard task into the target merging task, the processing unit 1203 is specifically configured to:

[0202] Instruct other computing nodes other than the computing node where the target merging task is located to execute a new shard task;

[0203] When it is determined that the memory occupied by a new shard task in other computing nodes is not greater than the remaining memory of the computing node where the target merging task is located, transfer a new shard task from other computing nodes to the computing node where the target merging task is located.

[0204] Based on the same inventive concept as the above - mentioned method embodiment of the present application, an apparatus for database pre - downsampling is further provided in the embodiment of the present application. The principle of the apparatus for solving problems is similar to that of the method in the above - mentioned embodiment. Therefore, the implementation of the apparatus can refer to the implementation of the above - mentioned method, and the repeated parts will not be described again.

[0205] Refer to Figure 13 As shown, an apparatus 1300 for database pre - downsampling provided in the embodiment of the present application is applied to a computing node of a time - series database cluster. The apparatus includes:

[0206] A task acquisition unit 1301, configured to acquire at least one merging task. Each merging task is determined by the following method: after writing time - series data into at least one storage shard of the storage node of the time - series database cluster, merging multiple shard tasks associated with the corresponding storage shard among a preset plurality of pre - downsampling tasks; where each pre - downsampling task is used to perform corresponding pre - downsampling processing on the time - series data, and at least one shard task included in each pre - downsampling task is respectively associated with a corresponding storage shard;

[0207] The processing unit 1302 is configured to perform the following operations respectively for at least one merging task: pull the target data in the written timing data from the storage slices associated with a merging task, perform corresponding pre-downsampling processing on the target data according to a merging task, and write the obtained pre-downsampled data into the storage node.

[0208] Optionally, the time sampling windows corresponding to multiple pre-downsampling tasks are the same;

[0209] When performing corresponding pre-downsampling processing on the target data according to a merging task, the processing unit 1302 is specifically configured to:

[0210] During the process of sampling the target data according to the time sampling window, for each piece of sampled data obtained within a time sampling window, perform corresponding pre-downsampling processing on the sampled data according to a merging task to obtain corresponding pre-downsampled sub-data;

[0211] Whenever the pre-downsampled sub-data corresponding to N time sampling windows are obtained, write the N pre-downsampled sub-data into the storage node; where N is an integer greater than or equal to 1.

[0212] Optionally, when N is an integer greater than 1, the processing unit 1302 is further configured to perform the following steps after obtaining the sampled data within the latest time sampling window:

[0213] If there are out-of-order data points in the sampled data within the latest time sampling window, and the time tags carried by the out-of-order data points are within any one of the previous N - 1 time sampling windows, insert the out-of-order data points into the sampled data within any one of the time sampling windows, and delete the out-of-order data points from the latest sampled data;

[0214] Perform corresponding pre-downsampling processing on the sampled data after inserting the out-of-order data points again according to a merging task, and perform corresponding pre-downsampling processing on the latest sampled data after deleting the out-of-order data points.

[0215] For the convenience of description, the above parts are divided into various modules (or units) according to functions and described separately. Of course, when implementing the present application, the functions of the various modules (or units) can be implemented in the same or multiple software or hardware.

[0216] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of that module or unit.

[0217] After introducing the database pre-downsampling method and apparatus of the exemplary embodiments of the present application, next, an electronic device according to another exemplary embodiment of the present application will be introduced.

[0218] Based on the same inventive concept as the above method embodiments, an electronic device is also provided in the embodiments of the present application. In one embodiment, the electronic device can be a server, such as Figure 1 the scheduling node 120 shown. Refer to Figure 14 As shown, the electronic device 1400 can at least include a processor 1401 and a memory 1402. Among them, the memory 1402 stores a computer program, and when the computer program is executed by the processor 1401, the processor 1401 is caused to execute the steps of any one of the above database pre-downsampling methods.

[0219] In some possible implementation manners, the electronic device of the embodiments of the present application can at least include at least one processor and at least one memory. Among them, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps in the database pre-downsampling method according to various exemplary embodiments of the present application described above in this specification. For example, the processor can execute the steps as Figure 2 shown therein.

[0220] Next, refer to Figure 15 to describe the electronic device 1500 according to this embodiment of the present application. Figure 15 The electronic device 1500 is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present application.

[0221] As Figure 15 shown, the electronic device 1500 is presented in the form of a general-purpose electronic device. The components of the electronic device 1500 can include but are not limited to: at least one processing unit 1510, at least one storage unit 1520, and a bus 1530 connecting different system components (including the storage unit 1520 and the processing unit 1510).

[0222] The bus 1530 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, a processor bus, or a local bus using any of the various bus architectures.

[0223] The storage unit 1520 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 1521 and / or cache memory 1522, and may further include read only memory (ROM) 1523.

[0224] The storage unit 1520 may further include a program / utility 1525 having a set (at least one) of program modules 1524. Such program modules 1524 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination thereof may include an implementation of a network environment.

[0225] The electronic device 1500 may also communicate with one or more external devices 1540 (such as a keyboard, a pointing device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 1500, and / or may communicate with any device that enables the electronic device 1500 to communicate with one or more other electronic devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 1550. In addition, the electronic device 1500 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 1560. As shown in the figure, the network adapter 1560 communicates with other modules for the electronic device 1500 through the bus 1530. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1500, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0226] In some possible embodiments, various aspects of the database pre-downsampling method provided by the present application may also be implemented in the form of a computer program product, which includes a computer program. When the computer program product runs on an electronic device, the computer program is used to cause the electronic device to execute the steps in the database pre-downsampling method according to various exemplary embodiments of the present application described above in this specification. For example, the electronic device may execute the steps as shown in Figure 2 shown.

[0227] A computer program product may employ any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0228] The program product of the embodiments of the present application may employ a portable compact disc read-only memory (CD-ROM) and include a computer program, and may run on an electronic device. However, the program product of the present application is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0229] The readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable computer program. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0230] The computer program contained on the readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0231] The computer program for performing the operations of the present application can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The computer program can be executed entirely on the user's electronic device, partially on the user's electronic device, executed as a stand-alone software package, partially on the user's electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In the case of a remote electronic device, the remote electronic device can be connected to the user's electronic device through any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external electronic device (e.g., connected through the Internet using an Internet service provider).

[0232] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0233] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the shown operations must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0234] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable computer programs.

[0235] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one or more flows and / or Figure 1 blocks specified in the flowchart and / or block diagram.

[0236] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device that realizes the functions specified in Figure 1 one or more flows and / or Figure 1 blocks specified in the flowchart and / or block diagram.

[0237] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the process Figure 1 in one process or a plurality of processes and / or boxes Figure 1 steps for the functions specified in one box or a plurality of boxes.

[0238] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present application.

[0239] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A database pre-downsampling method, characterized in that, Applied to a scheduling node in a time-series database cluster, the time-series database cluster further includes at least one storage node and at least one computing node, and the method includes: After writing time-series data into at least one storage shard of a storage node, obtaining a plurality of preset pre-downsampling tasks; wherein each pre-downsampling task is used to perform corresponding pre-downsampling processing on the time-series data, and at least one shard task included in each pre-downsampling task is respectively associated with a corresponding storage shard; Merging a plurality of shard tasks associated with the same storage shard among at least one shard task of each of the plurality of pre-downsampling tasks to obtain at least one merged task; For each of the at least one merged task, respectively perform the following operations: instruct a computing node to pull target data from the written time-series data in the storage shard associated with the merged task, and perform corresponding pre-downsampling processing on the target data according to the merged task, and write the obtained pre-downsampled data into the storage node.

2. The method according to claim 1, characterized in that, The time sampling windows corresponding to each of the plurality of pre-downsampling tasks are the same; The performing corresponding pre-downsampling processing on the target data according to the merged task includes: During the process of sampling the target data according to the time sampling window, each time sampling data within a time sampling window is obtained, corresponding pre-downsampling processing is performed on the sampling data according to the merged task to obtain corresponding pre-downsampled sub-data; Whenever pre-downsampled sub-data corresponding to N time sampling windows is obtained, write the N pre-downsampled sub-data into the storage node; wherein N is an integer greater than or equal to 1.

3. The method according to claim 2, wherein When N is an integer greater than 1, after obtaining the sampling data within the latest time sampling window, it further includes: If there are out-of-order data points in the sampling data within the latest time sampling window, and the time tag carried by the out-of-order data point is within any one of the first N - 1 time sampling windows, insert the out-of-order data point into the sampling data within the any one of the time sampling windows, and delete the out-of-order data point from the latest sampling data; According to the merged task, re-perform corresponding pre-downsampling processing on the sampling data after inserting the out-of-order data point, and perform corresponding pre-downsampling processing on the latest sampling data after deleting the out-of-order data point.

4. The method according to any one of claims 1 to 3, characterized in that, During the process of respectively performing corresponding pre-downsampling processing on the target data according to the at least one merged task, it further includes: For each of the at least one merged task, respectively perform the following operations: obtain the execution status of a merged task from the corresponding computing node every set period, and the execution status of the merged task includes: the execution status of each of the plurality of shard tasks in the merged task; For each of the plurality of pre-downsampling tasks, respectively perform the following operations: merge the execution status of each of the at least one shard task included in a pre-downsampling task every the set period to obtain the execution status of the pre-downsampling task.

5. The method according to claim 4, characterized in that The method further includes: When receiving a status query request for any pre-downsampling task sent by a query end, obtain the latest execution status of the any pre-downsampling task; Send the latest execution status of the any pre-downsampling task to the corresponding query end.

6. The method according to claim 5, characterized in that After sending the latest execution status of the any pre-downsampling task to the corresponding query end, it further includes: Receive a supplementary task of the any pre-downsampling task sent by the query end; wherein, the supplementary task is sent when there is abnormal pre-downsampled data in the latest execution status; Instruct a corresponding computing node to re-execute the corresponding pre-downsampling process on the first original data corresponding to the abnormal pre-downsampled data according to the supplementary task, and the first original data is included in the time series data.

7. The method according to any one of claims 4, characterized in that The method further includes: When an abnormal condition occurs in a computing node, obtain the execution status of each of at least one merging task in the one computing node; When it is determined according to the execution status of any merging task that there is missing pre-downsampled data in the any merging task, obtain the second original data corresponding to the missing pre-downsampled data from a storage node, and the second original data is included in the time series data; After the one computing node returns to normal, instruct the one computing node to re-execute the corresponding pre-downsampling process on the second original data according to the any merging task, and write the obtained pre-downsampled data into the storage node.

8. The method according to any one of claims 1 to 3, characterized in that, The method further includes: When a new pre-downsampling task is obtained, for each of at least one new sharding task included in the new pre-downsampling task, perform the following operations respectively: When a storage shard associated with a new sharding task corresponds to a target merging task, if it is determined that the memory used to execute the one new sharding task is not greater than the remaining memory of the computing node where the target merging task is located, divide the one new sharding task into the target merging task.

9. The method according to claim 8, wherein The if it is determined that the memory used to execute the one new sharding task is not greater than the remaining memory of the computing node where the target merging task is located, divide the one new sharding task into the target merging task includes: Instruct other computing nodes other than the computing node where the target merging task is located to execute the one new sharding task; When it is determined that the memory occupied by the one new sharding task in the other computing nodes is less than the remaining memory of the computing node where the target merging task is located, transfer the one new sharding task from the other computing nodes to the computing node where the target merging task is located.

10. A database pre-downsampling method, characterized in that Applied to a computing node of a time series database cluster, the method includes: Obtain at least one merging task, where each merging task is determined as follows: after writing time series data into at least one storage shard of a storage node in the time series database cluster, merge multiple shard tasks associated with the corresponding storage shard among a preset plurality of pre-downsampling tasks; wherein, each pre-downsampling task is used to perform corresponding pre-downsampling processing on the time series data, and at least one shard task included in each pre-downsampling task is respectively associated with a corresponding storage shard; For each of the at least one merging task, perform the following operations respectively: pull the target data in the written time series data from the storage shard associated with one merging task, and perform corresponding pre-downsampling processing on the target data according to the one merging task, and write the obtained pre-downsampled data into the storage node.

11. The method according to claim 10, wherein The time sampling windows corresponding to the plurality of pre-downsampling tasks are the same; The performing corresponding pre-downsampling processing on the target data according to the one merging task includes: During the process of sampling the target data according to the time sampling window, every time sampling data within a time sampling window is obtained, perform corresponding pre-downsampling processing on the sampling data according to the one merging task to obtain corresponding pre-downsampled sub-data; Whenever pre-downsampled sub-data corresponding to N time sampling windows is obtained, write the N pre-downsampled sub-data into the storage node; where N is an integer greater than or equal to 1.

12. The method according to claim 11, characterized in that When N is an integer greater than 1, after obtaining the sampling data within the latest time sampling window, it further includes: If there are out-of-order data points in the sampling data within the latest time sampling window, and the time tag carried by the out-of-order data points is within any one of the previous N - 1 time sampling windows, insert the out-of-order data points into the sampling data within the any one of the time sampling windows, and delete the out-of-order data points from the latest sampling data; According to the one merging task, re-perform corresponding pre-downsampling processing on the sampling data after inserting the out-of-order data points, and perform corresponding pre-downsampling processing on the latest sampling data after deleting the out-of-order data points.

13. A database pre-downsampling device, characterized in that, Applied to a scheduling node in a time series database cluster, the time series database cluster further includes at least one storage node and at least one computing node, and the apparatus includes: A task acquisition unit, configured to obtain a preset plurality of pre-downsampling tasks after writing time series data into at least one storage shard of a storage node; wherein, each pre-downsampling task is used to perform corresponding pre-downsampling processing on the time series data, and at least one shard task included in each pre-downsampling task is respectively associated with a corresponding storage shard; A merging unit, configured to merge multiple shard tasks associated with the same storage shard among at least one shard task of each of the plurality of pre-downsampling tasks to obtain at least one merging task; A processing unit for respectively performing the following operations for the at least one merging task: instructing a computing node to pull target data in the written time-series data from a storage shard associated with one merging task, performing corresponding pre-downsampling processing on the target data according to the one merging task, and writing the obtained pre-downsampled data into a storage node.

14. A database pre-downsampling device, characterized in that, A computing node applied to a time-series database cluster, the device comprising: A task acquisition unit for acquiring at least one merging task, each merging task being determined by the following method: after writing time-series data into at least one storage shard of a storage node of the time-series database cluster, merging a plurality of shard tasks associated with the corresponding storage shard among a plurality of preset pre-downsampling tasks; wherein each pre-downsampling task is used for performing corresponding pre-downsampling processing on the time-series data, and at least one shard task included in each pre-downsampling task is respectively associated with a corresponding storage shard; A processing unit for respectively performing the following operations for the at least one merging task: pulling target data in the written time-series data from a storage shard associated with one merging task, performing corresponding pre-downsampling processing on the target data according to the one merging task, and writing the obtained pre-downsampled data into a storage node.

15. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 9.

16. A computer-readable storage medium, characterized in that, It includes a computer program, and when the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the method according to any one of claims 1 to 9.

17. A computer program product, characterized in that, It includes a computer program, and the computer program is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, causing the electronic device to execute the steps of the method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Mass data display method, device and equipment based on browser client and medium

    CN121166997A

  • Chart display method and device for machine data and storage medium

    CN121901472A