Partition distribution method and device for time sequence database, equipment, medium and product

By introducing the concepts of sequence partition slots and time partition slots, the partition allocation method of timing databases is optimized, and the performance and efficiency problems of timing databases in the industrial Internet of Things application scenarios in the prior art are solved, and more efficient data access and write throughput is achieved.

CN120216503APending Publication Date: 2025-06-27TSINGHUA UNIVERSITY +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510205457.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the industrial IoT application scenarios, the performance and efficiency of the partition allocation method of existing timing databases are poor, resulting in low data access efficiency and reduced cluster write throughput.

Method used

Using the concepts of sequence partition slots and time partition slots, time series data is balancedly allocated to a limited sequence partition slot, and data partitions are formed through pairing, the data storage structure is optimized, active data migration is avoided, and data distribution of data replica groups is dynamically equalized.

Benefits of technology

Improve data access efficiency, reduce cluster burden, improve write throughput, and realize dynamic balance of data distribution to ensure cluster load balancing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216503A_ABST
    Figure CN120216503A_ABST
Patent Text Reader

Abstract

The invention provides a partition distribution method and device for a time sequence database, equipment, a medium and a product, and the method comprises the steps: analyzing received time sequence data, and obtaining a plurality of time sequence data points; for each time sequence data point, determining a sequence partition slot and a time partition slot corresponding to the time sequence data point; matching the sequence partition slots corresponding to the time sequence data points with the time partition slots corresponding to the time sequence data points to obtain data partitions; a data distribution table is inquired, a data copy group corresponding to the sequence partition slot is determined, and the data distribution table is used for recording a mapping relation between the sequence partition slot and the data copy group; and writing the data partitions into the data copy groups corresponding to the sequence partition slots, and updating a data partition table which is used for recording a mapping relationship between the data partitions and the data copy groups. According to the scheme, the performance and efficiency of the time sequence database are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer data management, and in particular, to a method, apparatus, device, medium and product for partitioning and allocating a time series database. Background Art

[0002] Time series data models are widely used in the field of industrial Internet of Things. Time series data is data that is collected in real time by sensors in an industrial production environment and recorded in timestamp order. Generally, data of the same time series has the same type, and each time series data point contains a timestamp and a data value. For example, metrics such as temperature, pressure, and speed generated by real-time monitoring of the operating status and performance parameters of sensors. Specifically, time series are usually used in scenarios such as fault diagnosis and production optimization.

[0003] Currently, most distributed database systems have a partitioning function. The common strategy for relational databases is to determine the data partition to which each row of data belongs based on the key of each row of data. For example, the consistent hashing algorithm used by Cassandra calculates a hash value for the data key, and then finds the next virtual node on the hash ring as the data partition for that row of data. Another example is that in HDFS, the data partitions to which each file block belongs are stored in the form of a directory tree. However, the classical partitioning strategies of relational databases have the following problems: (1) In an industrial Internet of Things production environment, tens of millions of time series are often deployed. The partitioning granularity of such strategies is relatively fine, making it difficult to load all the partitioning information generated at the time series granularity into memory, resulting in poor data access efficiency; (2) Such strategies commonly use data migration techniques to balance the cluster load. However, data migration will bring an additional burden to the cluster and reduce the cluster write throughput, which is important in industrial Internet of Things scenarios.

[0004] Therefore, in the industrial Internet of Things application scenario, the performance and efficiency of existing time series database partitioning and allocation methods are poor. Summary of the Invention

[0005] The present invention provides a method, apparatus, device, medium and product for partitioning and allocating a time series database, so as to solve the defect that the performance and efficiency of existing solutions are poor in the industrial Internet of Things application scenario.

[0006] The present invention provides a method for partitioning and allocating a time series database, including: parsing the received time series data to obtain a plurality of time series data points; wherein each time series data point includes a timestamp and a data value; for each time series data point, applying a predefined sequence partition slot algorithm to determine the sequence partition slot corresponding to the time series data point, and determining the time partition slot corresponding to the time series data point according to the timestamp of the time series data point; pairing the sequence partition slot corresponding to the time series data point with the time partition slot corresponding to the time series data point to obtain a data partition; querying a data allocation table to determine the data replica group corresponding to the sequence partition slot, the data allocation table being used to record the mapping relationship between the sequence partition slot and the data replica group, thereby guiding the construction of the data partition table; writing the data partition into the data replica group corresponding to the sequence partition slot, and updating the data partition table, the data partition table being used to record the mapping relationship between the data partition and the data replica group.

[0007] According to a method for partitioning and allocating a time series database provided by the present invention, the method further includes: setting the data survival time before the time series database cluster is started for the first time; after determining the time partition slot corresponding to the time series data point according to the timestamp of the time series data point, the method further includes: based on the data survival time, executing a sliding window algorithm to determine an expired time partition slot; clearing the data partition table records and data partitions related to the expired time partition slot.

[0008] According to a method for partitioning and allocating a time series database provided by the present invention, the method further includes: if the number of data replica groups in the time series database cluster changes, updating the data allocation table.

[0009] According to a method for partitioning and allocating a time series database provided by the present invention, the step of updating the data allocation table if the number of data replica groups in the time series database cluster changes includes: traversing each allocated sequence partition slot in a random order, and for each allocated sequence partition slot, if the sequence partition slot meets the clearing condition, clearing the mapping relationship between the sequence partition slot and the data replica group in the data allocation table, and taking the sequence partition slot as a sequence partition slot to be allocated; wherein the clearing condition is that the data replica group corresponding to the sequence partition slot is deleted or the number of sequence partition slots held by the data replica group corresponding to the sequence partition slot exceeds a dynamic threshold; traversing each sequence partition slot to be allocated in a random order, and for each sequence partition slot to be allocated, randomly allocating a data replica group that holds fewer sequence partition slots, and adding the mapping relationship between the sequence partition slot and the data replica group to the data allocation table.

[0010] A method for partitioning and allocating a time-series database according to the present invention, the updated data partition table includes: querying the data partition table, if there is no corresponding data replica group for the data partition, performing an allocation process, allocating a data replica group for the data partition, and writing the mapping relationship between the data partition and the data replica group into the data partition table; wherein, the allocation process includes: querying the data replica group corresponding to the sequence partition slot where the data partition is located in the data allocation table, and recording the allocation relationship in the data partition table.

[0011] A method for partitioning and allocating a time-series database according to the present invention, the method further includes: before the time-series database cluster is started for the first time, configuring the number of sequence partition slots in the time-series database cluster, the number of data replicas held by each node, the sequence partition slot algorithm, and the time interval between time partitions; after the time-series database cluster is started for the first time, initializing each sequence partition slot; whenever the number of data replica groups in the time-series database cluster changes, allocating the corresponding data replica group for each sequence partition slot, and generating the data allocation table.

[0012] The present invention also provides a device for partitioning and allocating a time-series database, the device includes: a parsing module, configured to parse the received time-series data to obtain a plurality of time-series data points; wherein, each time-series data point includes a timestamp and a data value; a determining module, configured to, for each time-series data point, apply a predefined sequence partition slot algorithm to determine the sequence partition slot corresponding to the time-series data point, and determine the time partition slot corresponding to the time-series data point according to the timestamp of the time-series data point; a pairing module, configured to pair the sequence partition slot corresponding to the time-series data point with the time partition slot corresponding to the time-series data point to obtain a data partition; a querying module, configured to query the data allocation table to determine the data replica group corresponding to the sequence partition slot, the data allocation table being used to record the mapping relationship between the sequence partition slot and the data replica group; a writing module, configured to write the data partition into the data replica group corresponding to the sequence partition slot, and update the data partition table, the data partition table being used to record the mapping relationship between the data partition and the data replica group.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor, where when the processor executes the computer program, it implements the method for partitioning and allocating a time-series database as described in any one of the above.

[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method for partitioning and allocating a time-series database as described in any one of the above.

[0015] The present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the partition allocation method of the time series database as described in any one of the above.

[0016] The partition allocation method, device, equipment, medium and product of the time series database provided by the present invention introduce the concepts of sequence partition slots and time partition slots, evenly allocate a large amount of time series data to a limited number of sequence partition slots to reduce the volume of partition information. By pairing sequence partition slots with time partition slots to form data partitions, the storage structure of the data is optimized, making the organization of the data more orderly, facilitating management and access, and improving the efficiency of data access. Further, this method avoids active data migration to reduce the burden on the cluster, especially in write-intensive scenarios such as industrial Internet of Things, which helps to improve the write throughput of the cluster. Further, by querying and updating the data allocation table, the data distribution between data replica groups can be dynamically balanced to ensure cluster load balancing. In summary, the solution of the present invention improves the performance and efficiency of the time series database in the industrial Internet of Things application scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is one of the flow diagrams of the partition allocation method of the time series database provided by the present invention.

[0019] Figure 2 is the flow diagram of the maintenance of the data allocation table provided by the present invention.

[0020] Figure 3 is the second flow diagram of the partition allocation method of the time series database provided by the present invention.

[0021] Figure 4 is the structural diagram of the partition allocation device of the time series database provided by the present invention.

[0022] Figure 5 is the structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts fall within the protection scope of the present invention.

[0024] The technical solutions of this application and how the technical solutions of this application solve the above technical problems will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following combines Figures 1 - 3 to describe the partition allocation method of the time series database of the present invention.

[0025] In practical applications, the execution subject of the partition allocation method of the time series database can be a partition allocation device of the time series database. There are various implementation manners of the partition allocation device of the time series database. For example, it can be implemented through a computer program, such as an application software, etc.; or, for example, a chip, etc. It can also be implemented as a medium storing relevant computer programs, such as a USB flash drive, a cloud disk, etc.; or, it can also be implemented through an entity device integrated or installed with relevant computer programs, such as a server, an intelligent device, etc. The following takes the partition allocation device of the time series database as the execution subject of the partition allocation method of the time series database for example.

[0026] Figure 1 is one of the flow diagrams of the partition allocation method of the time series database provided by the present invention. As Figure 1 shown, the method includes the following steps 101 to step 106.

[0027] Step 101: Parse the received time series data to obtain multiple time series data points; where each time series data point includes a timestamp and a data value.

[0028] Among them, the time series data refers to the data stream of the time series, and this data may come from multiple sensors in the industrial Internet of Things, device logs, transaction records, and other various real-time data sources. Exemplarily, the time series data can be received through various data acquisition tools or interfaces, such as data write APIs, MQTT protocols, database logs, etc.

[0029] Further, after the partition allocation device of the time series database receives the time series data, it parses the received time series data to extract multiple time series data points, and each time series data point contains a timestamp and a corresponding data value.

[0030] Specifically, a timestamp is a date and time marker for each data point, which indicates when the data value was observed or recorded. The data value is the actual measured or recorded value associated with the timestamp, which can be an index value such as temperature, pressure, speed, or any numerical value that needs to be recorded and analyzed.

[0031] Exemplarily, the parsing process may involve steps such as format conversion (e.g., converting from JSON, CSV to an internal format), error detection and correction, data cleaning, etc., to ensure the accuracy and usability of the data.

[0032] Optionally, according to the actual application requirements and application scenarios, the parsed data may need further preprocessing, such as unit conversion, data normalization, etc., to meet the subsequent partition allocation requirements.

[0033] Step 102: For each time-series data point, apply a predefined sequence partition slot algorithm to determine the sequence partition slot corresponding to the time-series data point.

[0034] Among them, the predefined sequence partition algorithm determines the mapping rule between the time-series data point and the sequence partition slot. The predefined sequence partition algorithm is an algorithm that evenly distributes time-series data into a finite number of sequence partition slots. In practice, the selection of the sequence partition algorithm can be based on various factors, such as the data access pattern, the distribution characteristics of the data, etc. The sequence partition algorithm includes but is not limited to hash functions, range partition algorithms, list partition algorithms. Exemplarily, the predefined sequence partition algorithm can be a hash algorithm selected by the user according to the specific application scenario.

[0035] In practical applications, when the time-series database cluster is started for the first time, set the number of sequence partition slots and the number of data replicas held by each node. Exemplarily, the number of sequence partition slots in the time-series database cluster is n , and denote the set of cluster sequence partition slots as . Assume that the number of data replica groups in the current time-series dataset database cluster is m , and denote the set of cluster data replica groups as . Further, the sequence partition slots are evenly distributed to the data replica groups currently held by the cluster, thereby forming a data allocation table for recording the mapping relationship between the sequence partition slots and the data replica groups.

[0036] Specifically, for each time-series data point, applying the predefined sequence partition slot algorithm can determine the sequence partition slot corresponding to the time-series data point. It can be understood that evenly distributing a large amount of time series to a limited number of sequence partition slots in this embodiment can reduce the partition information volume.

[0037] Step 103: According to the timestamp of the time-series data point, determine the time partition slot corresponding to the time-series data point.

[0038] Among them, the time partition slot is a mechanism for partitioning data based on timestamps. It divides time into multiple intervals (such as one slot per hour, day, or week), and groups data according to these intervals. Combining the above description, each time series data point contains a timestamp indicating the exact time when the data was recorded.

[0039] In practical applications, a time partition slot is divided at regular time intervals. Specifically, the partition allocation device of the time series database calculates which time partition slot a data point belongs to according to a preset time interval (such as per hour or per day) and the timestamp of the data point. That is, the time partition corresponding to the time series data point is determined according to the time partition to which the timestamp of the time series data point belongs. For example, assuming the time interval is per hour, if the timestamp of the time series data point is 9:10:00 and the time partition to which this timestamp belongs is from 9:00:00 to 10:00:00, then this time series data point is assigned to the time partition slot corresponding to the time partition from 9:00:00 to 10:00:00.

[0040] It should be noted that users can configure the interval of time partitions according to the access pattern of data and business requirements. For example, for data that needs to be accessed frequently, a smaller time interval (such as per hour) may be selected to create more time partitions and thus improve the access parallelism; for historical data that is less accessed, a larger time interval (such as per day) may be selected to reduce the number of created time partitions and lower the cache cost.

[0041] Step 104: Pair the sequence partition slot corresponding to the time series data point with the time partition slot corresponding to the time series data point to obtain a data partition.

[0042] Furthermore, a sequence partition slot and a time partition slot are paired to form a data partition. In practical applications, for each time series data point, the partition allocation device of the time series database pairs the sequence partition slot determined in step 102 with the time partition slot determined in step 103, thereby obtaining a unique data partition.

[0043] In this embodiment, the data partition is a combination of a sequence partition slot and a time partition slot, which represents the storage location of the time series data point in the database. This pairing mechanism ensures that the data organization method takes into account both the source (sequence) of the data and the time attributes of the data.

[0044] It can be understood that by pairing the sequence partition slot and the time partition slot, a two-dimensional data partition system is formed. The partition in the sequence dimension greatly reduces the cache cost of partition information and improves the performance of data operations; the partition in the time dimension provides support for the life cycle management of time series data.

[0045] Step 105: Query the data allocation table to determine the data replica group corresponding to the sequence partition slot. The data allocation table is used to record the mapping relationship between the sequence partition slot and the data replica group.

[0046] Step 106: Write the data partition into the data replica group corresponding to the sequence partition slot and update the data partition table. The data partition table is used to record the mapping relationship between the data partition and the data replica group.

[0047] Exemplarily, the data allocation table can be designed as a key-value pair structure, where the key is the identifier of the sequence partition slot and the value is the identifier of the corresponding data replica group. Further, to improve the query efficiency, the data allocation table can be implemented using efficient data structures such as hash tables or balanced trees to support fast lookup. Further, when the data replica groups in the cluster change (such as adding or deleting replica groups), the data allocation table needs to be dynamically updated to reflect the new mapping relationship.

[0048] It should be noted that by maintaining the data allocation table, the system can ensure that data is evenly distributed among different data replica groups, avoiding overloading of some replica groups while other replica groups are idle. The design of the data allocation table enables the system to flexibly adapt to changes in the number of data replica groups, improving the scalability of the system. In a distributed system, the data allocation table also helps with fault recovery because it records the storage location of data. Even if a data replica group fails, the system can quickly redistribute the data based on the information in the table.

[0049] It can be understood that the data allocation table is used to record the mapping relationship between the sequence partition slot and the data replica group. Therefore, by querying the data allocation table, the data replica group corresponding to the sequence partition slot can be determined, that is, it can quickly locate which data replica group the data of each sequence partition slot should be stored in.

[0050] For example, a distributed time-series database is used to store data from multiple sensors. The time-series database is configured with 6 sequence partition slots ( , , , , ) and 3 data replica groups ( , , ). Specifically, the data replica groups corresponding to the sequence partition slots and in the data allocation table are , the data replica groups corresponding to the sequence partition slots and are , and the data replica groups corresponding to the sequence partition slots and The corresponding data replica group is .

[0051] Specifically, the timestamp of the time series data point a is 12:01:00. The predefined sequence partition slot algorithm is used to determine that the sequence partition slot corresponding to the time series data point a is . According to the timestamp of the time series data point, it is determined that the time partition slot corresponding to the time series data point a is the time partition slot corresponding to the time partition from 12:00:00 to 13:00:00. . The sequence partition slot corresponding to the time series data point a is paired with the time partition slot corresponding to the time series data point to obtain a data partition ( , ). Further, query the data allocation table to determine that the data replica group corresponding to the sequence partition slot is .

[0052] Further, write the data partition ( , ) into the data replica group , and at the same time record the mapping relationship between the data partition written into the data replica group and the data partition table. Among them, the data replica group is an important part of the distributed system design, which refers to a group of storage nodes (or servers) that jointly store copies of the same data.

[0053] Preferably, in order to further improve the management efficiency of time series data, in one example, the data partitions with the same sequence partition slot are stored centrally. Specifically, when the number of cluster data replica groups remains unchanged, the data partitions with the same sequence partition slot will be written into the same data replica group under the guidance of the data allocation table and recorded in the data partition table. It can be understood that since the data with the same sequence partition slot is centrally stored in the same data replica group, queries can locate the relevant data faster, reducing the need to transfer data across physical nodes. Therefore, the method of writing data partitions with the same sequence partition slot into the same data replica group not only improves the data access efficiency but also reduces the system operation and maintenance costs.

[0054] In this embodiment, by introducing the concepts of sequence partition slots and time partition slots, a large amount of time series data is evenly distributed to a limited number of sequence partition slots to reduce the volume of partition information. By pairing sequence partition slots with time partition slots to form data partitions, the storage structure of the data is optimized, making the organization of the data more orderly, facilitating management and access, and improving the efficiency of data access. Further, this method avoids active data migration to reduce the burden on the cluster, especially in write-intensive scenarios such as industrial Internet of Things, which helps to improve the write throughput of the cluster. Further, by querying and updating the data allocation table, the data distribution between data replica groups can be dynamically balanced to ensure cluster load balancing. Therefore, the solution of this embodiment improves the performance and efficiency of the time series database in the industrial Internet of Things application scenario.

[0055] In addition, in a possible implementation manner, the above-mentioned partition allocation method of the time series database further includes: Before the time series database cluster is started for the first time, set the data survival time; After the above step 103, the above method further includes: Based on the data survival time, execute a sliding window algorithm to determine the expired time partition slots; Clean up the data partition table records and data partitions related to the expired time partition slots.

[0056] Among them, the data survival time (Time to Live, abbreviated as TTL) refers to the length of time that data is stored and remains in a valid state. In practical applications, before the time series database cluster is started for the first time, the data survival time can be configured by itself. Whenever the cluster runs into a new time partition slot After that, if , then delete all Related data partition table records and data partitions. The cluster ensures that only the data partition table within the range of Is cached in memory and only the data partitions within the range of Is saved on disk.

[0057] Exemplarily, assume that the user configures the data survival time , and the current cluster runs into a new time partition slot After that, if , delete the corresponding data partition table records and data partitions. For example, the time interval between time partitions is set to 1 hour. The user configures the TTL to be 3 time partition slots, which means that the data will be automatically deleted after 3 hours. The current cluster runs into a new time partition slot , therefore, any data in the partition slot before Will be deleted.

[0058] In this embodiment, the expired time partition slots are cleaned up in a timely manner, which can reduce the memory pressure. Moreover, the active migration of the data partitions corresponding to the premature time partition slots is reduced. Therefore, the present invention can achieve load balancing without affecting the write throughput of the cluster.

[0059] In practical applications, when the time series database cluster is started for the first time, it is necessary to configure the necessary parameters and the sequence partition slot algorithm. As an example, in some possible embodiments, the above-mentioned partition allocation method for the time series database further includes: Before the time series database cluster is started for the first time, configure the number of sequence partition slots in the time series database cluster, the number of data replicas held by each node, the sequence partition slot algorithm, and the time interval between partitions; After the time series database cluster is started for the first time, initialize each sequence partition slot; Whenever the number of data replica groups in the time series database cluster changes, assign the corresponding data replica groups to each sequence partition slot and generate a data allocation table.

[0060] Among them, the time series database cluster is a special type of database cluster, which is optimized for processing and storing time series data. Time series data is a series of data points recorded in chronological order, usually generated by sensors, devices, applications, etc. Each data point contains a timestamp and a data value.

[0061] In practical applications, when the time series database cluster is started for the first time, configure the number of sequence partition slots in the cluster according to the actual application scenario, actual requirements, expected data volume, data access mode, etc. Generally, the number of sequence partition slots will not be changed thereafter.

[0062] Furthermore, define an algorithm for determining the sequence partition slot corresponding to the time series data point. Specifically, the sequence partition algorithms include but are not limited to hash functions, range partition algorithms, and list partition algorithms. Exemplarily, the predefined sequence partition algorithm can be a hash algorithm selected by the user according to the specific application scenario. Further, initialize each sequence partition slot and mark each sequence partition slot as to be assigned.

[0063] Furthermore, configure the time interval between partitions. Specifically, determine a suitable time interval between partitions according to the business requirements and data access mode. For example, if the data needs to be queried hourly, one time partition slot per hour can be selected.

[0064] Furthermore, determine the data replica group corresponding to each sequence partition slot and create a data allocation table to record these mapping relationships.

[0065] Exemplarily, when the time series database cluster is started for the first time, configure the number of sequence partition slots in the time series database cluster to ben , denote the set of sequence partition slots of the cluster sequence as , assume the number of current configured data replica groups is m , denote the set of cluster data replica groups as . The data allocation table can be described as a mapping . When , it means that the sequence partition slot is to be allocated, otherwise the sequence partition slot is already allocated. In addition, denote , which means the number of sequence partition slots held by the allocated data replica group in the current data allocation table.

[0066] In this embodiment, on the one hand, by configuring the number of sequence partition slots and the number of data replicas held by each node when the cluster starts, resources can be reasonably allocated according to the expected data volume and access pattern, improving resource utilization. On the other hand, defining the sequence partition slot algorithm and initializing each sequence partition slot can ensure that data write operations are quickly and evenly distributed in the cluster, improving data write efficiency. On the other hand, configuring the time interval makes the data be managed in an orderly manner according to time, thus improving data query performance. On the other hand, whenever the number of data replica groups in the time series database cluster changes, allocating corresponding data replica groups for each sequence partition slot helps to balance the load of the cluster, avoiding some nodes being overloaded while other nodes are idle, achieving load balancing. On the other hand, generating a data allocation table to record the mapping relationship between sequence partition slots and data replica groups helps to maintain data consistency and reliability. And the data allocation table provides a clear view of data distribution, simplifying data management and maintenance work. On the other hand, reasonable partition and replica group configuration enables the system to be more easily extended by adding nodes to cope with the growth of data volume, improving the scalability of the system. On the other hand, users can customize the sequence partition slot algorithm and time interval according to specific business requirements, making the system have good adaptability. Therefore, this embodiment provides an efficient, scalable, reliable and easy-to-manage solution, which is especially suitable for scenarios that need to process large-scale time series data.

[0067] In practical applications, when the number of data replica groups changes due to the change of the number of nodes in the cluster, it is necessary to re-determine the data replica group corresponding to each sequence partition slot and update the data allocation table so that each data replica group holds an equal number of sequence partition slots to achieve load balancing.

[0068] As an example, in a possible embodiment, the above-mentioned partition allocation method of the time series database further includes: If the number of data replica groups in the time series database cluster changes, update the data distribution table.

[0069] In practical applications, according to actual business requirements and actual application scenarios, such as the amount of data, the number of nodes in the cluster may be increased or decreased, and then the number of data replica groups in the cluster is recalculated based on the number of data replicas held by each node. In a time series database cluster, any change (increase or decrease) in the number of data replica groups requires updating the data distribution table to maintain balanced data distribution and high availability.

[0070] Specifically, when adding new data replica groups, a portion of the sequence partition slots in the existing data replica groups need to be assigned to the new data replica groups to ensure balanced data distribution. In practice, after updating the data replica group corresponding to each sequence partition slot, record the mapping relationship between each updated sequence partition slot and the data replica group to update the data distribution table.

[0071] As an example, the sequence partition slots in the data replica group with more sequence partition slots can be assigned to the new data replica group. For example, according to the data distribution table, the time series database is configured to have 4 sequence partition slots ( , , , ) and 3 data replica groups ( , , ). Specifically, the data replica groups corresponding to the sequence partition slots and are recorded in the data distribution table as , the data replica group corresponding to the sequence partition slot is , and the data replica group corresponding to the sequence partition slot is . After that, a new data replica group is added to the time series database. According to the data distribution table, the data replica group corresponds to two sequence partition slots, namely and . Any one of the sequence partition slots and can be assigned to the data replica group . For example, assign the sequence partition slot to the data replica group . Further, delete the assignment record of the sequence partition slot assigned to the data replica group in the data distribution table, and add the assignment record of the sequence partition slot assigned to the data replica group in the data distribution table The allocation record is used to update the data allocation table. It can be understood that after the update, the record sequence partition slots of the data allocation table The corresponding data replica group is , the sequence partition slot The corresponding data replica group is , the sequence partition slot The corresponding data replica group is , the sequence partition slot The corresponding data replica group is .

[0072] In contrast, when removing a data replica group, it is necessary to allocate the sequence partition slots in the removed replica group to the remaining replica groups to ensure the balanced distribution of data. In practice, after updating the data replica group corresponding to each sequence partition slot, the mapping relationship between each updated sequence partition slot and the data replica group is recorded to update the data allocation table.

[0073] Specifically, in one example, if the number of data replica groups in the time series database cluster changes, updating the data allocation table includes: Traverse each allocated sequence partition slot in a random order. For each allocated sequence partition slot, if the sequence partition slot meets the clearing condition, clear the mapping relationship between the sequence partition slot and the data replica group in the data allocation table, and use the sequence partition slot as the sequence partition slot to be allocated; wherein, the clearing condition is that the data replica group corresponding to the sequence partition slot is deleted or the number of sequence partition slots held by the data replica group corresponding to the sequence partition slot exceeds the dynamic threshold; Traverse each sequence partition slot to be allocated in a random order. For each sequence partition slot to be allocated, randomly allocate a data replica group with fewer held sequence partition slots, and add the mapping relationship between the sequence partition slot and the data replica group in the data allocation table.

[0074] To better understand the maintenance process of the data allocation table, the following combines Figure 2 To illustrate the maintenance process of the data allocation table by way of example. Figure 2 is a schematic diagram of the maintenance process of the data allocation table provided by the present invention. As Figure 2 shown, the maintenance process of the data allocation table is as follows.

[0075] Step 201, Initialize n sequence partition slots and mark them as to be allocated.

[0076] Step 202, Determine whether the number of cluster data replica groups has changed.

[0077] Step 203, If the number of cluster data replica groups has changed, determine whether there are unexamined allocated sequence partition slots.

[0078] Step 204: If there are unexamined allocated sequence partition slots, select an unexamined allocated sequence partition slot.

[0079] Step 205: Determine whether the sequence partition slot meets the clearing condition.

[0080] Step 206: If the sequence partition slot meets the clearing condition, clear the mapping relationship between the sequence partition slot and the data replica group in the data allocation table, and use the sequence partition slot as the to-be-allocated sequence partition slot.

[0081] Step 207: If there are no unexamined allocated sequence partition slots, determine whether there are unexamined to-be-allocated sequence partition slots.

[0082] Step 208: If there are unexamined to-be-allocated sequence partition slots, select an unexamined to-be-allocated sequence partition slot, randomly allocate a data replica group with fewer held sequence partition slots, and add the mapping relationship between the sequence partition slot and the data replica group to the data allocation table.

[0083] Specifically, in combination with Figure 2 , the user can configure the number of cluster sequence partition slots before the first startup of the cluster n , and the cluster initializes n sequence partition slots at the first startup and marks all of them as to-be-allocated. After that, the number of sequence partition slots cannot be changed anymore. For example, assume the user sets the number of sequence partition slots .

[0084] Furthermore, the current number of cluster replica groups is m. For example, assume the original cluster data replica groups , and a data allocation table has been constructed previously: , and at this time the number of cluster data replica groups grows to .

[0085] Furthermore, when it is detected that a data replica group is added or deleted in the cluster, traverse each allocated sequence partition slot in a random order. For each allocated sequence partition slot, determine whether the sequence partition slot meets the clearing condition. If the sequence partition slot meets the clearing condition, clear the mapping relationship between the sequence partition slot and the data replica group in the data allocation table, and use the sequence partition slot as the to-be-allocated sequence partition slot; where the clearing condition is that the data replica group corresponding to the sequence partition slot is deleted or the number of sequence partition slots held by the data replica group corresponding to the sequence partition slot exceeds the dynamic threshold.

[0086] For example, traverse each allocated sequence partition slot in a random order. Currently, the traversed sequence partition slot is , if is deleted, or meets , then let . Among them, refers to the record corresponding to the sequence partition slot in the data allocation table, represents the number of sequence partition slots allocated to the data replica group in the current data allocation table .

[0087] For example, traverse the sequence partition slots in sequence . When traversing to , if , then let , that is, clear the allocation record of and mark it as to be allocated. When traversing to , the processing flow is the same as . For the remaining sequence partition slots, since , no processing is required.

[0088] Furthermore, traverse each to-be-allocated sequence partition slot in a random order. For each to-be-allocated sequence partition slot, randomly allocate a data replica group that holds fewer sequence partition slots, and add the mapping relationship between the sequence partition slot and the data replica group in the data allocation table.

[0089] Exemplarily, traverse the to-be-allocated sequence partition slots in a random order. Suppose the current traversed one is . Randomly allocate to a that meets the requirement, and let . In this example, . Since in this step, let .

[0090] It can be understood that after the above reallocation process ends, , it is ensured to generate a uniform data allocation table. In this example, after the above steps, there is , and a uniform data allocation table can be obtained.

[0091] Optionally, when the number of data replica groups in the cluster changes, immediately update the balanced data allocation table so that each data replica group holds an equal number of sequence partition slots. The data partition table starts to be updated from the next time partition slot after the data allocation table is updated. The data replica group to which each data partition belongs is the same as the record in the data allocation table, thus ensuring that the data partitions corresponding to all future time partition slots are evenly allocated.

[0092] In this embodiment, when the number of data replica groups in the time series database cluster changes, by redistributing data to adapt to the new number of replica groups, the read and write performance and high availability of time series data are improved. Further, updating the data distribution table can ensure that data is evenly distributed in the new replica group configuration, avoid the situation where some replica groups are overloaded while other replica groups have idle resources, achieve load balancing, and improve the reliability of the time series database system.

[0093] In practical applications, it is necessary to maintain the data partition table. As an example, in one possible embodiment, Figure 3 is the second schematic flow diagram of the partition allocation method of the time series database provided by the present invention. As shown in Figure 3 above, the above step 106 includes: step 301 and step 302.

[0094] Step 301: Write the data partition into the data replica group corresponding to the sequence partition slot.

[0095] Step 302: Query the data partition table. If the data partition does not have a corresponding data replica group, perform the allocation process, allocate a data replica group for the data partition, and write the mapping relationship between the data partition and the data replica group into the data partition table.

[0096] Among them, the allocation process includes: querying the data replica group corresponding to the sequence partition slot where the data partition is located in the data distribution table, allocating the data partition to the data replica group, and recording the allocation relationship in the data partition table.

[0097] Exemplarily, if , it means that the data partition has been allocated to a certain data replica group, and return as the write location of the data partition .

[0098] Further, if , then let . It can be understood that under the guidance of the data distribution table, the generated data partition table will, while ensuring the load balance of the cluster, as much as possible ensure that the data partitions under the same sequence partition slot belong to the same data replica group.

[0099] It can be understood that the above process ensures that when the time partition slot satisfies and there is , assuming that the number of cluster data replica groups corresponding to the time period of the time partition slot is , then , that is, the data partition table under any time partition slot is evenly distributed.

[0100] In this embodiment, the corresponding relationship between data partitions and data replica groups is confirmed by querying the data partition table to ensure the correctness and consistency of data writing. Further, for new data partitions or unallocated data partitions, data replica groups are intelligently allocated to maintain the balanced distribution of data in the cluster. Further, by preferentially allocating data partitions to the data replica groups corresponding to the previous time partition slots of the same sequence partition slot, the proximity of related data in physical locations is maintained, improving the query efficiency. Further, by dynamically allocating data partitions to data replica groups, storage and computing resources are utilized more effectively, improving the resource utilization rate. Therefore, this embodiment not only improves the performance and efficiency of data writing, but also enhances the flexibility of data management and the stability of the system, and also brings convenience to the operation and maintenance management of the time-series database.

[0101] In the partition allocation method of the time-series database provided by the present invention, by introducing the concepts of sequence partition slots and time partition slots, a large amount of time-series data is evenly allocated to a limited number of sequence partition slots to reduce the volume of partition information. By pairing sequence partition slots with time partition slots to form data partitions, the storage structure of data is optimized, making the organization of data more orderly, facilitating management and access, and improving the efficiency of data access. Further, this method avoids active data migration to reduce the burden on the cluster, especially in write-intensive scenarios such as industrial Internet of Things, which helps to improve the write throughput of the cluster. Further, by querying and updating the data allocation table, the data distribution between data replica groups can be dynamically balanced to ensure cluster load balancing. In summary, the solution of the present invention improves the performance and efficiency of the time-series database in the industrial Internet of Things application scenario.

[0102] The partition allocation device of the time-series database provided by the present invention will be described below. The partition allocation device of the time-series database described below can be correspondingly referred to the partition allocation method of the time-series database described above.

[0103] Figure 4 is a schematic structural diagram of the partition allocation device of the time-series database provided by the present invention. As Figure 4 shown, the partition allocation device of the time-series database includes: a parsing module 41, a determination module 42, a pairing module 43, a query module 44, and a writing module 45.

[0104] The parsing module 41 is used to parse the received time-series data to obtain a plurality of time-series data points; wherein each time-series data point includes a timestamp and a data value.

[0105] The determination module 42 is used for each time-series data point, applying a predefined sequence partition slot algorithm to determine the sequence partition slot corresponding to the time-series data point, and determining the time partition slot corresponding to the time-series data point according to the timestamp of the time-series data point.

[0106] A pairing module 43, configured to pair the sequence partition slots corresponding to the time series data points with the time partition slots corresponding to the time series data points to obtain data partitions.

[0107] A query module 44, configured to query a data distribution table to determine the data replica group corresponding to the sequence partition slot, where the data distribution table is used to record the mapping relationship between the sequence partition slot and the data replica group.

[0108] A writing module 45, configured to write the data partition into the data replica group corresponding to the sequence partition slot and update the data partition table, where the data partition table is used to record the mapping relationship between the data partition and the data replica group.

[0109] Wherein, the time series data refers to the data stream of the time series, and these data may come from multiple real-time data sources such as multiple sensors, device logs, transaction records, etc. in the industrial Internet of Things. Further, after receiving the time series data, the parsing module 41 parses the received time series data to extract a plurality of time series data points, and each time series data point includes a timestamp and a data value corresponding thereto.

[0110] Specifically, the timestamp is the date and time mark of each data point, which indicates when the data value is observed or recorded. The data value is the actual measured value or recorded value associated with the timestamp, and it can be an index value such as temperature, pressure, speed, or any numerical value that needs to be recorded and analyzed.

[0111] Wherein, the predefined sequence partitioning algorithm determines the mapping rule between the time series data points and the sequence partition slots. The predefined sequence partitioning algorithm is an algorithm that evenly distributes the time series data into a finite number of sequence partition slots. In practice, the selection of the sequence partitioning algorithm can be based on various factors such as the data access pattern and the data distribution characteristics. The sequence partitioning algorithm includes but is not limited to hash functions, range partitioning algorithms, and list partitioning algorithms.

[0112] In practical applications, when the time series database cluster is started for the first time, the number of sequence partition slots and the number of data replicas held by each node are set. Further, the sequence partition slots are evenly distributed to the data replica groups currently held by the cluster, thereby forming a data distribution table, where the data distribution table is used to record the mapping relationship between the sequence partition slots and the data replica groups, that is, the data distribution table includes records of allocating data replica groups for each sequence partition slot.

[0113] Specifically, for each time series data point, applying the predefined sequence partition slot algorithm can determine the sequence partition slot corresponding to the time series data point. It can be understood that in this embodiment, evenly distributing a large amount of time series data to a limited number of sequence partition slots can reduce the volume of partition information.

[0114] Among them, the time partition slot is a mechanism for partitioning data based on timestamps. It divides time into multiple intervals (such as one slot per hour, per day, or per week), and groups data according to these intervals. Combining the above description, each time-series data point contains a timestamp indicating the exact time when the data was recorded.

[0115] In practical applications, a time partition slot is divided at fixed time intervals. Specifically, the determination module 42 calculates which time partition slot the data point belongs to according to a preset time interval (such as per hour or per day) and the timestamp of the data point. That is, the time partition corresponding to the time-series data point is determined according to the time partition to which the timestamp of the time-series data point belongs.

[0116] Furthermore, a sequence partition slot and a time partition slot are paired to form a data partition. In practical applications, for each time-series data point, the pairing module 43 pairs the sequence partition slot and the time partition slot determined by the determination module 42, thereby obtaining a unique data partition.

[0117] In this embodiment, the data partition is a combination of a sequence partition slot and a time partition slot, which represents the unique storage location of the time-series data point in the database. This pairing mechanism ensures that the data is organized in a way that takes into account both the source of the data (sequence) and the time attributes of the data.

[0118] It can be understood that by pairing the sequence partition slot and the time partition slot, a multi-dimensional data partition system is formed, which not only improves the performance of data operations, but also provides support for data storage optimization and life cycle management.

[0119] It can be understood that the data allocation table is used to record the mapping relationship between the sequence partition slot and the data replica group. Therefore, by querying the data allocation table, the data replica group corresponding to the sequence partition slot can be determined, that is, it is possible to quickly locate which data replica group the data of each sequence partition slot should be stored in.

[0120] Preferably, in order to further improve the management efficiency of time-series data, in one example, the data partitions with the same sequence partition slot are stored centrally. Specifically, when the number of cluster data replica groups remains unchanged, the data partitions with the same sequence partition slot will be written into the same data replica group under the guidance of the data allocation table and recorded in the data partition table. It can be understood that since the data with the same sequence partition slot is centrally stored in the same data replica group, queries can locate the relevant data faster, reducing the need to transfer data across physical nodes. Therefore, the method of writing the data partitions with the same sequence partition slot into the same data replica group not only improves the data access efficiency, but also reduces the operation and maintenance costs of the system.

[0121] In addition, in a possible implementation manner, the above partition allocation device of the time series database further includes: a regular cleaning module; the regular cleaning module is used to: Set the data survival time before the time series database cluster is started for the first time; Execute a sliding window algorithm based on the data survival time to determine the expired time partition slot; Clean the data partition table records and data partitions related to the expired time partition slot.

[0122] In this implementation manner, timely cleaning of the expired time partition slot can reduce the memory pressure, and moreover, reduce the active migration of the data partitions corresponding to the premature time partition slots. Therefore, the present invention can achieve load balancing without affecting the cluster write throughput.

[0123] In addition, in a possible implementation manner, the above partition allocation device of the time series database further includes: a pre-configuration module; the pre-configuration module is used to: Configure the number of sequence partition slots, the number of data replicas held by each node, the sequence partition slot algorithm, and the time interval between time partitions in the time series database cluster before the time series database cluster is started for the first time; Initialize each sequence partition slot after the time series database cluster is started for the first time; Whenever the number of data replica groups in the time series database cluster changes, allocate the corresponding data replica group to each sequence partition slot and generate a data allocation table.

[0124] In this embodiment, on the one hand, by configuring sequence partition slots and the number of data replicas held by each node when the cluster starts, resources can be reasonably allocated according to the expected data volume and access patterns, improving resource utilization. On the other hand, defining the sequence partition slot algorithm and initializing each sequence partition slot can ensure that data write operations are quickly and evenly distributed across the cluster, improving data write efficiency. On the other hand, configuring the time interval makes it possible to manage data in an orderly manner according to time, thereby improving data query performance. On the other hand, whenever the number of data replica groups in the time series database cluster changes, allocating corresponding data replica groups to each sequence partition slot helps balance the load of the cluster, avoiding overloading of some nodes while other nodes are idle, and achieving load balancing. On the other hand, generating a data allocation table to record the mapping relationship between sequence partition slots and data replica groups helps maintain data consistency and reliability. And the data allocation table provides a clear view of data distribution, simplifying data management and maintenance work. On the other hand, reasonable partition and replica group configuration enables the system to be more easily extended by adding nodes to cope with the growth of data volume, improving the scalability of the system. On the other hand, users can customize the sequence partition slot algorithm and time interval according to specific business requirements, making the system highly adaptable. Therefore, this embodiment provides an efficient, scalable, reliable, and easy-to-manage solution, which is particularly suitable for scenarios that need to process large-scale time series data.

[0125] In addition, in a possible implementation, the partition allocation device of the above time series database further includes: an update module; The update module is used to update the data allocation table if the number of data replica groups in the time series database cluster changes.

[0126] Further, in an example, the update module is specifically used for: Traverse each allocated sequence partition slot in a random order. For each allocated sequence partition slot, if the sequence partition slot meets the clearing condition, clear the mapping relationship between the sequence partition slot and the data replica group in the data allocation table, and use the sequence partition slot as a to-be-allocated sequence partition slot; where the clearing condition is that the data replica group corresponding to the sequence partition slot is deleted or the number of sequence partition slots held by the data replica group corresponding to the sequence partition slot exceeds the dynamic threshold; Traverse each to-be-allocated sequence partition slot in a random order. For each to-be-allocated sequence partition slot, randomly allocate a data replica group that holds fewer sequence partition slots, and add the mapping relationship between the sequence partition slot and the data replica group to the data allocation table.

[0127] In this embodiment, when the number of data replica groups in the time series database cluster changes, by reallocating data to adapt to the new number of replica groups, the read and write performance and high availability of time series data are improved. Further, updating the data allocation table can ensure that data is evenly distributed in the new replica group configuration, avoid the situation where some replica groups are overloaded while other replica groups have idle resources, achieve load balancing, and improve the reliability of the time series database system.

[0128] In addition, in a possible implementation manner, when the above-mentioned writing module 45 is used to update the data partition table, it is specifically used for: Query the data partition table. If there is no corresponding data replica group for the data partition, perform the allocation process, allocate a data replica group for the data partition, and write the mapping relationship between the data partition and the data replica group into the data partition table; where the allocation process includes: query the data replica group corresponding to the sequence partition slot where the data partition is located in the data allocation table, allocate the data partition to the data replica group, and record the allocation relationship in the data partition table.

[0129] In this embodiment, by querying the data partition table to confirm the correspondence between the data partition and the data replica group, the correctness and consistency of data writing are ensured. Further, for new data partitions or unallocated data partitions, data replica groups are intelligently allocated to maintain the balanced distribution of data in the cluster. Further, by preferentially allocating the data partition to the data replica group corresponding to the previous time partition slot of the same sequence partition slot, the proximity of related data in the physical location is maintained, improving the query efficiency. Further, by dynamically allocating data partitions to data replica groups, storage and computing resources are utilized more effectively, improving the resource utilization rate. Therefore, this embodiment not only improves the performance and efficiency of data writing, but also enhances the flexibility of data management and the stability of the system, and at the same time brings convenience to the operation and maintenance management of the time series database.

[0130] In the partition allocation device of the time series database provided by the present invention, by introducing the concepts of sequence partition slots and time partition slots, a large amount of time series data is evenly allocated to a limited number of sequence partition slots to reduce the volume of partition information. By pairing the sequence partition slots with the time partition slots to form data partitions, the storage structure of the data is optimized, making the organization of the data more orderly, facilitating management and access, and improving the efficiency of data access. Further, this method avoids active data migration to reduce the burden on the cluster, especially in write-intensive scenarios such as industrial Internet of Things, which helps to improve the write throughput of the cluster. Further, by querying and updating the data allocation table, the data distribution between data replica groups can be dynamically balanced to ensure load balancing of the cluster. In summary, the solution of the present invention improves the performance and efficiency of the time series database in the industrial Internet of Things application scenario.

[0131] Figure 5 is a schematic structural diagram of an electronic device provided by the present invention. As Figure 5 shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete mutual communication through the communication bus 540. The processor 510 may call logical instructions in the memory 530 to execute a method for partitioning and allocating a time series database. The method includes: parsing the received time series data to obtain a plurality of time series data points; where each time series data point includes a timestamp and a data value; for each time series data point, applying a predefined sequence partition slot algorithm to determine the sequence partition slot corresponding to the time series data point, and determining the time partition slot corresponding to the time series data point according to the timestamp of the time series data point; pairing the sequence partition slot corresponding to the time series data point with the time partition slot corresponding to the time series data point to obtain a data partition; querying a data allocation table to determine the data replica group corresponding to the sequence partition slot, and the data allocation table is used to record the mapping relationship between the sequence partition slot and the data replica group; writing the data partition into the data replica group corresponding to the sequence partition slot, and updating a data partition table, and the data partition table is used to record the mapping relationship between the data partition and the data replica group.

[0132] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of a software functional unit and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0133] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the partition allocation method of the time series database provided by the above-mentioned various methods. The method includes: parsing the received time series data to obtain a plurality of time series data points; wherein each time series data point includes a timestamp and a data value; for each time series data point, applying a predefined sequence partition slot algorithm to determine the sequence partition slot corresponding to the time series data point, and determining the time partition slot corresponding to the time series data point according to the timestamp of the time series data point; pairing the sequence partition slot corresponding to the time series data point with the time partition slot corresponding to the time series data point to obtain a data partition; querying a data allocation table to determine the data replica group corresponding to the sequence partition slot, where the data allocation table is used to record the mapping relationship between the sequence partition slot and the data replica group; writing the data partition into the data replica group corresponding to the sequence partition slot, and updating the data partition table, where the data partition table is used to record the mapping relationship between the data partition and the data replica group.

[0134] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the partition allocation method of the time series database provided by the above-mentioned various methods. The method includes: parsing the received time series data to obtain a plurality of time series data points; wherein each time series data point includes a timestamp and a data value; for each time series data point, applying a predefined sequence partition slot algorithm to determine the sequence partition slot corresponding to the time series data point, and determining the time partition slot corresponding to the time series data point according to the timestamp of the time series data point; pairing the sequence partition slot corresponding to the time series data point with the time partition slot corresponding to the time series data point to obtain a data partition; querying a data allocation table to determine the data replica group corresponding to the sequence partition slot, where the data allocation table is used to record the mapping relationship between the sequence partition slot and the data replica group; writing the data partition into the data replica group corresponding to the sequence partition slot, and updating the data partition table, where the data partition table is used to record the mapping relationship between the data partition and the data replica group.

[0135] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0136] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A partition allocation method for a time series database, characterized in that: The method comprises: Parsing the received time series data to obtain multiple time series data points; wherein each time series data point includes a timestamp and a data value; For each time series data point, a predefined sequence partition slot algorithm is applied to determine the sequence partition slot corresponding to the time series data point, and according to the timestamp of the time series data point, the time partition slot corresponding to the time series data point is determined; Pairing the sequence partition slot corresponding to the time series data point with the time partition slot corresponding to the time series data point to obtain a data partition; Querying a data allocation table to determine a data replica group corresponding to the sequence partition slot, wherein the data allocation table is used to record a mapping relationship between the sequence partition slot and the data replica group; The data partition is written into the data replica group corresponding to the sequence partition slot, and a data partition table is updated, wherein the data partition table is used to record a mapping relationship between the data partition and the data replica group.

2. The partition allocation method for a time series database according to claim 2, characterized in that: The method further comprises: Before the time series database cluster is started for the first time, setting the data survival time; After determining the time partition slot corresponding to the time series data point according to the timestamp of the time series data point, the method further includes: Based on the data survival time, a sliding window algorithm is executed to determine an expiration time partition slot; Clean up the data partition table records and data partitions associated with the expired time partition slot.

3. The partition allocation method for a time series database according to claim 1, characterized in that: The method further comprises: If the number of data replica groups in the time series database cluster changes, the data allocation table is updated.

4. The partition allocation method for a time series database according to claim 3, characterized in that: If the number of data replica groups in the time series database cluster changes, updating the data allocation table includes: Traversing each allocated sequence partition slot in a random order, for each allocated sequence partition slot, if the sequence partition slot meets the clearing condition, clearing the mapping relationship between the sequence partition slot and the data replica group in the data allocation table, and using the sequence partition slot as the sequence partition slot to be allocated; wherein the clearing condition is that the data replica group corresponding to the sequence partition slot is deleted or the number of sequence partition slots held by the data replica group corresponding to the sequence partition slot exceeds a dynamic threshold; Each sequence partition slot to be allocated is traversed in a random order, and for each sequence partition slot to be allocated, a data replica group holding fewer sequence partition slots is randomly allocated, and a mapping relationship between the sequence partition slot and the data replica group is added in the data allocation table.

5. The partition allocation method for a time series database according to claim 3, characterized in that: The updating data partition table includes: Query the data partition table, and if there is no corresponding data copy group for the data partition, perform allocation processing, allocate a data copy group for the data partition, and write the mapping relationship between the data partition and the data copy group into the data partition table; wherein the allocation processing includes: querying the data allocation table for the data copy group corresponding to the serial partition slot where the data partition is located, and recording the allocation relationship in the data partition table.

6. The partition allocation method for a time series database according to any one of claims 1 to 5, characterized in that: The method further comprises: Before the time series database cluster is started for the first time, the number of sequence partition slots in the time series database cluster, the number of data copies held by each node, the sequence partition slot algorithm, and the time partition interval are configured; After the time series database cluster is started for the first time, each sequence partition slot is initialized; Whenever the number of data replica groups in the time series database cluster changes, a corresponding data replica group is allocated to each sequence partition slot, and the data allocation table is generated.

7. A partition allocation device for a time series database, characterized in that: The device comprises: A parsing module, used to parse the received time series data to obtain multiple time series data points; wherein each time series data point includes a timestamp and a data value; A determination module, configured to apply a predefined sequence partition slot algorithm to each time series data point to determine the sequence partition slot corresponding to the time series data point, and determine the time partition slot corresponding to the time series data point according to the timestamp of the time series data point; A pairing module, used to pair the sequence partition slot corresponding to the time series data point with the time partition slot corresponding to the time series data point to obtain a data partition; A query module, used to query a data allocation table to determine the data replica group corresponding to the sequence partition slot, wherein the data allocation table is used to record the mapping relationship between the sequence partition slot and the data replica group; The writing module is used to write the data partition into the data replica group corresponding to the sequence partition slot and update the data partition table, wherein the data partition table is used to record the mapping relationship between the data partition and the data replica group.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the partition allocation method for the time series database according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the partition allocation method for a time series database as claimed in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the partition allocation method for a time series database as claimed in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Method and system for realizing partition table of time sequence database in massive equipment scene

    CN117235183A

  • Data processing method and device, medium and electronic equipment

    CN118708609A