Data storage, query, management method, device, equipment, system and program
By dividing time partitions and setting routing information in a distributed storage system, the data migration problem when adding storage nodes is solved, and efficient data storage and query are achieved.
Patent Information
- Application Number
- CN202110729328.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-29
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-06-29
AI Technical Summary
When adding distributed storage nodes, the prior art requires large-scale migration of existing data, resulting in frequent migration of data between storage nodes.
By dividing multiple time partitions according to time granularity in a distributed storage system and setting routing information for each partition, the timing data is stored in multiple storage nodes to ensure that the new storage node does not change the storage location of the existing data.
It realizes that no need to migrate existing data when adding storage nodes, improves the efficiency and stability of data storage, and reduces data migration operations.
Smart Images

Figure CN113590675B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of distributed storage technology, and in particular, to a method, device, equipment, system and program for data storage, query and management. Background Art
[0002] Time series data is a series of metric data generated continuously based on a certain frequency. For example, when monitoring the air quality of a city, a series of data is generated by collecting the value of sulfur dioxide concentration every second.
[0003] There is a large amount of time series data in the Internet of Things and the Industrial Internet, and time series data can be stored using a distributed storage system. Usually, the proxy nodes in the distributed storage system use the same partitioning rule to partition the data that has been stored in the distributed storage system and the data that needs to be stored in the distributed storage system in the future, so as to store the data dispersedly on multiple storage nodes in the distributed storage system.
[0004] However, in this way, when adding storage nodes, it is necessary to recalculate the storage nodes for the existing data, and a large amount of data needs to be migrated between the storage nodes. Summary of the Invention
[0005] Embodiments of this application provide a method, device, equipment, system and program for data storage, query and management, so as to solve the problem that a large amount of data needs to be migrated between storage nodes when adding storage nodes in the prior art.
[0006] In a first aspect, an embodiment of this application provides a data storage method, including:
[0007] Obtain time series data to be stored;
[0008] Determine a target time partition corresponding to the time series data in the distributed storage system according to the time information corresponding to the time series data; there are multiple time partitions divided according to time granularity in the distributed storage system, and there is corresponding routing information for the time partition;
[0009] Store the time series data into at least one of the multiple storage nodes in the distributed storage system according to the target routing information corresponding to the target time partition.
[0010] In a second aspect, an embodiment of this application provides a data query method, including:
[0011] Obtain a query request for requesting to query time series data;
[0012] Determine the target time partition corresponding to the time series data requested to be queried in the distributed storage system according to the time information corresponding to the query request; there are multiple time partitions divided according to time granularity in the distributed storage system, and there is corresponding routing information for the time partitions;
[0013] Find the time series data from at least one storage node among multiple storage nodes in the distributed storage system according to the target routing information corresponding to the target time partition.
[0014] In a third aspect, an embodiment of the present application provides a data management method, including:
[0015] Divide multiple time partitions according to time granularity;
[0016] When reaching the target moment within the time partition, add corresponding routing information according to the current storage node situation of the distributed storage system, so as to be able to use the routing information corresponding to the time partition to disperse and store the time series data corresponding to the time partition into multiple storage nodes in the distributed storage system.
[0017] In a fourth aspect, an embodiment of the present application provides a data storage device, including:
[0018] An acquisition module, configured to acquire time series data to be stored;
[0019] A determination module, configured to determine the target time partition corresponding to the time series data in the distributed storage system according to the time information corresponding to the time series data; there are multiple time partitions divided according to time granularity in the distributed storage system, and there is corresponding routing information for the time partitions;
[0020] A storage module, configured to store the time series data into at least one storage node among multiple storage nodes in the distributed storage system according to the target routing information corresponding to the target time partition.
[0021] In a fifth aspect, an embodiment of the present application provides a data query device, including:
[0022] An acquisition module, configured to acquire a query request for requesting to query time series data;
[0023] A determination module, configured to determine the target time partition corresponding to the time series data requested to be queried in the distributed storage system according to the time information corresponding to the query request; there are multiple time partitions divided according to time granularity in the distributed storage system, and there is corresponding routing information for the time partitions;
[0024] A search module, configured to search for the time series data from at least one storage node among multiple storage nodes of the distributed storage system according to the target routing information corresponding to the target time partition.
[0025] In a sixth aspect, an embodiment of the present application provides a data management device, including:
[0026] A partitioning module, configured to partition multiple time partitions according to a time granularity;
[0027] An adding module, configured to, when reaching a target moment within the time partition, add corresponding routing information according to the current storage node situation of the distributed storage system, so as to be able to use the routing information corresponding to the time partition to disperse and store the time series data corresponding to the time partition into multiple storage nodes of the distributed storage system.
[0028] In a seventh aspect, an embodiment of the present application provides a computer device, including: a memory and a processor; wherein, the memory is used to store one or more computer instructions, and when the one or more computer instructions are executed by the processor, the method described in any one of the first aspects is implemented.
[0029] In an eighth aspect, an embodiment of the present application provides a computer device, including: a memory and a processor; wherein, the memory is used to store one or more computer instructions, and when the one or more computer instructions are executed by the processor, the method described in any one of the second aspects is implemented.
[0030] In a ninth aspect, an embodiment of the present application provides a computer device, including: a memory and a processor; wherein, the memory is used to store one or more computer instructions, and when the one or more computer instructions are executed by the processor, the method described in any one of the third aspects is implemented.
[0031] In a tenth aspect, an embodiment of the present application provides a distributed storage system, including: a proxy node and multiple storage nodes; the proxy node is configured to execute the method described in any one of the first aspects.
[0032] In an eleventh aspect, an embodiment of the present application provides a distributed storage system, including: a proxy node and multiple storage nodes; the proxy node is configured to execute the method described in any one of the second aspects.
[0033] In a thirteenth aspect, an embodiment of the present application provides a distributed storage system, including: a proxy node and multiple storage nodes; the proxy node is configured to execute the method described in any one of the third aspects.
[0034] An embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes at least one piece of code. The at least one piece of code can be executed by a computer to control the computer to execute the method described in any one of the first aspect.
[0035] An embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes at least one piece of code. The at least one piece of code can be executed by a computer to control the computer to execute the method described in any one of the second aspect.
[0036] An embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes at least one piece of code. The at least one piece of code can be executed by a computer to control the computer to execute the method described in any one of the third aspect.
[0037] An embodiment of the present application further provides a computer program. When the computer program is executed by a computer, it is used to implement the method described in any one of the first aspect.
[0038] An embodiment of the present application further provides a computer program. When the computer program is executed by a computer, it is used to implement the method described in any one of the second aspect.
[0039] An embodiment of the present application further provides a computer program. When the computer program is executed by a computer, it is used to implement the method described in any one of the third aspect.
[0040] In an embodiment of the present application, in combination with the time characteristics of time series data, a plurality of time partitions are divided in the time dimension, and there is corresponding routing information for the time partitions; in the node dimension, according to the routing information corresponding to the time partitions, the time series data within the same time partition is scattered and stored on multiple storage nodes in a distributed storage system for storage, realizing a multi-dimensional partitioning method of partitioning the time series data in the node dimension on the basis of partitioning in the time dimension. When adding a new storage node to the system using this data storage method, it will not cause a change in the storage location of the existing stock data. Therefore, when adding a new storage node, there is no need to migrate the stock data between the storage nodes. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0042] Figure 1 It is a schematic structural diagram of the distributed storage system according to an embodiment of the present application;
[0043] Figure 2A It is a schematic diagram of dispersedly storing time-series data into multiple storage nodes in the prior art;
[0044] Figure 2B It is for Figure 2A a schematic diagram of migrating data between storage nodes when adding storage nodes on the basis of
[0045] Figure 3 It is a schematic diagram of dispersedly storing time-series data into multiple storage nodes according to an embodiment of the present application;
[0046] Figure 4 It is a schematic flowchart of a data storage method provided by an embodiment of the present application;
[0047] Figure 5 It is a schematic flowchart of a data query method provided by an embodiment of the present application;
[0048] Figure 6 It is a schematic flowchart of a data management method provided by an embodiment of the present application;
[0049] Figure 7 It is a schematic structural diagram of a data storage device provided by an embodiment of the present application;
[0050] Figure 8 It is a schematic structural diagram of a computer device provided by an embodiment of the present application;
[0051] Figure 9 It is a schematic structural diagram of a data query device provided by an embodiment of the present application;
[0052] Figure 10 It is a schematic structural diagram of a computer device provided by another embodiment of the present application;
[0053] Figure 11 It is a schematic structural diagram of a data management device provided by an embodiment of the present application;
[0054] Figure 12 It is a schematic structural diagram of a computer device provided by yet another embodiment of the present application. Detailed implementation manners
[0055] To make the objectives, technical solutions and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are part of rather than all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0056] The terms used in the embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms "a", "the" and "said" used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. "Plural" generally includes at least two, but does not exclude the case of including at least one.
[0057] It should be understood that the term "and / or" used herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0058] Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "when...", "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" may be interpreted as "when determined", "in response to determining", "when detecting (stated condition or event)", or "in response to detecting (stated condition or event)".
[0059] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a commodity or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such commodity or system. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of another identical element in the commodity or system including the said element.
[0060] In addition, the time sequence of the steps in the following method embodiments is only an example and not a strict limitation.
[0061] To facilitate the understanding of the technical solutions provided by the embodiments of this application by those skilled in the art, the technical environment for implementing the technical solutions will be described below first.
[0062] In the related art, the commonly used method for storing time-series data in a distributed storage system mainly includes: the proxy nodes in the distributed storage system use the same partitioning rule to partition the data already stored in the distributed storage system and the data that needs to be stored in the distributed storage system in the future. However, when adding storage nodes, it is necessary to recalculate the storage nodes for the existing data, and a large amount of data needs to be migrated between the storage nodes. Therefore, there is an urgent need in the related art for a data storage method that can avoid data migration between storage nodes when adding storage nodes.
[0063] Based on the actual technical requirements similar to those described above, the data storage method provided in this application can use technical means to avoid data migration between storage nodes when adding storage nodes.
[0064] The following specifically describes the data storage method provided by each embodiment of this application through an exemplary application scenario.
[0065] Figure 1 It is a schematic structural diagram of a distributed storage system provided by an embodiment of this application. As Figure 1 shown, the distributed storage system may include a proxy node 11 and multiple storage nodes 12. Among them, the storage node 12 can be used to provide actual data access and storage functions, and the proxy node 11 can act as an agent for the outside world (such as a user) to access the storage node 12, and the proxy node 11 can execute the method provided by the embodiment of this application. Specifically, the proxy node 11 can execute the data storage method provided by the embodiment of this application to disperse and store time-series data in multiple storage nodes 12 of the distributed storage system. The proxy node 11 can also execute the data query method provided by the embodiment of this application to disperse and search for time-series data from multiple storage nodes 12 of the distributed storage system.
[0066] The method provided by the embodiment of this application can be applied to any type of scenario that uses a distributed storage system to store time-series data. For example, it can be applied to scenarios that use a distributed storage system to store time-series data in a power grid, a vehicle network, or an Internet of Things, etc.
[0067] It should be noted that Figure 1 in this example, the number of proxy nodes 11 is taken as one. In other embodiments, the number of proxy nodes 11 can also be multiple, and multiple proxy nodes 11 all execute the method provided by the embodiment of this application.
[0068] It should be noted that Figure 1 in this example, the proxy node 11 and the storage node 12 can be located in the same computer device, or they can also be located in different computer devices.
[0069] Generally, the proxy node 11 uses the same partitioning rule to partition the data already stored in the distributed storage system and the data that needs to be stored in the distributed storage system in the future, so as to disperse and store the time-series data in multiple storage nodes 12 of the distributed storage system. As Figure 2A shown, assuming that the number of storage nodes in the distributed storage system is 3, namely storage node 1, storage node 2, and storage node 3, the proxy node 11 can use a partitioning rule to disperse and store all the time-series data to be stored (i.e., the full amount of data) in storage nodes 1 to 3. The partitioning rule describes that the time-series data meeting condition 1 needs to be stored in storage node 1, the time-series data meeting condition 2 needs to be stored in storage node 2, and the time-series data meeting condition 3 needs to be stored in storage node 3. Among them, the time-series data meeting condition 1, the time-series data meeting condition 2, and the time-series data meeting condition 3 constitute the full amount of data.
[0070] On the Figure 2A basis of the distributed storage system shown, if a new storage node 4 is added, the structure of the distributed storage system can be as Figure 2B shown. Since the number of storage nodes has changed, the partitioning rule also needs to be updated. The new partitioning rule describes that the time-series data meeting condition 1' needs to be stored in storage node 1, the time-series data meeting condition 2' needs to be stored in storage node 2, the time-series data meeting condition 3' needs to be stored in storage node 3, and the time-series data meeting condition 4' needs to be stored in storage node 4. Among them, the time-series data meeting condition 1', the time-series data meeting condition 2', the time-series data meeting condition 3', and the time-series data meeting condition 4' constitute the full amount of data.
[0071] Since Figure 2A the time-series data meeting condition 1, the time-series data meeting condition 2, and the time-series data meeting condition 3 in Figure 2B constitute the full amount of data, and the time-series data meeting condition 1', the time-series data meeting condition 2', the time-series data meeting condition 3', and the time-series data meeting condition 4' in
[0072] constitute the full amount of data, and the number of conditions has changed, so at least two of conditions 1', 2', 3', and 4' will change compared with conditions 1, 2, and 3. Figure 2BAs shown, a part of the stored data in storage node 1 needs to be migrated to storage node 2, another part of the stored data needs to be migrated to storage node 3, and yet another part of the stored data needs to be migrated to storage node 4; a part of the stored data in storage node 2 needs to be migrated to storage node 1, another part of the stored data needs to be migrated to storage node 3, and yet another part of the stored data needs to be migrated to storage node 4; a part of the stored data in storage node 3 needs to be migrated to storage node 1, another part of the stored data needs to be migrated to storage node 2, and yet another part of the stored data needs to be migrated to storage node 4. It can be seen that when adding storage nodes, a large amount of data needs to be migrated between storage nodes.
[0073] It should be noted that Figure 2B the arrow direction in indicates the migration direction of the stored data.
[0074] To solve the technical problem that a large amount of data needs to be migrated between storage nodes when adding storage nodes, in the Figure 1 shown distributed storage system, the proxy node 11 combines the time characteristics of the time series data, divides multiple time partitions in the time dimension according to the time granularity, and there is corresponding routing information for the time partitions. Based on this, the time series data is first divided into different time partitions, and then in the node dimension, according to the routing information corresponding to the time partition, the time series data in the same time partition is scattered and stored on multiple storage nodes in the distributed storage system, realizing that the time series data in any time partition can be accessed using the routing information corresponding to the time partition, so that even if a new storage node is added, the stored data existing before the new storage node can still be accessed using the routing information corresponding to the corresponding time range, that is, the routing information used to store the stored data remains unchanged, which also means that the storage location of the stored data remains unchanged. Therefore, when adding a new storage node in the distributed storage system, it will not cause a change in the storage location of the existing stored data, and thus when adding a new storage node, there is no need to migrate the stored data between storage nodes.
[0075] Based on the above, when it is necessary to store time series data in the Figure 1 shown distributed storage system, the proxy node 1 can determine the target time partition corresponding to the time series data in the distributed system according to the time information corresponding to the time series data to be stored. There are multiple time partitions divided according to the time granularity in the distributed storage system and there is corresponding routing information for the time partitions, and use the target routing information corresponding to the target time partition to store the time series data in at least one of the multiple storage nodes in the distributed storage system.
[0076] It should be noted that the time ranges of different time partitions do not overlap. Since time series data all have corresponding time information, multiple time partitions can be divided according to the time granularity, so as to partition the time series data in the time dimension. For example, as Figure 3 shown, there can be multiple time partitions 1, time partition 2, time partition 3, time partition 4, time partition 5, time partition 6,... divided according to the time granularity in the distributed storage system.
[0077] Among them, the routing information corresponding to a time partition is used to partition the time series data corresponding to this time partition in the node dimension. Specifically, Figure 3 in, the routing information corresponding to time partition 1 can be used to partition the time series data corresponding to time partition 1 in the node dimension; the routing information corresponding to time partition 2 can be used to partition the time series data corresponding to time partition 2 in the node dimension;.... Among them, different time partitions can correspond to different routing information or the same routing information. In the case where different time partitions correspond to the same routing information, it means that the partitioning method of the time series data of these several time partitions in the node dimension is the same.
[0078] For example, Figure 3 shown, assume that time partition 1, time partition 2, and time partition 3 all correspond to routing information 1, and the storage nodes corresponding to time partition 1 to time partition 3 are three storage nodes, namely storage node 12A, storage node 12B, and storage node 12C. Then the routing information 1 can be used to partition the time series data corresponding to time partition 1 to time partition 3 among storage node 12A to storage node 12C. Through the data storage method provided by the embodiments of the present application, the routing information 1 can be used to disperse and store the time series data to be stored corresponding to time partition 1, time partition 2, and time partition 3 in the full amount of data among storage node 12A to storage node 12C. Correspondingly, the routing information 1 can also be used to disperse and search for the time series data corresponding to time partition 1, time partition 2, and time partition 3 among storage node 12A to storage node 12C.
[0079] Further assume that time partitions 4, 5, and 6 all correspond to routing information 2, and the storage nodes corresponding to time partitions 4 to 6 are three storage nodes, namely storage node 12A, storage node 12B, storage node 12C, and storage node 12D. Then, routing information 2 can be used to partition the time-series data corresponding to time partitions 4 to 6 among storage nodes 12A to 12D. Through the data storage method provided by the embodiments of the present application, routing information 2 can be used to disperse and store the time-series data to be stored in time partitions 4, 5, and 6 in the full amount of data among storage nodes 12A to 12D. Correspondingly, routing information 2 can also be used to disperse and search for the time-series data corresponding to time partitions 4, 5, and 6 among storage nodes 12A to 12D.
[0080] It can be understood that when adding storage node 12D, the time-series data corresponding to time partitions 1, 2, and 3 can still be dispersed and searched among storage nodes 12A to 12C using routing information 1, so that adding a new storage node does not cause a change in the storage location of the existing data.
[0081] The following will describe some embodiments of the present application in detail with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0082] Figure 4 The figure is a schematic flowchart of a data storage method provided by an embodiment of the present application. As Figure 4 shown, the method of this embodiment may include:
[0083] Step 41: Obtain the time-series data to be stored;
[0084] Step 42: Determine the target time partition corresponding to the time-series data in the distributed storage system according to the time information corresponding to the time-series data; there are multiple time partitions divided according to time granularity in the distributed storage system, and the time partitions have corresponding routing information;
[0085] Step 43: Store the time-series data into at least one of the multiple storage nodes in the distributed storage system according to the target routing information corresponding to the target time partition.
[0086] In the embodiments of the present application, the time information corresponding to the time series data may be the time when the reporting device (such as an Internet of Things device, a vehicle-to-everything device, a power grid device, etc.) collects the time series data. Exemplarily, the time information corresponding to the time series data may include the timestamp of collecting the time series data. Alternatively, optionally, the time information corresponding to the time series data may be the time when the receiving device (such as a device for receiving data reported by an Internet of Things device, a vehicle-to-everything device, or a power grid device, etc.) receives the time series data. Exemplarily, the time information corresponding to the time series data may include the timestamp of receiving the time series data. The time series data may be as shown in Table 1 below.
[0087] Table 1
[0088]
[0089] Among them, Metric is used to indicate the index of the monitoring data. In Table 1, temperature is taken as an example, and in other embodiments, it may also be wind force, etc. Tag is used to indicate the specific object targeted by the index item monitoring. A tag consists of a tag key (TagKey) and a corresponding tag value (TagValue). For example, "floor = 33" is a tag. In Table 1, the tag keys are floor, room, and device ID as examples, and in other embodiments, it may also be country, province, city, etc. The metric value (Value) refers to the value corresponding to the metric. In Table 1, the metric values are 26°C and 25.8°C as examples. The timestamp 202106031412 can represent 14:12 on June 3, 2021, and the timestamp 202106031413 can represent 14:13 on June 3, 2021. In Table 1, the timestamp corresponding to the metric value 26°C is 14:12 on June 3, 2021, and the timestamp corresponding to the metric value 25.8°C is 14:13 on June 3, 2021 as an example.
[0090] It should be noted that in Table 1, it is taken as an example that the time series data to be stored includes 2 data points. It can be understood that in other embodiments, the time series data may also include other numbers of data points. Among them, each metric value collected at a specific time interval (continuous timestamps) for a certain index (defined by metric and tag) of the monitoring object is a data point. "One metric + N tags (N >= 1) + one timestamp + one value" can define a data point.
[0091] It should be noted that in Table 1, it is taken as an example that the time information corresponding to the time series data to be stored is accurate to "minutes". It can be understood that in other embodiments, the time information corresponding to the time series data may also be accurate to "hours", "seconds", "milliseconds", etc.
[0092] In the distributed storage system provided by the embodiments of the present application, there are multiple time partitions divided according to time granularity. Among them, a time partition may have a corresponding start time and end time to define the time range of the time partition. By dividing multiple time partitions according to time granularity, time-series data can be partitioned in the time dimension. Among them, the time granularity can be flexibly implemented. The smaller the time granularity, the more routing information needs to be stored in the distributed storage system. In one embodiment, the time granularity can be one day.
[0093] Among them, the routing information corresponding to the time partition is used to partition the time-series data corresponding to the time partition in the node dimension.
[0094] For example, taking the time granularity of one day as an example, assume Figure 3 the start time of time partition 3 can be, for example, 00:00:01 on June 3, 2021, and the end time can be, for example, 24:00:00 on June 3, 2021. And if the time-series data to be stored obtained in step 41 is as shown in Table 1, it can be determined that the time-series data corresponds to time partition 3 in the distributed storage system, and time partition 3 is the target time partition corresponding to the time-series data. Among them, the routing information 1 corresponding to time partition 3 is used to partition the time-series data corresponding to time partition 3 in the node dimension.
[0095] It should be noted that Figure 3 in, the rectangular frames filled with different filling patterns can represent the time-series data corresponding to different time partitions, and the rectangular frames filled with the same filling pattern in different storage nodes can represent that the time-series data corresponding to the same time partition is scattered and stored in multiple storage nodes. Time partitions 1 to 6 can be 6 partitions with continuous time and in chronological order from early to late, Figure 3 specifically, it can be a scenario where storage node 12D is newly added on the basis of existing storage nodes 12A to 12C in the distributed storage system.
[0096] Optionally, the method provided by the embodiments of the present application may further include: instructing the storage nodes in the distributed storage system of the time granularity, so that the storage nodes perform data access according to the time granularity. When the storage nodes perform data access according to the time granularity, Figure 3 one rectangular frame in storage nodes 12A to 12D can represent a data shard. Among them, a data shard refers to a logical unit for data differentiation inside a storage node and can be considered a kind of indexing method. By instructing the storage nodes of the time granularity so that the storage nodes perform data access according to the time granularity, the storage nodes can quickly locate the data partitions that need to be scanned according to the time partitions when performing data queries, which is beneficial to improving the scanning efficiency.
[0097] Optionally, the routing information corresponding to the time partition in the distributed storage system may be manually added by an operation and maintenance personnel. For example, at the start time of the time partition, the operation and maintenance personnel may add the corresponding routing information according to the current storage node situation of the distributed storage system. Alternatively, optionally, the routing information corresponding to the time partition in the distributed storage system may be automatically added by a proxy node to reduce manual operations. The following mainly takes the automatic addition by the proxy node as an example for specific description.
[0098] In one embodiment, the method provided by the embodiments of the present application may further include: at the target time within the time partition, according to the current storage node situation of the distributed storage system, adding corresponding routing information. The current storage node situation may be used to indicate the storage nodes currently available for storing data. The current storage node situation may include, for example, the IP addresses of the storage nodes currently available for storing data.
[0099] For example, assume that at Figure 3 the target time of the time partition 3 shown in Figure 3 the storage nodes 12A to 12C can all be used for storing data. Then, at the target time of the time partition 3, the current storage node situation of the distributed storage system may include, for example, the IP address of the storage node 12A, the IP address of the storage node 12B, and the IP address of the storage node 12C. Thus, the routing information newly added according to the current storage node request of the distributed storage system can be used to partition the time series data corresponding to the time partition 3 among the storage nodes 12A to 12C.
[0100] For another example, assume that at Figure 3 the target time of the time partition 4 shown in Figure 3 the storage nodes 12A to 12D can all be used for storing data. Then, at the target time of the time partition 4, the current storage node situation of the distributed storage system may include, for example, the IP address of the storage node 12A, the IP address of the storage node 12B, the IP address of the storage node 12C, and the IP address of the storage node 12D. Thus, the routing information newly added according to the current storage node request of the distributed storage system can be used to partition the time series data corresponding to the time partition 4 among the storage nodes 12A to 12D.
[0101] The target time of the time partition refers to the time when the routing information corresponding to the newly added time partition is added, and the target time can be flexibly implemented according to requirements. Optionally, if it is necessary to add the routing information corresponding to the time partition at the start time of the time partition, the target time may include the start time of the time partition. That is, when reaching the start time of the time partition, the routing information corresponding to the time partition can be added according to the current storage node situation of the distributed storage system.
[0102] Alternatively, optionally, if it is necessary to add the routing information corresponding to the time partition when the routing information corresponding to the time partition is first required, the target time may include the time when the time series data corresponding to the time within the time partition is first obtained. That is, when the time series data corresponding to the time within the time partition is first obtained, the routing information corresponding to the time partition can be added according to the storage node situation before the dynamic change of the distributed storage system.
[0103] Optionally, the routing information added when reaching the target time may remain unchanged. In this case, if new storage nodes are added to the distributed storage system after the target time and before the next time partition of the time partition, the storage node situation after the new storage nodes are added to the distributed storage system can be reflected in the routing information corresponding to the next time partition of the time partition, that is, the new nodes can take effect in the routing information corresponding to the next time partition. For example, assume that Figure 3 after the target time of the middle time partition 3 and before the time partition 4, storage node 12D is added to the distributed storage system, then the storage node situation after adding the new storage node 12D can be reflected in the routing information corresponding to the time partition 4, that is, the newly added storage node 12D can take effect in the routing information corresponding to the time partition 4.
[0104] Alternatively, optionally, the routing information added when reaching the target time can be modified into multiple routing information corresponding to multiple time sub - partitions according to the need of adding new storage nodes. In one embodiment, after adding the corresponding routing information according to the current storage node situation of the distributed storage system when reaching the target time within the time partition, it may further include: modifying the added routing information into a first routing information corresponding to the first time sub - partition of the time partition and a second routing information corresponding to the second time sub - partition of the time partition.
[0105] Among them, the storage node situations corresponding to the first time sub - partition and the second time sub - partition are different. In one embodiment, the storage node situation corresponding to one of the first time sub - partition and the second time sub - partition is the storage node situation before the new storage nodes are added to the distributed storage system, and the storage node situation corresponding to the other one is the storage node situation after the new storage nodes are added to the distributed storage system. By modifying the newly added routing information into first routing information and second routing information, the newly added storage nodes can take effect earlier.
[0106] In one embodiment, the node addition moment corresponding to the newly added storage node can be used as the boundary between the first time sub - partition and the second time sub - partition. Based on this, the modification of the newly added routing information into the first routing information corresponding to the first time sub - partition of the time partition and the second routing information corresponding to the second time sub - partition of the time partition may specifically include: according to the node addition moment, modifying the newly added routing information into the first routing information corresponding to the start moment of the time partition to the moment before the node addition moment, and the second routing information corresponding to the moment from the node addition moment to the end moment of the time partition.
[0107] For example, assume Figure 3 the start time of the middle time partition 3 is 00:00:01 on June 3, 2021, the end time is 24:00:00 on June 3, 2021, and the node addition moment of the newly added storage node 12D is 10:20:10 on June 3, 2021. Then the routing information 1 corresponding to the time partition 3 can be modified into the routing information 1 corresponding to 00:00:01 on June 3, 2021 to 10:20:09 on June 3, 2021, and the routing information B1 (should be routing information 2 in the original text, seems to be a typo) corresponding to 10:20:10 on June 3, 2021 to 24:00:00 on June 3, 2021.
[0108] Alternatively, optionally, the node effective moment corresponding to the newly added storage node can be used as the boundary between the first time sub - partition and the second time sub - partition, where the node effective moment is a moment after the node addition moment. Based on this, the modification of the newly added routing information into the first routing information corresponding to the first time sub - partition of the time partition and the second routing information corresponding to the second time sub - partition of the time partition may specifically include: determining the node effective moment; and according to the node effective moment, modifying the newly added routing information into the first routing information corresponding to the start moment of the time partition to the moment before the node effective moment, and the second routing information corresponding to the moment from the node effective moment to the end moment of the time partition.
[0109] For example, assume Figure 3The start time of time partition 3 is 00:00:01 on June 3, 2021, and the end time is 24:00:00 on June 3, 2021. Given that the node addition time of the newly added storage node 12D is 10:20:10 on June 3, 2021, and the node effective time is 12:00:00 on June 3, 2021, the routing information 1 corresponding to time partition 3 can be modified to the routing information 1 corresponding to the period from 00:00:01 on June 3, 2021 to 11:59:59 on June 3, 2021, and the routing information 2 corresponding to the period from 12:00:00 on June 3, 2021 to 24:00:00 on June 3, 2021.
[0110] In one embodiment, the node effective time can be set by the operation and maintenance personnel. Based on this, the determination of the node effective time may specifically include: obtaining the node effective time set by the operation and maintenance personnel. That is, the node effective time can be set by the operation and maintenance personnel, which is beneficial to improving flexibility.
[0111] In another embodiment, the node effective time can be automatically determined by the proxy node. Based on this, the determination of the node effective time may specifically include: taking the sum of the node addition time of the storage node and the target duration as the node effective time. That is, the node effective time can be automatically determined by the proxy node, which is beneficial to simplifying the operation. Herein, the target duration can be a fixed value or a variable value.
[0112] In the embodiments of the present application, there is corresponding routing information (i.e., target routing information) for the target time partition, and the target routing information is used to partition the time series data of the target time partition at the node level. After determining the target time partition corresponding to the time series data, the time series data can be stored in at least one of the multiple storage nodes of the distributed storage system according to the target routing information corresponding to the target time partition. Among them, the partitioning method corresponding to the target routing information can be flexibly implemented. Exemplarily, the partitioning method may include a hash partitioning method or a range (Range) method. It should be noted that the partitioning methods corresponding to different routing information in the distributed storage system may be the same or different.
[0113] When the partitioning method corresponding to the target routing information is the hash partitioning method, the target routing information can specifically be used to disperse and store the time series data of the target time partition into the routing information of multiple storage nodes in the distributed storage system by using the hash partitioning method.
[0114] Example 1, assume the target routing information is Figure 3For the routing information 1 in , when using the target routing information to store the time-series data into at least one of multiple storage nodes in the distributed storage system, for example, it may include: calculating the hash value of the first label of the time-series data to be stored (for example, it may be "floor = 33" + "room = 3302" + "device ID = 123" in Table 1), and taking the remainder of the hash value of the first label modulo 3. If the remainder is 0, the time-series data can be stored in storage node 12A; if the remainder is 1, the time-series data can be stored in storage node 12B; if the remainder is 2, the time-series data can be stored in storage node 12C. It should be noted that the number of the first labels can be one or more. When there are multiple first labels, for each first label, according to the result of taking the remainder of the hash value of the first label modulo 3, the partial time-series data corresponding to the first label in the time-series data to be stored can be stored in the corresponding storage node.
[0115] Example 2, assume the target routing information is Figure 3 For the routing information 2 in , when using the target routing information to store the time-series data into at least one of multiple storage nodes in the distributed storage system, for example, it may include: calculating the hash value of the second label of the time-series data to be stored, and taking the remainder of the hash value of the second label modulo 4. If the remainder is 0, the time-series data can be stored in storage node 12A; if the remainder is 1, the time-series data can be stored in storage node 12B; if the remainder is 2, the time-series data can be stored in storage node 12C; if the remainder is 3, the time-series data can be stored in storage node 12C. It should be noted that the number of the second labels can be one or more. When there are multiple second labels, for each second label, according to the result of taking the remainder of the hash value of the second label modulo 4, the partial time-series data corresponding to the second label in the time-series data to be stored can be stored in the corresponding storage node.
[0116] It should be noted that the label keys of the first label in Example 1 and the second label in Example 2 can be the same or different, and the present application does not limit this.
[0117] When the partitioning method corresponding to the target routing information is the range partitioning method, the target routing information can specifically be used to adopt the range partitioning method to disperse and store the time-series data of the target time partition into the routing information of multiple storage nodes in the distributed storage system.
[0118] Example 3, assume the target routing information is Figure 3For the routing information 1 in , when using the target routing information to store the time-series data into at least one of multiple storage nodes in the distributed storage system, it may include: determining the reporting device ID of the time-series data to be stored. If the reporting device ID belongs to the range of [1, 400], the time-series data can be stored in storage node 12A. If the reporting device ID belongs to the range of [401, 800], the time-series data can be stored in storage node 12B. If the device ID belongs to the range of [801, 1200], the time-series data can be stored in storage node 12C. It should be noted that the number of reporting device IDs can be one or more. When multiple reporting device IDs span different range intervals, for each range interval that the reporting device IDs span, the part of the time-series data corresponding to that range interval in the time-series data to be stored can be stored in the storage node corresponding to that range interval.
[0119] Example 4, assume the target routing information is Figure 3 For the routing information 2 in , when using the target routing information to store the time-series data into at least one of multiple storage nodes in the distributed storage system, it may include: determining the reporting device ID of the reported time-series data. If the reporting device ID belongs to the range of [1, 300], the time-series data can be stored in storage node 12A. If the reporting device ID belongs to the range of [301, 600], the time-series data can be stored in storage node 12B. If the reporting device ID belongs to the range of [601, 900], the time-series data can be stored in storage node 12C. If the reporting device ID belongs to the range of [901, 1200], the time-series data can be stored in storage node 12D.
[0120] It should be noted that in Examples 3 and 4, the range partitioning is based on the device ID. It can be understood that in other embodiments, other methods can also be used for range partitioning, and the present application does not limit this.
[0121] The data storage method provided by the embodiments of the present application has multiple time partitions divided according to time granularity in the distributed storage system, and there is corresponding routing information for the time partitions. According to the time information corresponding to the time-series data to be stored, the target time partition corresponding to the time-series data is determined, and the target routing information corresponding to the target time partition is used to store the time-series data into at least one of multiple storage nodes, realizing a multi-dimensional partitioning method of partitioning the time-series data on the same time partition in the node dimension on the basis of partitioning the time-series data in the time dimension, so that adding new storage nodes does not cause changes in the storage locations of the existing data. Therefore, when adding new storage nodes, there is no need to migrate the existing data between storage nodes.
[0122] Figure 5The following is a schematic flowchart of a data query method provided by an embodiment of the present application. As Figure 5 shown, the method of this embodiment may include:
[0123] Step 51, obtaining a query request for requesting to query time-series data;
[0124] Step 52, determining a target time partition corresponding to the time-series data requested to be queried in the distributed storage system according to the time information corresponding to the query request; there are multiple time partitions divided according to time granularity in the distributed storage system, and there is corresponding routing information for the time partition;
[0125] Step 53, searching for the time-series data from at least one storage node among multiple storage nodes in the distributed storage system according to the target routing information corresponding to the target time partition.
[0126] In an embodiment of the present application, the obtaining of the query request for requesting to query time-series data may be, for example, receiving a query request for requesting to query time-series data sent by a user device. Of course, in other embodiments, the query request for requesting to query time-series data may also be obtained by other means, and the present application does not limit this.
[0127] The time information corresponding to the query request refers to the time range information of the time-series data requested to be queried by the query request. In one embodiment, the time information corresponding to the query request may be the time condition information in the query request.
[0128] Example 5, assuming that a certain query request A is used to request to query time-series data that meets the label condition A1 on June 3, 2021, then the time information corresponding to the query request A is June 3, 2021. When Figure 3 the start time of time partition 3 is 00:00:01 on June 3, 2021, and the end time is 24:00:00 on June 3, 2021, it can be determined that the time-series data requested to be queried by the query request A corresponds to time partition 3 in the distributed storage system. Further, the routing information 1 corresponding to time partition 3 can be used to search for the time-series data requested to be queried by the query request A from at least one storage node among storage nodes 12A to 12C.
[0129] Based on Example 1 and Example 5, further assume that the label condition A1 includes the label key of the aforementioned first label and its corresponding label value (hereinafter referred to as the first label'), then when using the routing information 1 corresponding to time partition 3 to search for the time-series data requested by the query request A from at least one storage node among storage nodes 12A to 12C, it may include: calculating the hash value of the first label', taking the remainder of the hash value of the first label' modulo 3. If the remainder is 0, the time-series data can be searched from storage node 12A; if the remainder is 1, the time-series data can be searched from storage node 12B; if the remainder is 2, the time-series data can be searched from storage node 12C. It should be noted that the number of the first label' can be one or more. When there are multiple first labels', for each first label', partial time-series data can be searched from the corresponding storage node according to the result of taking the remainder of the hash value of the first label' modulo 3.
[0130] Example 6, assume that a certain query request B is used to request to query the time-series data on June 4, 2021 that meets the label condition B1, then the time information corresponding to the query request B is June 4, 2021. When Figure 3 the start time of time partition 4 is 00:00:01 on June 4, 2021, and the end time is 24:00:00 on June 4, 2021, it can be determined that the time-series data requested by the query request B corresponds to time partition 4 in the distributed storage system. Further, the routing information 2 corresponding to time partition 4 can be used to search for the time-series data requested by the query request B from at least one storage node among storage nodes 12A to 12D.
[0131] Based on Example 2 and Example 6, further assume that the label condition B includes the label key of the aforementioned second label and its corresponding label value (hereinafter referred to as the second label'), then when using the routing information 2 corresponding to time partition 4 to search for the time-series data requested by the query request B from at least one storage node among storage nodes 12A to 12D, it may include: calculating the hash value of the second label', taking the remainder of the hash value of the second label' modulo 4. If the remainder is 0, the time-series data can be searched from storage node 12A; if the remainder is 1, the time-series data can be searched from storage node 12B; if the remainder is 2, the time-series data can be searched from storage node 12C; if the remainder is 3, the time-series data can be searched from storage node 12D. It should be noted that the number of the second label' can be one or more. When there are multiple second labels', for each second label', partial time-series data can be searched from the corresponding storage node according to the result of taking the remainder of the hash value of the second label' modulo 4.
[0132] Example 7. Suppose a query request C is used to request the query of time-series data that meets the device ID condition C1 on June 3, 2021. Then the time information corresponding to the query request C is June 3, 2021. In Figure 3 when the start time of time partition 3 is 00:00:01 on June 3, 2021 and the end time is 24:00:00 on June 3, 2021, it can be determined that the time-series data requested by the query request C corresponds to time partition 3 in the distributed storage system. Further, the routing information 1 corresponding to time partition 3 can be used to search for the time-series data requested by the query request C from at least one storage node among storage nodes 12A to 12C.
[0133] Based on Examples 3 and 7, using the routing information 1 corresponding to time partition 3 to search for the time-series data requested by the query request C from at least one storage node among storage nodes 12A to 12C may include: determining a first target ID that meets the device ID condition C1. If the first target ID belongs to the range [1, 400], the time-series data can be searched from storage node 12A. If the first target ID belongs to the range [401, 800], the time-series data can be searched from storage node 12B. If the first target ID belongs to the range [801, 1200], the time-series data can be searched from storage node 12C. It should be noted that the number of first target IDs can be one or more. When multiple first target IDs span different ranges, for each range spanned by the first target IDs, partial time-series data can be searched from the storage node corresponding to that range.
[0134] Example 8. Suppose a query request D is used to request the query of time-series data that meets the device ID condition D1 on June 4, 2021. Then the time information corresponding to the query request D is June 4, 2021. In Figure 3 when the start time of time partition 4 is 00:00:01 on June 4, 2021 and the end time is 24:00:00 on June 4, 2021, it can be determined that the time-series data requested by the query request D corresponds to time partition 4 in the distributed storage system. Further, the routing information 2 corresponding to time partition 4 can be used to search for the time-series data requested by the query request D from at least one storage node among storage nodes 12A to 12D.
[0135] Based on Example 4 and Example 8, using the routing information 2 corresponding to time partition 4, searching for the timing data requested by the query request D from at least one storage node among storage nodes 12A to 12D may include: determining a second target ID that meets the device ID condition D1. If the second target ID belongs to the range [1, 300], the timing data can be searched from storage node 12A. If the target ID belongs to the range [301, 600], the timing data can be searched from storage node 12B. If the target ID belongs to the range [601, 900], the timing data can be searched from storage node 12C. If the target ID belongs to the range [901, 1200], the timing data can be searched from storage node 12D. It should be noted that the number of second target IDs can be one or more. When multiple second target IDs span different range intervals, for each range interval that the second target ID spans, partial timing data can be searched from the storage node corresponding to that range interval.
[0136] The data query method provided by the embodiments of the present application, through multiple time partitions divided according to time granularity existing in the distributed storage system and corresponding routing information existing for the time partitions, determines the target time partition corresponding to the timing data requested by the query request according to the time information corresponding to the query request, and uses the target routing information corresponding to the target time partition to search for the timing data from at least one storage node among multiple storage nodes in the distributed storage system, realizing the function of providing data decentralized query externally on the basis of data partition storage using a multi-dimensional partition method of time dimension partition + node dimension partition, enabling the outside world to query the timing data stored using the multi-dimensional partition method according to requirements.
[0137] Figure 6 It is a schematic flowchart of a data management method provided by an embodiment of the present application, as Figure 6 shown, the method of this embodiment may include:
[0138] Step 61, dividing multiple time partitions according to time granularity;
[0139] Step 62, when reaching the target moment within the time partition, according to the current storage node situation of the distributed storage system, adding corresponding routing information, so as to be able to use the routing information corresponding to the time partition to disperse and store the timing data corresponding to the time partition into multiple storage nodes of the distributed storage system.
[0140] It should be noted that for the specific content of the time granularity, the time partition, the routing information corresponding to the time partition, and adding the corresponding routing information, reference can be made to the relevant descriptions in the foregoing embodiments, and details are not elaborated herein.
[0141] The data management method provided by the embodiments of the present application divides multiple time partitions according to time granularity. When reaching the target moment within the time partition, corresponding routing information is added according to the current storage node situation of the distributed storage system, so as to be able to use the routing information corresponding to the time partition to disperse and store the time-series data of the corresponding time partition into multiple storage nodes of the distributed storage system. It realizes that there are multiple time partitions divided according to time granularity in the distributed storage system and there is corresponding routing information for the time partitions, enabling data partition storage to be carried out in a multi-dimensional partition manner of time dimension partition + node dimension partition.
[0142] Figure 7 It is a schematic structural diagram of a data storage device provided by an embodiment of the present application; refer to the appendix Figure 7 As shown, this embodiment provides a data storage device, and this device can execute the method provided by the above Figure 4 shown embodiment. Specifically, this device may include:
[0143] An acquisition module 71, configured to acquire the time-series data to be stored;
[0144] A determination module 72, configured to determine the target time partition corresponding to the time-series data in the distributed storage system according to the time information corresponding to the time-series data; there are multiple time partitions divided according to time granularity in the distributed storage system, and there is corresponding routing information for the time partitions;
[0145] A storage module 73, configured to store the time-series data into at least one of the multiple storage nodes of the distributed storage system according to the target routing information corresponding to the target time partition.
[0146] Optionally, this device may further include an addition module, configured to add corresponding routing information according to the current storage node situation of the distributed storage system when reaching the target moment within the time partition.
[0147] Optionally, the target moment includes the starting moment of the time partition.
[0148] Optionally, the target moment includes the moment when the time-series data with the corresponding time within the time partition is first acquired.
[0149] Optionally, after the target moment and before the next time partition of the time partition, a new storage node is added to the distributed storage system; the addition module is further configured to:
[0150] Modify the routing information corresponding to the time partition to the first routing information corresponding to the first time sub-partition of the time partition and the second routing information corresponding to the second time sub-partition of the time partition.
[0151] Optionally, the adding module is configured to modify the routing information corresponding to the time partition into first routing information corresponding to a first time sub - partition of the time partition and second routing information corresponding to a second time sub - partition of the time partition, specifically including: determining the node effective time; and, according to the node effective time, modifying the routing information corresponding to the time partition into first routing information corresponding to the time from the start time of the time partition to the moment before the node effective time, and second routing information corresponding to the time from the node effective time to the end time of the time partition.
[0152] Optionally, the adding module is configured to determine the node effective time, specifically including: obtaining the node effective time set by the operation and maintenance personnel.
[0153] Optionally, the adding module is configured to determine the node effective time, specifically including: taking the sum of the node adding time and the target duration as the node effective time.
[0154] Optionally, the apparatus may further include an indicating module, configured to indicate the time granularity to the storage nodes in the distributed storage system, so that the storage nodes perform data access according to the time granularity.
[0155] Optionally, the time granularity includes days.
[0156] Figure 7 The shown apparatus can execute Figure 4 The method of the shown embodiment. For parts not described in detail in this embodiment, reference can be made to the relevant descriptions of the Figure 4 shown embodiment. The execution process and technical effects of this technical solution are referred to the descriptions in the Figure 4 shown embodiment and will not be elaborated here.
[0157] In a possible implementation, Figure 7 The structure of the shown apparatus can be implemented as a computer device. As Figure 8 shown, the computer device may include: a processor 81 and a memory 82. Among them, the memory 82 is used to store the program related to supporting the computer device to execute the method provided in the above Figure 4 shown embodiment, and the processor 81 is configured to execute the program stored in the memory 82.
[0158] The program includes one or more computer instructions. Among them, when one or more computer instructions are executed by the processor 81, the following steps can be implemented:
[0159] Obtain the time - series data to be stored;
[0160] Determine the target time partition corresponding to the time series data in the distributed storage system according to the time information corresponding to the time series data; there are multiple time partitions divided according to time granularity in the distributed storage system, and there is corresponding routing information for the time partitions;
[0161] Store the time series data into at least one of the multiple storage nodes in the distributed storage system according to the target routing information corresponding to the target time partition.
[0162] Optionally, the processor 81 is further configured to execute all or part of the steps in the foregoing Figure 4 illustrated embodiments.
[0163] Among them, the structure of the computer device may further include a communication interface 83 for the computer device to communicate with other devices or communication networks.
[0164] Figure 9 It is a schematic structural diagram of a data query device provided in an embodiment of the present application; referring to the appendix Figure 9 As shown, this embodiment provides a data query device, and this device can execute the method provided in the foregoing Figure 5 illustrated embodiment. Specifically, this device may include:
[0165] An acquisition module 91, configured to acquire a query request for requesting to query time series data;
[0166] A determination module 92, configured to determine the target time partition corresponding to the time series data requested to be queried in the distributed storage system according to the time information corresponding to the query request; there are multiple time partitions divided according to time granularity in the distributed storage system, and there is corresponding routing information for the time partitions;
[0167] A search module 93, configured to search for the time series data from at least one of the multiple storage nodes in the distributed storage system according to the target routing information corresponding to the target time partition.
[0168] Figure 9 The illustrated device can execute Figure 5 the method in the illustrated embodiment. For parts not described in detail in this embodiment, reference may be made to the relevant descriptions of the Figure 5 illustrated embodiment. The execution process and technical effects of this technical solution are referred to the description in the Figure 5 illustrated embodiment and will not be elaborated here.
[0169] In a possible implementation, Figure 9 the structure of the illustrated device can be implemented as a computer device. As Figure 10As shown, the computer device may include: a processor 101 and a memory 102. Among them, the memory 102 is used to store programs related to the method provided in the above Figure 5 shown embodiments, and the processor 101 is configured to execute the programs stored in the memory 102.
[0170] The program includes one or more computer instructions. Among them, when the one or more computer instructions are executed by the processor 101, the following steps can be implemented:
[0171] Obtain a query request for requesting to query time series data;
[0172] According to the time information corresponding to the query request, determine the target time partition corresponding to the time series data to be queried in the distributed storage system; there are multiple time partitions divided according to time granularity in the distributed storage system, and there is corresponding routing information for the time partition;
[0173] According to the target routing information corresponding to the target time partition, search for the time series data from at least one storage node among multiple storage nodes in the distributed storage system.
[0174] Optionally, the processor 101 is further configured to execute all or part of the steps in the foregoing Figure 5 shown embodiments.
[0175] Among them, the structure of the computer device may further include a communication interface 103, which is used for the computer device to communicate with other devices or communication networks.
[0176] Figure 11 is a schematic structural diagram of a data management device provided in an embodiment of the present application; referring to the attached Figure 11 As shown, this embodiment provides a data management device, and the device can execute the above Figure 6 method provided in the shown embodiments. Specifically, the device may include:
[0177] A partitioning module 111, configured to partition multiple time partitions according to time granularity;
[0178] An adding module 112, configured to, when reaching the target moment within the time partition, add corresponding routing information according to the current storage node situation of the distributed storage system, so as to be able to use the routing information corresponding to the time partition to disperse and store the time series data corresponding to the time partition into multiple storage nodes in the distributed storage system.
[0179] Figure 11 The device shown can execute Figure 6 the method in the shown embodiments. For parts not described in detail in this embodiment, reference can be made to Figure 6Related description of the illustrated embodiment. For the execution process and technical effects of this technical solution, please refer to Figure 6 the description in the illustrated embodiment, which will not be elaborated here.
[0180] In a possible implementation, Figure 11 the structure of the illustrated device can be implemented as a computer device. As Figure 12 shown, the computer device may include: a processor 121 and a memory 122. Among them, the memory 122 is used to store programs related to the method provided in the above Figure 6 illustrated embodiment, and the processor 121 is configured to execute the programs stored in the memory 122.
[0181] The program includes one or more computer instructions. When one or more computer instructions are executed by the processor 121, the following steps can be achieved:
[0182] Divide multiple time partitions according to the time granularity;
[0183] When reaching the target moment within the time partition, according to the current storage node situation of the distributed storage system, add corresponding routing information, so as to be able to use the routing information corresponding to the time partition to disperse and store the time-series data corresponding to the time partition into multiple storage nodes of the distributed storage system.
[0184] Optionally, the processor 121 is further configured to execute all or part of the steps in the foregoing Figure 6 illustrated embodiment.
[0185] Among them, the structure of the computer device may further include a communication interface 123 for the computer device to communicate with other devices or communication networks.
[0186] In addition, an embodiment of the present application provides a computer storage medium for storing computer software instructions used by a computer device, which includes a program related to the method in the above Figure 4 illustrated method embodiment.
[0187] An embodiment of the present application provides a computer storage medium for storing computer software instructions used by a computer device, which includes a program related to the method in the above Figure 5 illustrated method embodiment.
[0188] An embodiment of the present application provides a computer storage medium for storing computer software instructions used by a computer device, which includes a program related to the method in the above Figure 6 illustrated method embodiment.
[0189] The embodiments of the present application also provide a distributed storage system, including: a proxy node and multiple storage nodes; the proxy node is used to execute Figure 4 the method described in the method embodiment shown.
[0190] The embodiments of the present application also provide a distributed storage system, including: a proxy node and multiple storage nodes; the proxy node is used to execute Figure 5 the method described in the method embodiment shown.
[0191] The embodiments of the present application also provide a distributed storage system, including: a proxy node and multiple storage nodes; the proxy node is used to execute Figure 6 the method described in the method embodiment shown.
[0192] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0193] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, can also be implemented by a combination of hardware and software. Based on such an understanding, the above technical solutions essentially or the part that contributes to the prior art can be embodied in the form of a computer product. The present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0194] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0195] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the acts Figure 1 or acts and / or boxes Figure 1 specified in one or more boxes.
[0196] These computer program instructions can also be loaded onto a computer or other programmable device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, thereby providing steps for implementing the functions specified in one or more of the acts Figure 1 or acts and / or boxes Figure 1 specified in one or more boxes.
[0197] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0198] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0199] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0200] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A data storage method, applied to a distributed storage system. There are multiple partitions in the distributed storage system. Each partition has a corresponding start time, end time, and multiple storage nodes. The multiple partitions include multiple first partitions with the same corresponding storage nodes, and the method includes: Obtain time-series data to be stored; According to the time information corresponding to the time-series data, determine that the target partition corresponding to the time-series data in the distributed storage system is a first partition; According to the target routing information corresponding to the first partition, disperse and store the time-series data into multiple storage nodes corresponding to the first partition; After adding storage nodes to the distributed storage system, the multiple partitions further include multiple second partitions with the same corresponding storage nodes. The multiple storage nodes corresponding to the multiple second partitions include the multiple storage nodes corresponding to the first partition and the added storage nodes; The method further includes: Obtain time-series data to be stored; According to the time information corresponding to the time-series data, determine that the target partition corresponding to the time-series data in the distributed storage system is a second partition; According to the target routing information corresponding to the second partition, disperse and store the time-series data into the multiple storage nodes corresponding to the second partition and the added storage nodes.
2. The method according to claim 1, the method further includes: When reaching the target moment within the partition, add corresponding routing information according to the current storage node situation of the distributed storage system.
3. The method according to claim 2, the target moment includes the start moment of the partition, or the moment when time-series data corresponding to the time within the partition is first obtained.
4. The method according to claim 2, after the target moment and before the next partition of the partition, storage nodes are added to the distributed storage system; after adding the corresponding routing information according to the current storage node situation of the distributed storage system when reaching the target moment within the partition, it further includes: Modify the routing information corresponding to the partition into a first routing information corresponding to the first time sub-partition of the partition and a second routing information corresponding to the second time sub-partition of the partition.
5. The method according to claim 4, the modifying the routing information corresponding to the partition into a first routing information corresponding to the first time sub-partition of the partition and a second routing information corresponding to the second time sub-partition of the partition includes: Determine the node effective moment; According to the node effective moment, modify the routing information corresponding to the partition into a first routing information corresponding to the time from the start moment of the partition to the moment before the node effective moment, and a second routing information corresponding to the time from the node effective moment to the end moment of the partition.
6. The method according to claim 5, wherein the determining the effective moment of the node includes: Obtain the node effective moment set by the operation and maintenance personnel, or use the sum of the node addition moment and the target duration as the node effective moment.
7. The method according to any one of claims 1-6, the method further includes: Instruct a storage node in the distributed storage system of a time granularity, so that the storage node performs data access according to the time granularity.
8. The method according to any one of claims 1-6, wherein the time granularity of the partition includes days.
9. A data query method applied to a distributed storage system, where there are multiple partitions in the distributed storage system, each partition has a corresponding start time, end time, and multiple storage nodes, the multiple partitions include multiple first partitions with the same corresponding storage nodes and multiple second partitions with the same corresponding storage nodes, and the multiple storage nodes corresponding to the multiple second partitions include the multiple storage nodes corresponding to the first partitions and newly added storage nodes, and the method includes: Obtain a query request for requesting to query time series data; When it is determined according to the time information corresponding to the query request that the target partition corresponding to the time series data to be queried in the distributed storage system is a first partition, search for the time series data from the multiple storage nodes corresponding to the first partition according to the target routing information corresponding to the first partition; When it is determined according to the time information corresponding to the query request that the target partition corresponding to the time series data to be queried in the distributed storage system is a second partition, search for the time series data from the multiple storage nodes corresponding to the second partition and the newly added storage nodes according to the target routing information corresponding to the second partition.
10. A data management method applied to a distributed storage system, where there are multiple partitions in the distributed storage system, each partition has a corresponding start time, end time, and multiple storage nodes, and the multiple partitions include multiple first partitions with the same corresponding storage nodes, and the method includes: Divide the start and end times according to a time granularity; When reaching a target moment within the start and end times, according to the current storage node situation after newly adding storage nodes to the distributed storage system, add corresponding routing information, so that the multiple partitions further include multiple second partitions with the same corresponding storage nodes, and the multiple storage nodes corresponding to the multiple second partitions include the multiple storage nodes corresponding to the first partitions and the newly added storage nodes.
11. A data management device applied to a distributed storage system, where there are multiple partitions in the distributed storage system, each partition has a corresponding start time, end time, and multiple storage nodes, and the multiple partitions include multiple first partitions with the same corresponding storage nodes, and the device includes: A division module for dividing the start and end times according to a time granularity; A new addition module for, when reaching a target moment within the start and end times, adding corresponding routing information according to the current storage node situation after newly adding storage nodes to the distributed storage system, so that the multiple partitions further include multiple second partitions with the same corresponding storage nodes, and the multiple storage nodes corresponding to the multiple second partitions include the multiple storage nodes corresponding to the first partitions and the newly added storage nodes.
12. A computer device, comprising: A memory and a processor; wherein the memory is used to store one or more computer instructions, and when the one or more computer instructions are executed by the processor, the method described in any one of claims 1 to 10 is implemented.
13. A distributed storage system, comprising: A proxy node and a plurality of storage nodes; the proxy node is used to execute the method described in any one of claims 1 to 10.
14. A computer program product, which is used to implement the method described in any one of claims 1 to 10 when the computer program is executed by a computer.
Citation Information
Patent Citations
Industrial time series data query processing method and system
CN107894997A
Distributed time sequence database, storage method and device and storage medium
CN112199419A