Method for injecting time series data, method for querying time series data and database system

By injecting time series data with mixed row and column storage in the database, the problems of low storage efficiency and poor query efficiency in the prior art are solved, and higher storage performance and query efficiency are achieved.

CN113868267BActive Publication Date: 2025-05-23HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010617592.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-30
Publication Date
2025-05-23
Estimated Expiration
2040-06-30

AI Technical Summary

Technical Problem

The prior art injects and querys time series data, low storage efficiency and poor query efficiency lead to waste of database performance.

Method used

Time sequence data is injected using mixed row and column storage. By storing data source parameters that do not change with time in a row storage format, time-changing indicators and timestamps are stored in a column storage format, and row identification is used to associate row and column storage.

Benefits of technology

It improves the storage performance and query efficiency of the database system, saves storage space, and quickly locates the required parameters during query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113868267B_ABST
    Figure CN113868267B_ABST
Patent Text Reader

Abstract

The present application discloses a method for injecting time series data, comprising: receiving time series data, the time series data including at least one parameter for identifying a data source generating the time series data, and an indicator and a timestamp representing at least one attribute of the data source, the timestamp indicating the time when the indicator was generated; storing a first parameter group in a row storage format, the first parameter group including at least one parameter for identifying a data source generating the time series data; storing a second parameter group in a column storage format, the second parameter group including an indicator and a timestamp representing at least one attribute of the data source. The embodiment of the present application injects time series data in a row-column hybrid storage method, which not only improves the storage performance (such as storage capacity, throughput) of the database, but also improves the efficiency of time series data query.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of database technology, and in particular to a method for injecting time series data, a method and device for querying time series data, and a database system. Background Art

[0002] With the development of various industries, the demand for databases is increasing. Currently, there are many types of databases, such as relational databases and time series databases. Among them, the demand for time series databases has increased significantly.

[0003] The data stored in a time series database is usually called time series data. Time series data is usually stored in rows, and a series of performance indicators of a device or an event at a time indicated by a timestamp are stored in a row. When querying some performance indicators of time series data, you only need to query one by one from the starting position of each time series data in order until you find the value of the performance indicator you need.

[0004] The current method of injecting time series data by row has low storage efficiency, and querying by row is not only inefficient but also wastes database performance. Summary of the invention

[0005] The present application provides a method for injecting time series data and a method for querying time series data, which are used to improve the storage performance (eg, storage capacity, throughput) and query efficiency of a database system. The present application also provides a corresponding device and a database system.

[0006] A first aspect of the present application provides a method for injecting time series data, comprising: receiving time series data, the time series data including at least one parameter for identifying a data source that generates the time series data, and an indicator and a timestamp representing at least one attribute of the data source, the timestamp indicating the time when the indicator was generated; storing a first parameter group in a row storage format, the first parameter group including at least one parameter for identifying a data source that generates the time series data; storing a second parameter group in a column storage format, the second parameter group including an indicator and a timestamp representing at least one attribute of the data source.

[0007] The method provided in the first aspect is applied to a database system, and is specifically applied to a time series database. A time series database stores time series data, and the time series data are all time-tagged. Usually, a time series data consists of a data source (tags), an indicator (field) and a timestamp (timestamp). In the present application, an indicator refers to the value of the attribute of the data source at the time indicated by the timestamp. Because the indicator changes over time, but the data source does not change over time, a data source will obtain many time series data over time, and the data source can also be called a "timeline (timeseries)". The data source refers to the source of the time series data, and the parameters in the first parameter group may include the name of the device, the identifier of the device, and the Internet protocol (IP) address of the device. In the present application, "at least one" includes "one" or "multiple", and "multiple" includes "two". "Multiple" can also be described as "at least two". The attributes of the data source described by the indicator can be output power, wind speed, throughput, frequency, input / output (I / O) and idle rate, etc. The attributes of the data source described by the specific indicator are related to the type of device. The index and the timestamp are interrelated, and the index corresponding to different timestamps are usually different. In the first aspect, at least one parameter that does not change with time and is used to identify the data source that generates the time series data is stored in a row storage format, and the index of the attribute that changes with time and the corresponding timestamp are stored in a column storage format. Because at least one parameter used to identify the data source that generates the time series data does not change with time, there is no need to repeatedly store at least one parameter for the time series data of the same data source. Storing at least one parameter in a row is not only conducive to saving storage space, but also to improving the storage performance of the database system. In addition, because at least one parameter needs to be read during query, storing at least one parameter in a row storage manner not only does not waste query resources, but also helps to quickly locate at least one parameter during query. Column storage is conducive to quickly finding the index of the corresponding attribute to be queried, so the present application scheme adopts a mixed row and column storage method to inject time series data, which not only improves the storage performance of the database, but also improves the efficiency of time series data query.

[0008] In a possible implementation of the first aspect, the above-mentioned step: determining the first parameter group and the second parameter group from the time series data, includes: determining a row identifier based on the first parameter group, and generating a data row based on the row identifier, at least one indicator and a timestamp; correspondingly, the above-mentioned step: storing the second parameter group in a column storage format, includes: storing at least one indicator contained in the data row in at least one column storage unit CU, at least one indicator corresponds one-to-one to at least one CU, and each CU in the at least one CU contains a row identifier; storing the timestamp in a CU, and the CU storing the timestamp contains the row identifier.

[0009] In this possible implementation, because the first parameter group is stored in a row storage format and the second parameter group is stored in a column storage format, the two parameter groups of a time series data will be stored in different locations. In order to realize the corresponding query of the two parameter groups, it is necessary to associate the two parameter groups through a row identifier (tagID). The row identifier is determined by the first parameter group. The first parameter group for all time series data of the same data source is the same, so the row identifier of all time series data of the same data source is also the same. After the row identifier is determined, for at least one indicator and timestamp to be stored in the column, the row identifier can be added to form a new data row (data row). For data rows, when storing in columns, an indicator will be stored in a column storage unit (compression unit, CU), and the timestamp will be placed in an independent CU. Each CU storing the indicator and the CU storing the timestamp contain the row identifier. A CU only stores an indicator of one attribute or only stores a timestamp, but a CU can store indicators of the same attribute in multiple data rows at different timestamps. For example: a CU stores 60 rows of wind speed values ​​of wind turbine 1, and each row of wind speed values ​​corresponds to a timestamp. It can be seen that the row identifier can ensure that after the time series data is mixedly stored in rows and columns, the data corresponding to the first parameter group in the row storage in the column storage can be found, thereby ensuring the accuracy of data query.

[0010] In a possible implementation manner of the first aspect, the step of storing the first parameter group in a row storage format includes storing the first parameter group and the row identifier in a row storage format.

[0011] In this possible time series mode, the row identifier is also stored in the row storage, so that the association between the row storage and the column storage can be established, so that when querying the time series data, the corresponding row identifier can be determined according to the first parameter group.

[0012] In a possible implementation of the first aspect, the above steps: determining the row identifier based on the first parameter group, include: determining an index value corresponding to at least one parameter in the first parameter group; querying the global cache to obtain the row identifier corresponding to the index value, the global cache storing the correspondence between the index value and the row identifier.

[0013] In this possible implementation, the index value is obtained by at least one parameter, and there are many ways to obtain the index value, such as: when there is a parameter in the first parameter group, the parameter can be used as the index value, or the parameter can be transformed, such as adding some numerical values ​​or symbols to obtain the index value. When there are multiple parameters in the first parameter group, these parameters can be spliced ​​together to obtain the index value, such as: including the name of the device and the IP address of the device, then the name of the device and the IP address of the device can be spliced ​​together to form an index value (key). Of course, it is not limited to this method, and the index value can also be obtained according to some parameters in the first parameter group, as long as the index value can uniquely identify a data source in the world, and the specific method of obtaining is not limited in this application. The global cache refers to the cache of the database system. If the global cache includes the index value, it means that the time series data of the data source has been stored before, and the row identifier corresponding to the data source has been generated. The row identifier can be determined directly according to the corresponding relationship between the index value and the row identifier, and then the time series data can be mixed row and column storage. This possible implementation method maintains the corresponding relationship between the row identifier and the index value through the global cache, which is conducive to improving the efficiency of injecting time series data.

[0014] In a possible implementation of the first aspect, the above steps: determining the row identifier based on the first parameter group, include: determining index values ​​corresponding to multiple parameters in the first parameter group; assigning row identifiers to the index values, and storing the correspondence between the index values ​​and the row identifiers in a global cache.

[0015] In this possible implementation, if the index value is not stored in the global cache, it means that the data of the data source has not been stored before, and a row identifier needs to be assigned and the corresponding relationship between the index value and the row identifier needs to be stored for subsequent use when there is time series data from the data source. This possible implementation maintains the corresponding relationship between the row identifier and the index value through the global cache, which is conducive to improving the efficiency of injecting time series data.

[0016] In a possible implementation manner of the first aspect, at least one CU storing at least one indicator and the CU storing a timestamp are located in a partition corresponding to a first time range, and the time indicated by the timestamp is within the first time range.

[0017] In this possible implementation, the first time range can be one month, one week, one day, multiple days, several hours or tens of minutes, and the first time range can be pre-configured. The specific value of the first time range is not limited in this application. A partition refers to a set of data within the first time range. The specific partition in which the time series data is suitable for storage can be determined by the timestamp in the data row, so that the partition can be quickly locked by the query time during subsequent queries, thereby improving the efficiency of the query.

[0018] In a possible implementation of the first aspect, the partition includes multiple data sets, each data set corresponds to a time range that does not completely overlap, at least one CU and a CU storing a timestamp are located in the same data set, and the time indicated by the timestamp is within the time range corresponding to the same data set.

[0019] In this possible implementation, a partition may include multiple data sets (parts), each of which corresponds to a time range, which can be represented by the minimum time and maximum time in the data set. The time range of the data set is smaller than the first time range of the partition. This is conducive to further improving the search efficiency of the data corresponding to the query row identifier.

[0020] In a possible implementation manner of the first aspect, the method further includes: merging at least two data sets from the multiple data sets to obtain a merged data set, wherein a second time range corresponding to the merged data set includes time ranges corresponding to at least two data sets, and the second time range is included in the first time range; and writing the merged data set into a data storage device.

[0021] In this possible implementation, the data storage is different from the cache and usually refers to a disk. The data set can be represented by small parts, and merging at least two data sets can be understood as merging small parts into large parts. In this way, when writing to the data storage, only a larger merged data set needs to be written once, and small data does not need to be written frequently, which can reduce the injection overhead of time series data and improve the injection performance of time series data. In addition, the second time range is usually smaller than the first time range. By narrowing the time range to narrow the search range, the query speed can be further improved.

[0022] In a possible implementation of the first aspect, the method further includes: compressing data in multiple CUs in the merged data set, and compressing multiple CUs to obtain a compressed merged data set; the above-mentioned step of: writing the merged data set into a data storage device includes: writing the compressed merged data set into the data storage device.

[0023] In this possible implementation, writing the merged data set to the data storage device can also be called "dropping to disk". Before the data is dropped to disk, not only the data in the CU is compressed, but also the CU is compressed. Through the double compression method, the compression ratio can be improved, and the throughput of the database system can be further improved.

[0024] The second aspect of the present application provides a method for querying time series data, including: receiving a query for time series data, the query including at least one parameter identifying a data source that generates the time series data, and at least one column identifier, the at least one column identifier indicating at least one target column, and the at least one target column containing an indicator representing at least one attribute of the data source; determining a row identifier corresponding to the at least one parameter; determining multiple CUs based on the row identifier, each CU in the multiple CUs containing the row identifier; determining at least one target CU from the multiple CUs based on the at least one column identifier, each target CU corresponding to a target column; and generating a query result based on the at least one target CU.

[0025] The method provided in the second aspect is applied to a database system, and specifically to a time series database. Some of the features involved in the second aspect can be understood by referring to the corresponding introduction of the first aspect, and will not be repeated here. It should be noted that at least one column identifier, at least one target column, at least one attribute indicator, and at least one target CU are all one-to-one corresponding. In the second aspect, the corresponding row identifier can be determined by identifying at least one parameter of the data source that generates the time series data, and then multiple CUs corresponding to the row identifier can be found in the column storage, and then the target CU is determined from the multiple CUs according to the column identifier. It can be seen that by identifying at least one parameter of the data source that generates the time series data and the column identifier, the target CU of the column corresponding to the column identifier can be quickly queried, and the query result is generated and returned to the client, which improves the efficiency of time series data query.

[0026] In a possible implementation manner of the second aspect, the above-mentioned step of: generating a query result according to at least one target CU includes: generating a query result according to an indicator of at least one attribute and at least one parameter in the at least one target CU.

[0027] In this possible implementation, an indicator of at least one attribute and at least one parameter in at least the target CU may be injected into the timing structure according to the timing data to form corresponding timing data, and then returned to the client.

[0028] In a possible implementation manner of the second aspect, the query further includes query time, and the method further includes:

[0029] Determine a target partition corresponding to a first time range including a query time, the target partition including multiple merged data sets; determine a first merged data set from the multiple merged data sets, the second time range corresponding to the first merged data set includes the query time, and the first merged data set includes multiple CUs; ​​determine multiple first CUs from the multiple CUs included in the first merged data set, the third time range corresponding to the first CU includes the query time, the first time range includes the second time range, and the second time range includes the third time range; correspondingly, the above step: determine multiple CUs based on a row identifier, includes: determine multiple CUs corresponding to the row identifier from the multiple first CUs.

[0030] In this possible implementation, the first time range can be pre-configured, and the second time range and the third time range can be represented by the minimum and maximum time values ​​in the corresponding merged data set or CU. The corresponding partition is determined by query time, and the query range is further narrowed to the merged data set, and then further narrowed to the CU. By narrowing the query range in three levels, there is no need to search for the target CU from a large amount of data, which improves the efficiency of data query.

[0031] In a possible implementation of the second aspect, the above-mentioned steps of: determining a row identifier corresponding to at least one parameter, include: determining an index value corresponding to at least one parameter; querying a global cache to obtain a row identifier corresponding to the index value, the global cache storing a correspondence between the index value and the row identifier.

[0032] In a possible implementation, the row identifier may also be determined according to a correspondence between at least one parameter in the row storage table and the row identifier.

[0033] In this possible implementation, the row identifier can be determined by the correspondence between the index value and the row identifier, and then data query can be performed, which can improve the efficiency of data query.

[0034] A third aspect of the present application provides a database system, including: a coordination node and a data node communicatively connected to the coordination node, the coordination node being used to receive time series data from a client; the data node being used to: obtain time series data from the coordination node (for example: receive time series data from the coordination node), the time series data including time series data, the time series data including at least one parameter for identifying a data source generating the time series data, and an indicator and a timestamp representing at least one attribute of the data source, the timestamp indicating the time when the indicator was generated; storing a first parameter group in a row storage format, the first parameter group including at least one parameter for identifying a data source generating the time series data; storing a second parameter group in a column storage format, the second parameter group including an indicator and a timestamp representing at least one attribute of the data source.

[0035] In a possible implementation of the third aspect, the data node is further used to: determine a row identifier based on the first parameter group, and generate a data row based on the row identifier, at least one indicator and a timestamp; store at least one indicator contained in the data row in at least one column storage unit CU, at least one indicator corresponds one-to-one to at least one CU, and each CU in the at least one CU contains the row identifier; store the timestamp in a CU, and the CU storing the timestamp contains the row identifier.

[0036] In a possible implementation manner of the third aspect, the data node is used to: store the first parameter group and the row identifier in a row storage format.

[0037] In a possible implementation of the third aspect, the data node is used to: determine an index value corresponding to at least one parameter in the first parameter group; query a global cache to obtain a row identifier corresponding to the index value, and the global cache stores a correspondence between the index value and the row identifier.

[0038] In a possible implementation manner of the third aspect, the data node is used to: determine index values ​​corresponding to multiple parameters in the first parameter group; assign row identifiers to the index values, and store the correspondence between the index values ​​and the row identifiers in a global cache.

[0039] In a possible implementation manner of the third aspect, at least one CU storing at least one indicator and the CU storing a timestamp are located in a partition corresponding to a first time range, and the time indicated by the timestamp is within the first time range.

[0040] In a possible implementation of the third aspect, the partition includes multiple data sets, wherein each data set corresponds to a time range that does not completely overlap, at least one CU and a CU storing a timestamp are located in the same data set, and the time indicated by the timestamp is within the time range corresponding to the same data set.

[0041] In a possible implementation of the third aspect, the data node is further used to: merge at least two data sets from the multiple data sets to obtain a merged data set, wherein a second time range corresponding to the merged data set includes time ranges corresponding to at least two data sets, and the second time range is included in the first time range; and write the merged data set into a data storage device.

[0042] In a possible implementation manner of the third aspect, the data node is further used to: compress data in multiple CUs in the merged data set, and compress multiple CUs to obtain a compressed merged data set; and write the compressed merged data set into a data storage device.

[0043] The third aspect and any possible implementation of the third aspect can be understood by referring to the first aspect and the corresponding possible implementation of the first aspect.

[0044] The fourth aspect of the present application provides a database system, including: a coordination node and a data node communicatively connected to the coordination node, the coordination node is used to receive queries for time series data from a client; the data node is used to: obtain queries for time series data from the coordination node, the query including at least one parameter identifying a data source that generates the time series data, and at least one column identifier, the at least one column identifier indicating at least one target column, and the at least one target column containing an indicator representing at least one attribute of the data source; determine a row identifier corresponding to at least one parameter; determine multiple CUs based on the row identifier, each CU in the multiple CUs containing the row identifier; determine at least one target CU from the multiple CUs based on the at least one column identifier, each target CU corresponding to a target column; generate a query result based on at least one target CU.

[0045] In a possible implementation manner of the fourth aspect, the data node is used to: generate a query result according to an indicator of at least one attribute in at least one target CU and at least one parameter.

[0046] In a possible implementation of the fourth aspect, the data node is also used to: determine a target partition corresponding to a first time range including a query time, the target partition including multiple merged data sets; determine a first merged data set from multiple merged data sets, the second time range corresponding to the first merged data set includes the query time, and the first merged data set includes multiple CUs; ​​determine multiple first CUs from the multiple CUs included in the first merged data set, the third time range corresponding to the first CU includes the query time, the first time range includes the second time range, and the second time range includes the third time range; determine multiple CUs corresponding to row identifiers from the multiple first CUs.

[0047] In a possible implementation of the fourth aspect, the data node is used to: determine an index value corresponding to at least one parameter; query a global cache to obtain a row identifier corresponding to the index value, and the global cache stores a correspondence between the index value and the row identifier.

[0048] The fourth aspect and any possible implementation of the fourth aspect can be understood by referring to the second aspect and the corresponding possible implementation of the second aspect.

[0049] In a fifth aspect of the present application, a device for injecting timing data is provided, which includes a module or unit for executing the method in the above-mentioned first aspect or any possible implementation of the first aspect, such as: a receiving unit, a determination unit and a processing unit. It should be noted that the determination unit and the processing unit can be implemented by one processing unit.

[0050] In a sixth aspect of the present application, a device for querying time series data is provided, which includes a module or unit for executing the method in the above-mentioned second aspect or any possible implementation of the second aspect, such as: a receiving unit, a first processing unit, a second processing unit and a third processing unit. It should be noted that the functions performed by these three processing units can also be implemented by one or two processing units.

[0051] In a seventh aspect of the present application, a device for injecting timing data is provided. The device may include at least one processor, a memory, and a communication interface. The processor is coupled to the memory and the communication interface. The memory is used to store instructions, the processor is used to execute the instructions, and the communication interface is used to communicate with other network elements under the control of the processor. When the instruction is executed by the processor, the processor executes the method in the first aspect or any possible implementation of the first aspect.

[0052] In an eighth aspect of the present application, a device for querying time series data is provided. The device may include at least one processor, a memory, and a communication interface. The processor is coupled to the memory and the communication interface. The memory is used to store instructions, the processor is used to execute the instructions, and the communication interface is used to communicate with other network elements under the control of the processor. When the instruction is executed by the processor, the processor executes the method in the second aspect or any possible implementation of the second aspect.

[0053] In a ninth aspect of the present application, a computer-readable storage medium is provided, which stores a program, wherein the program enables a processor to execute the method of injecting timing data in the first aspect and any one of its various implementations.

[0054] In a tenth aspect of the present application, a computer-readable storage medium is provided, which stores a program, and the program enables a processor to execute the method for querying time series data in the above-mentioned second aspect and any one of its various implementation methods.

[0055] In the eleventh aspect of the present application, a computer program product is provided, which includes computer execution instructions, which are stored in a computer-readable storage medium; at least one processor of the device can read the computer execution instructions from the computer-readable storage medium, and at least one processor executes the computer execution instructions so that the device implements a method for injecting timing data provided by the above-mentioned first aspect or any possible implementation of the first aspect.

[0056] In the twelfth aspect of the present application, a computer program product is provided, which includes computer execution instructions, which are stored in a computer-readable storage medium; at least one processor of the device can read the computer execution instructions from the computer-readable storage medium, and at least one processor executes the computer execution instructions so that the device implements a method for querying time series data provided by the above-mentioned second aspect or any possible implementation of the second aspect.

[0057] The thirteenth aspect of the present application provides a chip system, which includes a processor for supporting a device for injecting timing data to implement the functions involved in the above-mentioned first aspect or any possible implementation of the first aspect. In one possible design, the chip system may also include a memory, which is used to store the necessary program instructions and data for the device for injecting timing data. The chip system can be composed of chips, or it can include chips and other discrete devices.

[0058] In a fourteenth aspect of the present application, a chip system is provided, which includes a processor for supporting a device for querying time series data to implement the functions involved in the second aspect or any possible implementation of the second aspect. In a possible design, the chip system may also include a memory, which is used to store program instructions and data necessary for the device for querying time series data. The chip system may be composed of a chip, or may include a chip and other discrete devices.

[0059] It can be understood that any of the above-mentioned devices for injecting time series data, devices for querying time series data, computer storage media or computer program products are used to execute the corresponding methods for injecting time series data and methods for querying time series data provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1A It is a schematic diagram of the architecture of the database system;

[0061] Figure 1B It is a schematic diagram of the architecture of a distributed database system provided in an embodiment of the present application;

[0062] Figure 1C is another schematic diagram of the architecture of the distributed database system provided in an embodiment of the present application;

[0063] Figure 2 is another schematic diagram of the architecture of the distributed database system provided in an embodiment of the present application;

[0064] Figure 3 It is a schematic diagram of an embodiment of a method for injecting time series data provided by an embodiment of the present application;

[0065] Figure 4 is a schematic diagram of another embodiment of the method for injecting time series data provided by an embodiment of the present application;

[0066] Figure 5A is a schematic diagram of another embodiment of the method for injecting time series data provided by an embodiment of the present application;

[0067] Figure 5B is a schematic diagram of an example of a column storage unit provided in an embodiment of the present application;

[0068] Figure 5C This is a schematic diagram of a scenario example provided in an embodiment of the present application;

[0069] Figure 5D is another example schematic diagram of a scenario provided in an embodiment of the present application;

[0070] Figure 5E is another example schematic diagram of a scenario provided in an embodiment of the present application;

[0071] Fig. 5F is another example schematic diagram of a scenario provided in an embodiment of the present application;

[0072] Figure 6 is another example schematic diagram of a scenario provided in an embodiment of the present application;

[0073] Fig. 7A This is an effect comparison diagram provided by an embodiment of the present application;

[0074] Figure 7B is another effect comparison diagram provided by the embodiment of the present application;

[0075] Figure 8 It is a schematic diagram of an embodiment of a method for querying time series data provided by an embodiment of the present application;

[0076] Fig. 9A is another scenario schematic diagram provided by an embodiment of the present application;

[0077] Fig. 9B is another scenario schematic diagram provided by an embodiment of the present application;

[0078] Fig. 9C It is another effect schematic diagram provided by an embodiment of the present application;

[0079] Fig.10 It is a schematic diagram of an embodiment of a device for injecting time series data provided in an embodiment of the present application;

[0080] Fig.11 It is a schematic diagram of an embodiment of a device for querying time series data provided in an embodiment of the present application;

[0081] Fig.12 is a structural schematic diagram of a computer device provided in an embodiment of the present application;

[0082] Fig.13 It is another structural diagram of the database system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0083] The following describes the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of the present application, rather than all embodiments. It is known to those skilled in the art that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0084] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0085] The embodiment of the present application provides a method for injecting time series data and a method for querying time series data, which are used to improve the storage performance (such as storage capacity, throughput) and query efficiency of the database system. The embodiment of the present application also provides a corresponding device and a database system. The following are detailed descriptions.

[0086] The method provided in the embodiment of the present application can be applied to a database system. Figure 1A Shows a typical logical architecture of a database system. Figure 1A The database system 100 includes a database 110 and a database management system (DBMS) 130 .

[0087] Among them, the database 110 is an organized data set stored in the data storage 120, that is, a collection of associated data organized, stored and used according to a specific data model. According to the different data models used to organize data, data can be divided into multiple types, such as relational data, graph data, time series data, etc. Relational data is data modeled using a relational model, usually represented as a table, and the rows in the table represent a set of related values ​​of an object or entity. Graph data, referred to as "graph", is used to represent the relationship between objects or entities, such as social relationships. Time series data, referred to as time series data, is a data column recorded and indexed in chronological order, used to describe the state change information of an object in the time dimension.

[0088] The database management system 130 is the core of the database system and is a system software for organizing, storing and maintaining data. The client 200 can access the database 110 through the database management system 130, and the database administrator also performs database maintenance work through the database management system. The database management system 130 provides multiple functions for the client 200 to establish, modify and query the database, wherein the client 200 can be an application or a user device. The functions provided by the database management system 130 may include but are not limited to the following: (1) Data definition function. The database management system 130 provides a data definition language (DDL) to define the structure of the database 110. DDL is used to describe the database framework and can be saved in the data dictionary; (2) Data access function. The database management system 130 provides a data manipulation language (DML) to implement basic access operations on the database 110, such as retrieval, insertion, modification and deletion; (3) Database operation management function. The database management system 130 provides a data control function to effectively control and manage the operation of the database 110 to ensure that the data is correct and valid; (4) Database establishment and maintenance function, including loading of initial database data, database dumping, recovery, reorganization, system performance monitoring, analysis and other functions; (5) Database transmission. The database management system provides transmission of processed data to achieve communication between the client and the database management system, which is usually coordinated with the operating system.

[0089] The data storage 120 includes, but is not limited to, solid state drives (SSDs), disk arrays, cloud storage, or other types of non-transient computer-readable storage media. Figure 1AFewer or more components than those shown in the figure, or including Figure 1A The components shown are different components, Figure 1A Only components more relevant to the implementation disclosed in the embodiment of the present invention are shown.

[0090] The database system provided in the embodiment of the present application may be a distributed database system (DDBS). It may be a database system with a massively parallel processor (MPP) architecture. The database system with the MPP structure is also a DDBS. Figure 1B and Figure 1C Introducing DDBS.

[0091] Figure 1BThe schematic diagram of a distributed database system using a shared-storage architecture includes one or more coordinator nodes (CN), multiple data nodes (DN). Of course, the DDBS may also include other devices, such as a global transaction manager (GTM). CN and DN communicate through a network channel. DN may include a timing engine, which can implement timing data injection, timing data query and other timing data-related functions. CN may include a computing engine, which can determine an execution plan for a query and then distribute the query to a corresponding DN for execution. In one embodiment, the network channel may be composed of network devices such as switches, routers and gateways. CN and DN jointly implement the functions of a database management system and provide database retrieval, insertion, modification and deletion services to the client. In one embodiment, a database management system is deployed on each CN and DN. The shared data storage stores data that can be shared by multiple DNs, and DN can perform read and write operations on the data in the data storage through the network channel. The shared data storage may be a shared disk array. The CN and DN in the distributed database system can be physical machines, such as database servers, or virtual machines (VM) or containers running on abstract hardware resources. In one embodiment, the CN and DN are virtual machines or containers, and the network channel is a virtual switching network, which includes a virtual switch. The database management system deployed in the CN and DN is a DBMS instance, which can be a process or a thread, and these DBMSs cooperate to complete the functions of the database relational system. In another embodiment, the CN and DN are physical machines, and the network channel includes one or more switches, which are storage area network (SAN) switches, Ethernet switches, fiber switches or other physical switching devices.

[0092] Figure 1C This is a schematic diagram of a distributed database system using a shared-nothing architecture. Each DN has its own exclusive hardware resources (such as data storage), operating system, and database. The CN and DN communicate through a network channel. The network channel can be found in the above Figure 1B Under this system, data will be distributed to each DN according to the database model and application characteristics. The query task will be divided into several parts by CN and executed in parallel on all DNs. They will coordinate calculations with each other and provide database services as a whole. All communication functions are implemented on a high-bandwidth network interconnection system. Figure 1B Like the distributed database system with shared-storage architecture described in the previous section, the CN and DN here can be either physical machines or virtual machines.

[0093] In all embodiments of the present application, the data storage of the database system includes but is not limited to solid state drives (SSDs), disk arrays, or other types of non-transitory computer-readable media. Figure 1B-Figure 1C Although the database is not shown in the figure, it should be understood that the database is stored in the data storage. A person skilled in the art can understand that a database system may include Figure 1A-Figure 1C Fewer or more components than those shown in the figure, or including Figure 1A-Figure 1C The components shown in the figure are different components. Figure 1A-Figure 1C Only components more relevant to the implementation disclosed in the embodiments of the present application are shown. However, those skilled in the art can understand that a distributed database system can include any number of CNs and DNs. The database management system functions of each CN and DN can be implemented by a suitable combination of software, hardware and / or firmware running on each CN and DN.

[0094] Above Figure 1B and Figure 1C The database system in FIG. 1 shows a computing engine and a timing engine. In the solution of the present application, the functions of the computing engine and the timing engine can be referred to in Figure 2 Understand. Figure 2 The DBMS 130 in the database system shown includes a storage engine 170 and a computing engine 132 .

[0095] The computing engine 132 supports at least one type of query language, such as structured query language (SQL), and other queries that support time series data. The main function of the computing engine 130 is to generate a corresponding execution plan based on the query (Query) submitted by the client 200, and perform data operations according to the execution plan to generate query results. For the database system of the time series database, the computing engine mainly includes a query engine and an execution engine. Among them, the query engine mainly completes the parsing of queries, the rewriting of queries, and the generation of execution plans; the execution engine consists of operation operators and their related execution environments. Commonly used operation operators include scan, hash join, aggregate, etc., and the execution environment is mainly composed of an execution framework and a resource manager.

[0096] The storage engine 170 is responsible for providing the computing engine with an interface for accessing data on top of the file system, and also provides index management, and management of cache, transactions, logs, and other data during runtime. For example, the storage engine 170 can write the execution results of the computing engine 132 to the data storage 120 through physical I / O.

[0097] The storage engine 170 includes a timing engine 171, an adapter 172, a row storage engine 173, and a column storage engine 174. The timing engine 171 is used to manage timing data. The adapter 172 provides an interface between the timing engine 171 and the row storage engine 173, as well as the column storage engine 174, and adapts the row storage engine 173 or the column storage engine 174 according to the parameter types of different parts in the timing data. The row storage engine 173 is used to store the timing data in a row storage format, and the column storage engine 174 is used to store the timing data in a column storage format.

[0098] The computing engine 132 and the storage engine 170 are used in the process of querying time series data, and the storage engine 170 is used in the process of injecting time series data.

[0099] When injecting time series data, the time series engine 171 will split the time series data, and then use the row storage engine 173 to store the part of the time series data that needs row storage in row storage format, and use the column storage engine 174 to store the part of the time series data that needs column storage in column storage format.

[0100] When querying time series data, the computing engine 132 receives a query from the client 200 (the query is also called a "query statement" or "statement" in some scenarios), analyzes the query, and establishes a time series scanning operator. The time series engine 171 starts to scan the data after receiving the analysis results of the computing engine 132. The time series engine 171 completes the time series data query in a mixed row and column scanning manner. The computing engine 132 receives the time series data returned by the timing engine 171, that is, the query result, and returns the query result to the client 200.

[0101] The data storage 120 includes a row storage file, a column storage file, and an inverted index file. The row storage file contains time series data stored in a row storage format, the column storage file contains time series data stored in a column storage format, and the inverted index file contains the correspondence between the index value and the row identifier. The row identifier can be found by the index value, which can be understood as an inverted index, and the file containing the inverted index is called an inverted index file.

[0102] Based on the distributed database system introduced above, the following introduces the method of injecting time series data and the method of querying time series data respectively.

[0103] For ease of understanding, it is noted in advance that in the present application in the embodiments of the present application, "at least one" includes "one" or "multiple", and "multiple" includes "two". "Multiple" can also be described as "at least two".

[0104] Figure 3 A schematic diagram of an embodiment of a method for injecting timing data provided in an embodiment of the present application.

[0105] like Figure 3 As shown, an embodiment of the method for injecting timing data provided by an embodiment of the present application includes:

[0106] 201. Receive timing data.

[0107] The time series data includes at least one parameter for identifying a data source that generates the time series data, and an indicator and a timestamp representing at least one attribute of the data source, wherein the timestamp indicates the time when the indicator is generated.

[0108] It can be understood that the first parameter group and the second parameter group can be a logical division of the parameters, indicators and timestamps contained in the time series data.

[0109] The time series database stores time series data, and the time series data are all time-tagged. Usually, a time series data consists of a data source (tags), an indicator (field) and a timestamp (timestamp). In this application, an indicator refers to the value of the attribute of the data source at the time indicated by the timestamp. Because the indicator changes over time, but the data source does not change over time, many time series data will be obtained for a data source over time, and the data source can also be called a "timeline (timeseries)". The data source refers to the source of the time series data. The parameters in the first parameter group may include the name of the device, the identifier of the device, and the Internet protocol (IP) address of the device. The attributes of the data source described by the indicator can be output power, wind speed, throughput, frequency, input / output (I / O) and idle rate, etc. The attributes of the data source described by the specific indicator are related to the type of device. The indicator and the timestamp are interrelated, and the indicators corresponding to different timestamps are usually different.

[0110] For more information about the time series data, please refer to Table 1.

[0111] Table 1: Timing data

[0112] Name of the device The IP address of the device I / O Idle rate Timestamp ombi 10.73.73.3 21.15% 94.5% 2020-03-07 12:01:01 sds 10.93.19.141 1% 99% 2020-03-07 12:01:01 sds 10.93.20.138 5% 98.9% 2020-03-07 12:01:01 nsp 10.1.142.176 0.8% 91% 2020-03-07 12:01:01

[0113] As shown in Table 1 above, Table 1 includes 4 time series data, and each row is a time series data. The first two columns in Table 1 can be understood as the data source (tags) of the time series data, and tags can also be called timelines. I / O and idle rate (idle) are two attributes of the data source, and the values ​​of the I / O column and the idle rate column are indicators. The fifth column timestamp is the time when the indicator of the row is obtained.

[0114] In addition, it should be noted that a device can have multiple data sources. For example, if the device name is the same but the IP address is different, it also represents two different data sources. For example, in the second and third rows of Table 1, the device name is the same, both are sds, but the IP address of the device is different. The timing data in the second row and the timing data in the third row come from different data sources. Only when all parameters in the first parameter group are the same, it means that different timing data come from the same data source.

[0115] 202. Store the first parameter group in a row storage format.

[0116] The first parameter group includes at least one parameter for identifying a data source that generates time series data. As shown in Table 1 above, the name of the device and the IP address of the device can be classified into the first parameter group.

[0117] 203. Store the second parameter group in a column storage format.

[0118] The second parameter group includes an index representing at least one attribute of the data source and a timestamp. As shown in Table 1 above, the I / O volume, idle rate and timestamp can be classified into the second parameter group.

[0119] In a possible embodiment, before storing the time series data, the time series data may be segmented to separate at least one parameter for identifying a data source generating the time series data from an indicator and a timestamp representing at least one attribute of the data source.

[0120] In the embodiment of the present application, at least one parameter that does not change with time and is used to identify the data source that generates time series data is stored in a row storage format, and the index of the attribute that changes with time and the corresponding timestamp are stored in a column storage format. Because at least one parameter used to identify the data source that generates time series data does not change with time, there is no need to repeatedly store at least one parameter for the time series data of the same data source. Row storage of at least one parameter is not only conducive to saving storage space, but also to improving the storage performance of the database system. In addition, because at least one parameter needs to be read during query, storing at least one parameter in a row storage manner not only does not waste query resources, but also helps to quickly locate at least one parameter during query. Column storage is conducive to quickly finding the index of the corresponding attribute to be queried. Therefore, the present application scheme adopts a mixed row and column storage method to inject time series data, which not only improves the storage performance of the database, but also improves the efficiency of time series data query.

[0121] The solution for injecting time series data provided in the present application may include the following two solutions: 1. Segmentation of mixed row and column storage; 2. Two-layer cache of the column storage part.

[0122] 1. Splitting of mixed row and column storage.

[0123] The process can be found in Figure 4 Understand. Figure 4 As shown in the figure, the splitting of time series data for mixed row and column storage includes:

[0124] 301. Determine an index value according to at least one parameter used to identify a data source that generates time series data.

[0125] The index value is obtained through at least one parameter. There are many ways to obtain the index value, such as: when there is one parameter, the parameter can be used as the index value, or the parameter can be transformed, such as adding some numerical values ​​or symbols to obtain the index value. When there are multiple parameters, these parameters can be spliced ​​together to obtain the index value, such as: multiple parameters include the name of the device and the IP address of the device, then the name of the device and the IP address of the device can be spliced ​​together to form an index value (key). Of course, it is not limited to this method, and the index value can also be obtained based on multiple parameters. As long as the index value can uniquely identify a data source in the world, the specific method of obtaining it is not limited in this application.

[0126] 302. Query the global cache according to the index value. If it exists, execute step 303. If it does not exist, execute step 304.

[0127] The global cache refers to the cache of the database system.

[0128] 303. If the index value exists in the global cache, determine the row identifier corresponding to the index value. The global cache stores the corresponding relationship between the index value and the row identifier.

[0129] The row identifier (tagID) can be represented by numerical values ​​such as 1, 2, 3, or by characters such as a, b, c, d. Of course, the row identifier is not limited to these two representations, and can also be represented in other forms, as long as each row identifier can uniquely identify a data source.

[0130] The correspondence between index values ​​and row identifiers can be maintained in a table or in other mapping ways.

[0131] 304. If the index value does not exist in the global cache, a row identifier is assigned to the index value, and a corresponding relationship between the index value and the row identifier is stored in the global cache.

[0132] 305. Combining a row identifier, a value of at least one indicator, and a timestamp into a data row.

[0133] After the travel identifier is determined, a correspondence between the first parameter group and the travel identifier may be established, and then a correspondence between the second parameter group and the travel identifier may be established.

[0134] The corresponding relationship between the first parameter group and the row identifier can be expressed in the form of Table 2 in combination with the time series data in Table 1 above:

[0135] Table 2: Row storage data

[0136] Name of the device The IP address of the device Row ID ombi 10.73.73.3 1 sds 10.93.19.141 2 sds 10.93.20.138 3 nsp 10.1.142.176 4

[0137] The corresponding relationship between the second parameter group and the row identifier can be expressed in the form of Table 3 in combination with the time series data in Table 1 above:

[0138] Table 3: Data of column storage part

[0139] I / O Idle rate Timestamp Row ID 21.15% 94.5% 2020-03-07 12:01:01 1 1% 99% 2020-03-07 12:01:01 2 5% 98.9% 2020-03-07 12:01:01 3 0.8% 91% 2020-03-07 12:01:01 4

[0140] Each row in the table 3 may be referred to as a data row.

[0141] The above process completes the segmentation of the mixed row and column storage of time series data. For example, the four time series data in Table 1 can be segmented into the forms in Table 2 and Table 3 according to the above-mentioned principle. For example, the part of Table 2 stores the device name, device IP address and row identifier in row storage format. The column storage part can be written to disk through two-layer cache. Writing to disk means writing to data storage. The process of two-layer cache is introduced below.

[0142] The solution provided in the embodiment of the present application maintains the correspondence between the row identifier and the index value through a global cache, which is conducive to improving the efficiency of injecting time series data.

[0143] 2. Two-layer cache of the column storage part.

[0144] The two-layer caching process of the column storage part can be found in Figure 5A To understand.

[0145] like Figure 5A As shown, the process includes:

[0146] 401. Store the data row into the first-level cache.

[0147] If the first-layer cache is not full and the timer refresh time has not arrived, the process can be terminated. If the first-layer cache is full or the timer refresh time has arrived, step 402 can be executed. Of course, it is also possible to directly execute step 402 without waiting for the first-layer cache to be full or the timer refresh time to arrive. The waiting method is only conducive to reducing I / O overhead.

[0148] The first level of cache can be a global cache.

[0149] 402. Divide the data row into partitions corresponding to the first time range according to the timestamp in the data row.

[0150] At least one CU storing at least one indicator and the CU storing a timestamp are located in a partition corresponding to a first time range, and the time indicated by the timestamp is within the first time range.

[0151] The first time range can be one month, one week, one day, multiple days, several hours or tens of minutes, and the first time range can be pre-configured. The specific value of the first time range is not limited in this application. A partition refers to the data set corresponding to the first time range. The timestamp in the data row can be used to determine the specific partition in which the time series data is suitable for storage, so that the partition can be quickly locked by query time during subsequent queries, thereby improving the efficiency of the query.

[0152] A partition includes multiple column storage units (compression units, CUs). In column storage, indicators of the same attribute at different times can be stored in one CU, and the timestamp will be placed in a separate CU. Then, the association between multiple CUs of the same data source is established through row identifiers. The association can be that each CU that stores the indicators in the data row and the CU that stores the timestamp contain the row identifier. A CU only stores one column of data, but can store indicators of the same attribute in multiple data rows, such as: a CU stores the indicators of I / O at different times in Table 3 above.

[0153] The data stored in a CU can be found in Figure 5B Understand. Figure 5B As shown, the CU stores the index of the I / O column with tagID=1.

[0154] The association between CUs with the same tagID can be found in Figure 5C Understand. Figure 5C As shown, CU1 stores the I / O index of tagID=1, CU2 stores the idle rate index of tagID=1, and CU3 stores the timestamp data of tagID=1.

[0155] The timestamp in the data row can be used to determine the specific partition where the time series data is suitable for storage. This makes it convenient to quickly lock the partition by query time during subsequent queries, thereby improving query efficiency.

[0156] 403. According to the row identifier, search for a CU containing the row identifier from multiple CUs.

[0157] If the row identifier = 1, CU1, CU2 and CU3 can be found.

[0158] 404. Store each indicator in the data row in a different CU.

[0159] If the I / O of the data row is 35%, the idle rate is 92%, and the timestamp is 2020-03-07 14:15:23, then the values ​​of each column of this data row are stored in CU1, CU2, and CU3 respectively. After adding, the data in CU1, CU2, and CU3 can participate in Figure 5D To understand.

[0160] 405. Multiple CUs associated with the row identifiers are divided into the same data set. The data set (small part) can also be understood as a second-level cache.

[0161] The partition includes multiple data sets, each of which corresponds to a time range that does not completely overlap, at least one CU and a CU storing a timestamp are located in the same data set, and the time indicated by the timestamp is within the time range corresponding to the same data set. The time range can be represented by the minimum time and the maximum time in the data set.

[0162] The data set is a collection of CUs at the next level of partitioning. The relationship between small parts and CUs can be found in Figure 5E To understand. Figure 5E Each small part shown in the figure includes CU description information, which records the time range of each CU and the corresponding tagID information.

[0163] In the same partition, CUs containing the same row identifier may also be divided into the same data set (part), which is helpful to further improve the search efficiency of the CU corresponding to the query row identifier.

[0164] 406. Merge at least two data sets from the multiple data sets to obtain a merged data set, and write the merged data set into a data storage device.

[0165] The data storage is another storage medium different from the global cache, such as a disk.

[0166] The second time range corresponding to the merged data set includes the time ranges corresponding to at least two data sets, and the second time range is included in the first time range;

[0167] This process can also be understood as merging small parts into large parts, such as Fig. 5F As shown, small part 1 and small part 2 are combined into large part 1. In this way, when writing to the data storage, only a larger merged data set needs to be written once, and small data does not need to be written frequently, which can reduce the injection overhead of time series data and improve the injection performance of time series data. In addition, the second time range is usually smaller than the first time range. By narrowing the time range to narrow the search range, the query speed can be further improved.

[0168] For the above two-layer caching of data rows, please refer to Figure 6 The scenario shown in the figure can be understood. Figure 6 As shown in the figure, the process includes: storing the data row into the global cache. The global cache includes multiple partitions, such as partition 1 and partition 2. According to the timestamp in the data row, it is determined that the data row should be divided into partition 1. Partition 1 includes multiple small parts, such as small part 1, small part 2 and small part 3. Small parts belong to the second-level cache. Each small part includes multiple CUs and CU description files. For the relationship between small parts and CUs, please refer to Figure 5E According to the time range, the small parts are combined into a large part. For example, small part 1, small part 2 and small part 3 are combined into a large part. The process can be referred to in Fig. 5F Then the large part is put into disk, that is, stored in data storage.

[0169] In addition, in the embodiment of the present application, the data in the multiple CUs in the merged data set are compressed, and then the multiple CUs are compressed to obtain a compressed merged data set; the compressed merged data set is written into the data storage device. Through the double compression method, the compression ratio can be improved, and the throughput of the database system can be further improved.

[0170] The solution provided in the embodiment of the present application is far better than the time series database in the prior art in terms of database storage performance, such as throughput, and data compression performance. Fig. 7A and Figure 7B It is a bar graph generated from experimental data recorded by engineers. Fig. 7A As shown in the throughput comparison chart, there are two parameters in the first parameter group (the number of tags is 2) and one indicator in the second parameter group (the number of fields is 1). When a batch of 10,000 records (the batch size is 10K time series data) are injected, the comparison of the cardinality of 1, 10, 100, 1000 and 10,000 shows that the solution for injecting time series data provided by the present application increases the throughput of the database system by at least 30% compared with the throughput of the existing influxdb. Figure 7B As shown in the compression performance comparison chart, when the number is 1 million (w), 1000w, 10000w and 100000w, the solution for injecting time series data provided by the present application improves the data compression performance by at least 50% compared with Influxdb.

[0171] The above describes the process of injecting time series data. The following describes the method for querying time series data provided in an embodiment of the present application in conjunction with the accompanying drawings.

[0172] like Figure 8 As shown, an embodiment of the method for querying time series data provided by an embodiment of the present application includes:

[0173] 501. Receive a query for time series data.

[0174] The query includes at least one parameter identifying a data source that generates time series data, and at least one column identifier, wherein the at least one column identifier indicates at least one target column, and the at least one target column contains an indicator representing at least one attribute of the data source.

[0175] The meaning and relationship of at least one parameter, attribute, indicator, etc. for identifying the data source that generates the time series data can be understood by referring to the relevant content of the above-mentioned embodiment of injecting time series data, and will not be repeated here.

[0176] 502. Determine a row identifier corresponding to at least one parameter.

[0177] This step may be determining the index value according to at least one parameter; and determining the row identifier corresponding to the index value according to the index value and the corresponding relationship between the index value and the row identifier.

[0178] The process can be found in the above Figure 4 The corresponding steps 301 and 303 of the embodiment part are understood.

[0179] 503. Determine multiple CUs according to the row identifier, where each CU in the multiple CUs includes the row identifier.

[0180] For the relationship between row identifiers and multiple CUs, see Figure 5C and Figure 5D The contents of the data row can be understood by referring to Table 3.

[0181] 504. Determine at least one target CU from multiple CUs according to at least one column identifier; and generate a query result according to the at least one target CU.

[0182] The column identifier can be represented by a number, such as: taking Table 3 as an example, if the column identifier = 1, it corresponds to the I / O column, and if the column identifier = 2, it corresponds to the idle rate column. Of course, the column identifier can also be represented by other methods, which are not limited in this application.

[0183] It should be noted that at least one column identifier, at least one target column, at least one attribute indicator, and at least one target CU are all in one-to-one correspondence.

[0184] In a possible embodiment, generating a query result according to at least one target CU includes: generating a query result according to an indicator of at least one attribute and at least one parameter in the at least one target CU.

[0185] The index of at least one attribute and at least one parameter in at least the target CU may be injected into the timing structure according to the timing data to form corresponding timing data, and then returned to the client.

[0186] The solution provided by the embodiment of the present application can determine the corresponding row identifier by identifying at least one parameter of the data source that generates the time series data, and then find multiple CUs corresponding to the row identifier in the column storage, and then determine the target CU from the multiple CUs according to the column identifier. It can be seen that by identifying at least one parameter of the data source that generates the time series data and the column identifier, the target CU of the column corresponding to the column identifier can be quickly queried, and the query result is generated and returned to the client, which improves the efficiency of time series data query.

[0187] In one embodiment, the query also includes query time, and the method also includes: the method also includes: determining a target partition corresponding to a first time range including the query time, the target partition including multiple merged data sets; determining a first merged data set from multiple merged data sets, the second time range corresponding to the first merged data set includes the query time, and the first merged data set includes multiple CUs; ​​determining multiple first CUs from the multiple CUs included in the first merged data set, the third time range corresponding to the first CU includes the query time, the first time range includes the second time range, and the second time range includes the third time range; correspondingly, determining multiple CUs corresponding to the row identifier from the multiple first CUs.

[0188] This process can also be understood as a three-layer tailoring process, such as Fig. 9A As shown in the figure, the corresponding partition is determined by the query time, and the query range is further narrowed down to the part, and then the query range is further narrowed down to the CU. By narrowing the query range through three levels, there is no need to search from a large amount of data, which improves the efficiency of data query.

[0189] The database system provided in the embodiment of the present application generates a query plan according to the query in the process of querying time series data, and then submits the execution plan to the scanning operator, and then the scanning operator calls the storage interface to query the corresponding time series data. The entire query process can be referred to Fig. 9B To understand, such as Fig. 9B The process shown includes:

[0190] 1. Initialization of the timing scan operator

[0191] When the time series scan operator is initialized, a scan state (TsStoreScanState) object is created. The object implements the interaction between the execution layer and the storage layer. The object stores context information, including at least one parameter for identifying the data source that generates the time series data, column identifiers, and query time, etc. During initialization, the tagId is queried using at least one parameter for identifying the data source that generates the time series data, and is saved in the scan state object.

[0192] 2. Timing Scan Operator Execution

[0193] The execution of the timing scan operator is driven by the upper-level partition operator. The storage layer provides an interface to implement data scanning through the search chain of global search (TsStoreSearch) -> partition search (PartitionSearch) -> data set search (PartSearch). PartSearch, as the smallest search unit, completes specific data scanning. The specific process includes: scanning the CU description file (cudesc) according to the tagId set in the TsStoreScanState object, finding the corresponding column identifier (columnId), obtaining the data in the corresponding target CU through the columnId and returning it, and then splicing the CU into vector rows (VectorBatch) and giving it to the timing scan operator.

[0194] 3. Timing Scan Operator Reset

[0195] When the upper-level partition operator finishes scanning the current partition data, it will call the reset interface of the lower-level operator to notify the partition switch. The timing scanning operator reinstallation mainly resets PartitionSearch and PartSearch to realize the function of switching partitions.

[0196] 4. TsStoreScan operator execution ends

[0197] When all the time series data is scanned, call the operator end interface of TsStoreScan to release the related resources.

[0198] The solution provided by the embodiments of the present application is far better than the time series database in the prior art in terms of database query performance. Fig. 9C It is a bar graph generated from experimental data recorded by engineers. Fig. 9C As shown in the comparison chart of normalized query delays of Influxdb and GuassDB TSDB in the basic resource monitoring scenario of the consumer cloud, it can be seen that in the elastic load balancing (ELB) scenario, as well as high hashing degree and large data volume such as single-group I / O, the database system (GaussDB TSDB) provided in the embodiment of the present application is 2 times better than Influxdb in query delay performance.

[0199] The above multiple embodiments describe a method for injecting time series data and a method for querying time series data. The following describes an apparatus for injecting time series data and an apparatus for querying time series data provided in embodiments of the present application in conjunction with the accompanying drawings.

[0200] like Fig.10 As shown, an embodiment of the device 60 for injecting time series data provided in an embodiment of the present application includes:

[0201] The receiving unit 601 is used to receive time series data, which includes at least one parameter for identifying a data source that generates the time series data, and an indicator and a timestamp representing at least one attribute of the data source, wherein the timestamp indicates the time when the indicator is generated.

[0202] Processing unit 602 is used to store a first parameter group in a row storage format, the first parameter group includes at least one parameter for identifying a data source that generates time series data; and store a second parameter group in a column storage format, the second parameter group includes an indicator and a timestamp representing at least one attribute of the data source.

[0203] In the embodiment of the present application, at least one parameter that does not change with time and is used to identify the data source that generates time series data is stored in a row storage format, and the index of the attribute that changes with time and the corresponding timestamp are stored in a column storage format. Because at least one parameter used to identify the data source that generates time series data does not change with time, there is no need to repeatedly store at least one parameter for the time series data of the same data source. Row storage of at least one parameter is not only conducive to saving storage space, but also to improving the storage performance of the database system. In addition, because at least one parameter needs to be read during query, storing at least one parameter in a row storage manner not only does not waste query resources, but also helps to quickly locate at least one parameter during query. Column storage is conducive to quickly finding the index of the corresponding attribute to be queried. Therefore, the present application scheme adopts a mixed row and column storage method to inject time series data, which not only improves the storage performance of the database, but also improves the efficiency of time series data query.

[0204] Optionally, the device 60 further includes: a determining unit 603, configured to determine a row identifier according to the first parameter group, and generate a data row according to the row identifier, at least one indicator and a timestamp.

[0205] Processing unit 602 is used to store at least one indicator contained in the data row into at least one column storage unit CU, at least one indicator corresponds to at least one CU one by one, and each CU in at least one CU contains a row identifier; store the timestamp into a CU, and the CU storing the timestamp contains a row identifier.

[0206] Optionally, the processing unit 602 is configured to store the first parameter group and the row identifier in a row storage format.

[0207] Optionally, the determination unit 603 is used to: determine an index value corresponding to at least one parameter in the first parameter group; query a global cache to obtain a row identifier corresponding to the index value, and the global cache stores a correspondence between the index value and the row identifier.

[0208] Optionally, the determination unit 603 is used to: determine index values ​​corresponding to multiple parameters in the first parameter group; assign row identifiers to the index values, and store the correspondence between the index values ​​and the row identifiers in the global cache.

[0209] Optionally, at least one CU storing at least one indicator and the CU storing a timestamp are located in a partition corresponding to a first time range, and the time indicated by the timestamp is within the first time range.

[0210] Optionally, the partition includes multiple data sets, wherein each data set corresponds to a time range that does not completely overlap, at least one CU and a CU storing a timestamp are located in the same data set, and the time indicated by the timestamp is within the time range corresponding to the same data set.

[0211] Optionally, the processing unit 602 is further used to merge at least two data sets from the multiple data sets to obtain a merged data set, wherein the second time range corresponding to the merged data set includes the time ranges corresponding to at least two data sets, and the second time range is included in the first time range; and write the merged data set into a data storage device.

[0212] Optionally, the processing unit 602 is further configured to compress data in multiple CUs in the merged data set, and compress multiple CUs to obtain a compressed merged data set; and write the compressed merged data set into a data storage device.

[0213] The above-mentioned contents related to the device 60 for injecting time series data can be understood by referring to the contents related to the embodiment of the method for injecting time series data, which will not be repeated here.

[0214] like Fig.11 As shown, an embodiment of the device 70 for querying time series data provided in an embodiment of the present application includes:

[0215] The receiving unit 701 is used to receive a query for time series data, where the query includes at least one parameter that identifies a data source that generates the time series data, and at least one column identifier, where the at least one column identifier indicates at least one target column, and the at least one target column contains an indicator representing at least one attribute of the data source.

[0216] The first processing unit 702 is configured to determine a row identifier corresponding to the at least one parameter received by the receiving unit 701 .

[0217] The second processing unit 703 is configured to determine a plurality of CUs according to the row identifier determined by the first processing unit 702, wherein each CU in the plurality of CUs includes a row identifier.

[0218] The third processing unit 704 is used to determine at least one target CU from the multiple CUs determined by the second processing unit 703 according to at least one column identifier, each target CU corresponds to a target column; and generate a query result according to the at least one target CU.

[0219] The solution provided by the embodiment of the present application can determine the corresponding row identifier by identifying at least one parameter of the data source that generates the time series data, and then find multiple CUs corresponding to the row identifier in the column storage, and then determine the target CU from multiple CUs based on the column identifier. It can be seen that by identifying at least one parameter of the data source that generates the time series data and the column identifier, the column to be searched can be quickly queried, thereby improving the efficiency of time series data query.

[0220] Optionally, the third processing unit 704 is configured to generate a query result according to an index of at least one attribute in at least one target CU and at least one parameter.

[0221] Optionally, the query also includes query time, and the first processing unit 702 is further used to: determine a target partition corresponding to a first time range including the query time, the target partition including multiple merged data sets; determine a first merged data set from multiple merged data sets, the second time range corresponding to the first merged data set includes the query time, and the first merged data set includes multiple CUs; ​​determine multiple first CUs from the multiple CUs included in the first merged data set, the third time range corresponding to the first CU includes the query time, the first time range includes the second time range, and the second time range includes the third time range.

[0222] The second processing unit 703 is configured to determine a plurality of CUs corresponding to the row identifiers from the plurality of first CUs.

[0223] Optionally, the first processing unit 702 is configured to determine an index value corresponding to at least one parameter; query a global cache to obtain a row identifier corresponding to the index value, and the global cache stores a correspondence between the index value and the row identifier.

[0224] As mentioned above, the relevant contents of the device 70 for querying time series data can be understood by referring to the relevant contents of the aforementioned method embodiment for querying time series data, which will not be repeated here.

[0225] Fig.12As shown, a possible logical structure diagram of the computer device 80 involved in the above-mentioned embodiments provided in the embodiments of the present application, the computer device 80 may be a device for injecting time series data or a device for querying time series data. The computer device 80 includes: a processor 801, a communication interface 802, a memory 803 and a bus 804. The processor 801, the communication interface 802 and the memory 803 are interconnected via the bus 804. In the embodiments of the present application, the processor 801 is used to control and manage the actions of the device for injecting time series data or the device for querying time series data 80. For example, the processor 801 is used to execute Figures 3 to 9C The steps related to determination in the embodiment of the present invention are as follows: steps 202 to 203, steps 301 to 305, steps 401 to 406, and steps 502 to 504. The communication interface 802 is used to support the computer device 80 to communicate, for example: the communication interface 802 can perform the steps related to receiving or sending in the embodiment of the above method. The memory 803 is used to store the program code and data of the database server.

[0226] Among them, the processor 801 can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a transistor logic device, a hardware component or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the contents disclosed in this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. The bus 804 can be a peripheral component interconnect standard (Peripheral Component Interconnect, PCI) bus or an extended industry standard architecture (Extended Industry Standard Architecture, EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.12 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0227] See also Fig.13 The embodiment of the present application also provides a distributed database system, including: a hardware layer 1007 and a virtual machine monitor (VMM) 1001 running on the hardware layer 1007, and multiple virtual machines 1002. A virtual machine can be used as a data node of the distributed database system. Optionally, a virtual machine can also be designated as a coordination node.

[0228] Specifically, virtual machine 1002 is a virtual computer simulated on public hardware resources by virtual machine software, on which operating system and application programs can be installed, and network resources can be accessed by the virtual machine. For application programs running in the virtual machine, the virtual machine is like working in a real computer.

[0229] Hardware layer 1007: The hardware platform on which the virtualized environment runs can be abstracted from the hardware resources of one or more physical hosts. The hardware layer may include a variety of hardware, such as a processor 1004 (e.g., CPU) and a memory 1005, and may also include a network card 1003 (e.g., an RDMA network card), high-speed / low-speed input / output (I / O, Input / Output) devices, and other devices with specific processing functions.

[0230] The virtual machine 1002 runs the executable program based on the VMM and the hardware resources provided by the hardware layer 1007 to achieve the above Figures 3 to 9C Part or all of the functions of the device for injecting time series data or the device for querying time series data in the relevant embodiments will not be described in detail for the sake of brevity.

[0231] Furthermore, the distributed database system may also include a host machine (Host): as a management layer, it is used to complete the management and allocation of hardware resources; present a virtual hardware platform to the virtual machine; and implement the scheduling and isolation of the virtual machine. Among them, the Host may be a virtual machine monitor (VMM); it may also be a combination of a VMM and a privileged virtual machine. Among them, the virtual hardware platform provides various hardware resources to each virtual machine running on it, such as providing a virtual processor (such as VCPU), virtual memory, virtual disk, virtual network card, etc. Among them, the virtual disk may correspond to a file or a logical block device of the Host. The virtual machine runs on the virtual hardware platform prepared for it by the Host, and one or more virtual machines run on the Host. The VCPU of the virtual machine 1002 implements or executes the method steps described in the above-mentioned method embodiments of the present invention by executing the executable program stored in its corresponding virtual memory. For example, to implement the above Figures 3 to 9C In the related embodiments, the device for injecting time series data or the device for querying time series data may have partial or complete functions.

[0232] In another embodiment of the present application, a computer-readable storage medium is further provided, wherein the computer-readable storage medium stores computer-executable instructions. When at least one processor of the device executes the computer-executable instructions, the device executes the above Figures 3 to 9C Some embodiments describe a method for injecting time series data or a method for querying time series data.

[0233] In another embodiment of the present application, a computer program product is also provided. The computer program product includes computer-executable instructions, which are stored in a computer-readable storage medium; at least one processor of the device can read the computer-executable instructions from the computer-readable storage medium, and at least one processor executes the computer-executable instructions so that the device performs the above Figures 3 to 9C Some embodiments describe a method for injecting time series data or a method for querying time series data.

[0234] In another embodiment of the present application, a chip system is further provided. The chip system includes a processor, which is used to support the device for injecting time series data or the device for querying time series data to implement the above Figures 3 to 9C Some embodiments describe a method for managing transactions. In a possible design, the chip system may also include a memory, which is used to store necessary program instructions and data for the device for injecting timing data or the device for querying timing data. The chip system may be composed of a chip or may include a chip and other discrete devices.

[0235] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present application.

[0236] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0237] In the several embodiments provided in the embodiments of the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0238] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0239] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0240] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program codes.

[0241] The above are only specific implementations of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or replacements within the technical scope disclosed in the embodiments of the present application, which should be included in the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application shall be based on the protection scope of the claims.

Claims

1. A method for injecting time series data, It is characterized in that include: Receive time series data, the time series data including at least one parameter for identifying a data source generating the time series data, and an indicator representing multiple attributes of the data source and a timestamp, the timestamp indicating a time when the indicator was generated; storing a first parameter group in a row storage format, wherein the first parameter group includes at least one parameter for identifying a data source generating the time series data; The indicators representing the multiple attributes of the data source and the timestamp included in the second parameter group are stored in different column storage units CU in a column storage format.

2. The method according to claim 1, It is characterized in that The method further comprises: determining a row identifier according to the first parameter group, and generating a data row according to the row identifier, the multiple indicators and the timestamp; The storing the indicators representing the multiple attributes of the data source and the timestamp included in the second parameter group in different CUs in a column storage format includes: Storing the multiple indicators included in the data row into multiple CUs, where the multiple indicators correspond to the multiple CUs one by one, and each CU in the multiple CUs includes the row identifier; The timestamp is stored in a CU, and the CU storing the timestamp includes the row identifier.

3. The method according to claim 2, It is characterized in that The storing the first parameter group in a row storage format includes: The first parameter group and the row identifier are stored in a row storage format.

4. The method according to claim 2 or 3, It is characterized in that The determining the row identifier according to the first parameter group includes: Determine an index value corresponding to the at least one parameter in the first parameter group; A global cache is queried to obtain a row identifier corresponding to the index value, wherein the global cache stores a correspondence between the index value and the row identifier.

5. The method according to claim 2 or 3, It is characterized in that The determining the row identifier according to the first parameter group includes: Determine index values ​​corresponding to the multiple parameters in the first parameter group; The row identifier is allocated to the index value, and a corresponding relationship between the index value and the row identifier is stored in a global cache.

6. The method according to claim 2 or 3, It is characterized in that The method further comprises: The multiple CUs storing the multiple indicators and the CU storing the timestamp are located in a partition corresponding to a first time range, and the time indicated by the timestamp is within the first time range.

7. The method according to claim 6, It is characterized in that The partition includes multiple data sets, wherein each data set corresponds to a time range that does not completely overlap, the multiple CUs and the CU storing the timestamp are located in the same data set, and the time indicated by the timestamp is within the time range corresponding to the same data set.

8. The method according to claim 7, It is characterized in that The method further comprises: Merging at least two of the multiple data sets to obtain a merged data set, wherein a second time range corresponding to the merged data set includes the time ranges corresponding to the at least two data sets, and the second time range is included in the first time range; The merged data set is written to a data storage.

9. The method according to claim 8, It is characterized in that The method further comprises: Compressing data in multiple CUs in the merged data set, and compressing the multiple CUs to obtain a compressed merged data set; The writing the merged data set into a data storage device includes: writing the compressed merged data set into the data storage device.

10. A method for querying time series data, It is characterized in that include: Receive a query for time series data, the query including at least one parameter identifying a data source generating the time series data, and at least one column identifier, the at least one column identifier indicating at least one target column, the at least one target column including an indicator representing at least one attribute among a plurality of attributes of the data source; Determining a row identifier corresponding to the at least one parameter; Determine a plurality of CUs according to the row identifier, each CU in the plurality of CUs including the row identifier; Determine at least one target CU from the multiple CUs according to the at least one column identifier, each target CU corresponding to a target column; Generate a query result according to the at least one target CU.

11. The method according to claim 10, It is characterized in that The generating a query result according to the at least one target CU includes: The query result is generated according to the indicator of the at least one attribute in the at least one target CU and the at least one parameter.

12. The method according to claim 10 or 11, It is characterized in that The query further includes query time, and the method further includes: Determine a target partition corresponding to a first time range including the query time, wherein the target partition includes a plurality of merged data sets; Determine a first merged data set from the multiple merged data sets, the second time range corresponding to the first merged data set includes the query time, and the first merged data set includes multiple CUs; Determine a plurality of first CUs from the plurality of CUs included in the first merged data set, wherein a third time range corresponding to the first CU includes the query time, the first time range includes the second time range, and the second time range includes the third time range; The determining a plurality of CUs according to the row identifier includes: determining the plurality of CUs corresponding to the row identifier from the plurality of first CUs.

13. The method according to claim 10 or 11, It is characterized in that The determining a row identifier corresponding to the at least one parameter includes: Determining an index value corresponding to the at least one parameter; A global cache is queried to obtain a row identifier corresponding to the index value, wherein the global cache stores a correspondence between the index value and the row identifier.

14. A database system, It is characterized in that It includes: a coordination node and a data node connected to the coordination node for communication, The coordination node is used to receive time series data from the client; The data node is used to: Acquire the time series data from the coordination node, where the time series data includes at least one parameter for identifying a data source generating the time series data, and an index and a timestamp representing multiple attributes of the data source, where the timestamp indicates a time when the index is generated; storing a first parameter group in a row storage format, wherein the first parameter group includes at least one parameter for identifying a data source generating the time series data; The indicators representing the multiple attributes of the data source and the timestamp included in the second parameter group are stored in different CUs in a column storage format.

15. The database system according to claim 14, It is characterized in that The data nodes are also used to: determining a row identifier according to the first parameter group, and generating a data row according to the row identifier, the multiple indicators and the timestamp; Storing the multiple indicators included in the data row into multiple CUs, where the multiple indicators correspond to the multiple CUs one by one, and each CU in the multiple CUs includes the row identifier; The timestamp is stored in a CU, and the CU storing the timestamp includes the row identifier.

16. The database system according to claim 15, It is characterized in that The data node is used to store the first parameter group and the row identifier in a row storage format.

17. The database system according to claim 15 or 16, It is characterized in that The data node is used to: Determine an index value corresponding to the at least one parameter in the first parameter group; A global cache is queried to obtain a row identifier corresponding to the index value, wherein the global cache stores a correspondence between the index value and the row identifier.

18. The database system according to claim 15 or 16, It is characterized in that The data node is used to: Determine index values ​​corresponding to the multiple parameters in the first parameter group; The row identifier is allocated to the index value, and a corresponding relationship between the index value and the row identifier is stored in a global cache.

19. The database system according to claim 15 or 16, It is characterized in that The multiple CUs storing the multiple indicators and the CU storing the timestamp are located in a partition corresponding to a first time range, and the time indicated by the timestamp is within the first time range.

20. The database system according to claim 19, It is characterized in that The partition includes multiple data sets, wherein each data set corresponds to a time range that does not completely overlap, the multiple CUs and the CU storing the timestamp are located in the same data set, and the time indicated by the timestamp is within the time range corresponding to the same data set.

21. The database system according to claim 20, It is characterized in that The data nodes are also used to: Merging at least two of the multiple data sets to obtain a merged data set, wherein a second time range corresponding to the merged data set includes the time ranges corresponding to the at least two data sets, and the second time range is included in the first time range; The merged data set is written to a data storage.

22. The database system according to claim 21, It is characterized in that The data nodes are also used to: Compressing data in multiple CUs in the merged data set, and compressing the multiple CUs to obtain a compressed merged data set; The compressed merged data set is written into the data storage.

23. A database system, It is characterized in that include: a coordinating node and a data node in communication connection with the coordinating node, The coordination node is used to receive queries for time series data from the client; The data node is used to: Acquire the query for the time series data from the coordination node, the query including at least one parameter identifying a data source generating the time series data, and at least one column identifier, the at least one column identifier indicating at least one target column, the at least one target column including an indicator representing at least one attribute among multiple attributes of the data source; Determining a row identifier corresponding to the at least one parameter; Determine a plurality of CUs according to the row identifier, each CU in the plurality of CUs including the row identifier; Determine at least one target CU from the multiple CUs according to the at least one column identifier, each target CU corresponding to a target column; Generate a query result according to the at least one target CU.

24. The database system according to claim 23, It is characterized in that The data node is used to generate the query result according to the index of the at least one attribute in the at least one target CU and the at least one parameter.

25. The database system according to claim 23 or 24, It is characterized in that The query also includes query time, The data nodes are also used to: Determine a target partition corresponding to a first time range including the query time, wherein the target partition includes a plurality of merged data sets; Determine a first merged data set from the multiple merged data sets, the second time range corresponding to the first merged data set includes the query time, and the first merged data set includes multiple CUs; Determine a plurality of first CUs from the plurality of CUs included in the first merged data set, wherein a third time range corresponding to the first CUs includes the query time, the first time range includes the second time range, and the second time range includes the third time range; The plurality of CUs corresponding to the row identifiers are determined from the plurality of first CUs.

26. The database system according to claim 23 or 24, It is characterized in that The data node is used to: Determining an index value corresponding to the at least one parameter; A global cache is queried to obtain a row identifier corresponding to the index value, wherein the global cache stores a correspondence between the index value and the row identifier.

27. A device for injecting time series data, It is characterized in that The device includes at least one processor, a memory, and instructions stored in the memory and executable by the at least one processor, wherein the at least one processor executes the instructions to implement the steps of the method described in any one of claims 1 to 9.

28. A device for querying time series data, It is characterized in that The device includes at least one processor, a memory, and instructions stored in the memory and executable by the at least one processor, wherein the at least one processor executes the instructions to implement the steps of the method described in any one of claims 10 to 13.

29. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

30. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the method according to any one of claims 10 to 13 is implemented.

Citation Information

Patent Citations

  • Data storage method and apparatus

    CN106326220A