Data downsampling and data querying methods, systems, and storage media
By pre-downsampling the original data in the timing database and storing the downsampling data, the problems of low data query efficiency and high resource consumption in the prior art are solved, and the latest downsampling data are efficiently queried.
Patent Information
- Application Number
- CN202111501316.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-09
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-09
AI Technical Summary
The prior art performs real-time downsampling in a timing database, resulting in low data query efficiency, high resource consumption, and the inability to obtain the latest downsampling data in real time.
During the process of writing raw data from memory to the persistent storage medium, the data is pre-downsampled according to the preset downsampling rules, and the downsampled data is stored in the second persistent storage medium so that the pre-downsampling result is directly queried during query.
It improves data query efficiency, reduces resource consumption, and ensures that the latest downsampling data can be queried, solving the problems of low downsampling query efficiency and high resource consumption in the prior art.
Smart Images

Figure CN114328601B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a data downsampling and data query method, system, and storage medium. Background Art
[0002] Time series data is a series of data continuously generated based on a certain frequency. There is a large amount of time series data in fields such as Application Performance Monitor (APM), Internet of Things, and Industrial Internet. A time series database is designed to efficiently store and query such time series data. One type of requirement in a time series database is to perform downsampling on the original data.
[0003] In the prior art, real-time downsampling is generally performed during data query. This downsampling method needs to scan the original data from the disk file corresponding to the time series database. For queries with a relatively large time span, a large amount of original data needs to be scanned, and the data query efficiency is low. Summary of the Invention
[0004] Multiple aspects of this application provide a data downsampling and data query method, system, and storage medium to improve data query efficiency.
[0005] An embodiment of this application provides a data downsampling method, including:
[0006] Writing the obtained original data into the memory;
[0007] When the original data in the memory reaches a set data volume, writing the original data in the memory into a first persistent storage medium;
[0008] During the process of writing the original data into the first persistent storage medium, performing downsampling on the target original data written into the first persistent storage medium according to a preset downsampling rule to obtain downsampled data;
[0009] Writing the downsampled data into a second persistent storage medium.
[0010] An embodiment of this application also provides a data query method, including:
[0011] Obtaining a query request; the query request is used for aggregated query;
[0012] According to the query request, querying the memory and the persistent storage medium storing the downsampled data;
[0013] In the case where there is data in the memory that meets the query request, respectively obtaining first original data and first downsampled data that meet the query request from the memory and the persistent storage medium;
[0014] Downsample the first original data according to the query request to obtain second downsampled data;
[0015] Determine the query result of the query request based on the first downsampled data and the second downsampled data.
[0016] An embodiment of the present application also provides a computing system, including: a memory and a processor; the memory includes: a memory and a persistent storage medium;
[0017] The processor is communicatively connected to the memory and the persistent storage medium and is configured to execute the steps in the above data downsampling method and / or the above data query method.
[0018] An embodiment of the present application also provides a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to execute the steps in the above data downsampling method and / or the above data query method.
[0019] In an embodiment of the present application, during the process of writing the original data from the memory to the persistent storage medium, according to a preset downsampling rule, downsample the target original data written to the persistent storage medium; and store the downsampled data obtained by the downsampling process, thereby realizing pre-downsampling of the original data. In this way, during downsampling query, the pre-downsampled result can be directly queried without performing real-time downsampling on the original data during downsampling query, which helps improve the subsequent downsampling query efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0021] Figure 1a is a flowchart of the data downsampling method provided by an embodiment of the present application;
[0022] Figure 1b is a schematic diagram of the data downsampling process provided by an embodiment of the present application;
[0023] Figure 2 is a schematic diagram of the field structure provided by an embodiment of the present application;
[0024] Figure 3 is a flowchart of the data query method provided by an embodiment of the present application;
[0025] Figure 4 is a schematic diagram of the data query process provided by an embodiment of the present application;
[0026] Figure 5 Schematic diagram of the downsampling file merging process provided by the embodiment of the present application;
[0027] Figure 6 Schematic diagram of the structure of the computing system provided by the embodiment of the present application. Detailed implementation manners
[0028] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part rather than all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0029] In practical applications, due to the large amount of original data, users have the need for downsampling queries when querying data. The temperature sensor reports the temperature at a frequency of once per minute. When querying data, it is necessary to query the average temperature per hour in the past 7 days. This scenario requires downsampling the original temperature data per minute into the average temperature data per hour. In some solutions, downsampling is performed in real time during data query. This downsampling method needs to scan the original data from the disk file corresponding to the original data. For queries with a relatively large time span, a large amount of original data needs to be scanned, resulting in low data query efficiency; moreover, querying a large amount of original data consumes a large amount of memory resources, and real-time downsampling calculation of a large amount of original data also consumes a large amount of CPU resources.
[0030] In other solutions, downsampling is periodically performed through the Continuous Queries (CQ) method. This downsampling method has the following defects: (1) High resource consumption. Each time the CQ downsampling is executed, a large number of indexes need to be queried, including the forward index and the inverted index, consuming a large amount of memory resources and CPU resources; (2) When querying data, the latest downsampled data cannot be queried. Since the CQ downsampling is executed periodically and not in real time, the newly written original data on the disk cannot be immediately downsampled to obtain the latest downsampled data, resulting in the inability to query the most recent downsampled data when querying data; (3) Since the original data and the downsampled data are stored in different data tables, deleting the original data cannot synchronously delete the downsampled data, resulting in out-of-sync between the original data and the downsampled data.
[0031] In view of the technical problem that downsampling during real-time query leads to low data query efficiency, in some embodiments of the present application, during the process of writing original data from memory to a persistent storage medium, according to a preset downsampling rule, downsampling processing is performed on the target original data written to the persistent storage medium; and the downsampled data obtained by the downsampling processing is stored, realizing pre-downsampling of the original data. In this way, during downsampling query, the pre-downsampled result can be directly queried, without performing real-time downsampling processing on the original data during downsampling query, which helps to improve the subsequent downsampling query efficiency.
[0032] The following will describe in detail the technical solutions provided by each embodiment of the present application with reference to the accompanying drawings.
[0033] It should be noted that the same reference numerals represent the same object in the following drawings and embodiments. Therefore, once an object is defined in one drawing or embodiment, it does not need to be further discussed in the subsequent drawings and embodiments.
[0034] Figure 1a It is a schematic flowchart of the data downsampling method provided by an embodiment of the present application. As Figure 1a shown, the method includes:
[0035] 101. Write the obtained original data into memory.
[0036] 102. When the original data in memory reaches a set data volume, write the original data in memory to a first persistent storage medium.
[0037] 103. During the process of writing the original data to the first persistent storage medium, according to a preset downsampling rule, perform downsampling processing on the target original data written to the first persistent storage medium to obtain downsampled data.
[0038] 104. Write the downsampled data to a second persistent storage medium.
[0039] In an embodiment of the present application, the original data can be time-series data, that is, a series of data continuously generated based on a certain frequency. For a physical machine, original data can be obtained. In an embodiment of the present application, the physical machine can be a terminal device such as a computer, or a single server device, or a cloud-based server array. In addition, the physical machine can also refer to other computing devices with corresponding service capabilities, such as a computer and other terminal devices (running service programs), etc.
[0040] In this embodiment, the physical machine can provide data management services. Optionally, the physical machine can provide data storage, data processing, and data query services, etc. In some embodiments, the physical machine can maintain a database. In this embodiment, the database can be a time-series database for storing time-series data and providing time-series data query services.
[0041] In step 101 of this embodiment, the obtained original data can be written into the memory of the physical machine. Specifically, the original data can be written into the MenStore space of the memory. Since the storage space of the memory is limited, when the amount of data stored in the memory reaches the set amount of data, it is necessary to write the data stored in the memory into a persistent storage medium for preservation. Accordingly, as Figure 1a shown in step 102 and Figure 1b shown, when the original data in the memory reaches the set amount of data, the original data in the memory can be written into a persistent storage medium. In the embodiments of the present application, the persistent storage medium mainly refers to a non-volatile storage medium, such as a disk, a floppy disk, a hard disk, a digital versatile disc (DVD) or other optical storage, a magnetic tape or a compact disc read-only memory (CD-ROM), etc.
[0042] In the embodiments of the present application, the persistent storage medium can be deployed on the same physical machine as the memory, or on a different physical machine from the memory. For an embodiment where the storage system mounted on the physical machine is a centralized storage system, the persistent storage medium and the memory belong to the same physical machine; for an embodiment where the storage system mounted on the physical machine is a distributed storage system, the persistent storage medium and the memory can belong to the same physical machine or to different physical machines.
[0043] In this embodiment, in order to improve the data query efficiency, the original data can be pre-downsampled. In this way, when performing a downsampling query, the downsampled data can be directly queried without having to perform downsampling on the original data during the data query process, which can effectively improve the data query efficiency. Based on this, in this embodiment, in order to implement the pre-downsampling of the original data, as Figure 1a shown in step 103 and Figure 1b shown, during the process of writing the original data in the memory into the persistent storage medium, according to the preset downsampling rule, the target original data written into the persistent storage medium can be downsampled to obtain downsampled data.
[0044] In the embodiments of the present application, the specific implementation manner of obtaining the downsampling rule is not limited. In some embodiments, the downsampling rule can be set independently by the user or provider of the original data, etc. Optionally, the storage system can provide an interactive interface for users to access; the user (such as the user or provider of the original data) can set the downsampling rule independently through this interactive interface. General downsampling rules may include: sampling time interval and aggregation operator. Among them, the sampling time interval mainly refers to the time interval at which the original data is downsampled. The aggregation operator refers to the downsampling method used for the original data within the sampling time interval. Among them, the aggregation operator can be an index aggregation operator, a bucket aggregation operator, a matrix aggregation operator, a pipeline aggregation operator, etc. The index aggregation operator may include: maximum value (max), minimum value (min), sum, average value (avg), value statistics, distinct aggregation, percentage statistics, percentage ranking aggregation, and so on.
[0045] For example, the following statement can be used to express the downsampling rule:
[0046]
[0047] The above downsampling rule means that the sum is calculated for the original data in the database "db" at sampling time intervals of 5s (5 seconds) and 5min (5 minutes) respectively.
[0048] Based on the preset downsampling rule, step 103 can be implemented as follows: obtain the sampling time interval and the aggregation operator from the preset downsampling rule; for the target original data currently written to the persistent storage medium, obtain the target original data within each sampling time interval; and perform aggregation processing on the target original data within each sampling time interval according to the aggregation operator in the downsampling rule to obtain the downsampled data within that sampling time interval.
[0049] In practical applications, data is often stored in data tables. A data table may include: fields. A field may include: a field name and a field value. The corresponding field value can be indexed by the field name. In some embodiments, the field values with the same field name can be stored by column or by row; in this way, all field values of this field can be indexed by the field name. For example, as Figure 2 shown, Temperature can be the field name; Timestamp and Value can be the field values corresponding to the field name Temperature.
[0050] Considering that the data object attributes corresponding to different field names are different, during downsampling, the original data with the same attributes can be aggregated; the original data with different attributes cannot be aggregated. For example, when detecting a certain physical space, temperature time series data, humidity time series data, air pollution index, etc. are obtained. Since temperature and humidity are attributes of different dimensions, it is meaningless to aggregate the temperature time series data and the humidity time series data. Based on this, in this embodiment, when downsampling the target original data written to the persistent storage medium, the target original data can be divided into at least one data unit according to the field name of the target original data. Optionally, according to the field name of the target original data, the field values corresponding to the same field name in the target original data can be divided into one data unit to obtain at least one data unit. Correspondingly, one data unit can be one field. In the embodiments of the present application, the specific number of data units can be determined by the number of field names included in the target original data.
[0051] Further, according to the preset downsampling rule, at least one data unit can be downsampled respectively to obtain the downsampled data corresponding to each data unit, and then the downsampled data corresponding to the target original data can be obtained.
[0052] Optionally, based on the above preset downsampling rule, the sampling time interval and the aggregation operator can be obtained from the preset downsampling rule; for any data unit A, the target original data within each sampling time interval can be obtained from the data unit A. Specifically, for any data unit A, the target original data within each sampling time interval can be obtained according to the timestamp information in the data unit A. Further, the target original data within each sampling time interval can be aggregated according to the aggregation operator to obtain the downsampled data corresponding to the data unit A.
[0053] After obtaining the downsampled data corresponding to the target original data, in step 104, the downsampled data can also be written to the persistent storage medium for storage. In the embodiments of the present application, for the convenience of description and distinction, the persistent storage medium for storing the original data is defined as the first persistent storage medium; the persistent storage medium for storing the downsampled data is defined as the second persistent storage medium.
[0054] Among them, the first persistent storage medium and the second persistent storage medium may be the same storage medium or different persistent storage media. In the case where the first persistent storage medium and the second persistent storage medium are different persistent storage media, the first persistent storage medium and the second persistent storage medium may be mounted on the same physical machine or on different physical machines. The number of the first persistent storage medium and the second persistent storage medium may each be one or more. "Multiple" means two or more. Multiple first persistent storage media may be mounted on the same physical machine or on different physical machines. Of course, multiple second persistent storage media may also be mounted on different physical machines.
[0055] In this embodiment, during the process of writing the original data from the memory to the persistent storage medium, according to the preset downsampling rule, the target original data written to the persistent storage medium is downsampled; and the downsampled data obtained by the downsampling process is stored, realizing the pre-downsampling of the original data. In this way, during the downsampling query, the pre-downsampling result can be directly queried without performing real-time downsampling on the original data during the downsampling query, which helps to improve the subsequent downsampling query efficiency.
[0056] On the other hand, the data downsampling provided in the embodiment is performed during the memory flush (MemStore Flush) phase, that is, during the process of writing the data in the memory to the first persistent storage medium, the target original data written to the first persistent storage medium is downsampled. Compared with CQ downsampling, it does not need to query the inverted data and forward index of the original data to obtain the original data, which can reduce the consumption of memory and CPU resources.
[0057] For the downsampling query, in the embodiment of the present application, the original data and downsampled data in the memory can be queried. On the one hand, the original data in the memory is downsampled in real time. For the downsampled data, the downsampled data that meets the query request can be directly obtained to obtain the data query result. Since the original data in the memory is the latest original data, plus the downsampled data query result, it can achieve the query of all downsampled data, solving the disadvantage that CQ downsampling cannot query the latest downsampled data. On the other hand, for directly querying the downsampled data part, no downsampling process is required during the data query, which helps to improve the data query efficiency compared with the real-time downsampling query.
[0058] The storage system maintained in the embodiment of the present application can not only provide downsampling queries but also non-downsampling queries. For non-downsampling query requests, the original data in the memory and the original data in the first persistent storage medium can be queried. This query process is the same as or similar to the data query of the existing storage system, which is not the focus of the present application. Therefore, below, taking the aggregation query (i.e., downsampling query) as an example, the data query method provided in the embodiment of the present application is exemplarily described.
[0059] Figure 3 This is a schematic flowchart of the data query method provided by the embodiment of the present application. As Figure 3 shown, the data query method includes:
[0060] 301. Obtain a query request; the query request is used for aggregated query.
[0061] 302. Query the memory and the second persistent storage medium according to the query request.
[0062] 303. For the case where there is data in the memory that meets the query request, obtain the first original data and the first downsampled data that meet the query request from the memory and the second persistent storage medium respectively.
[0063] 304. Downsample the first original data according to the query request to obtain the second downsampled data.
[0064] 305. Determine the query result of the query request based on the first downsampled data and the second downsampled data.
[0065] In the embodiment of the present application, the query request can be a non-aggregated query or an aggregated query. The embodiment of the present application focuses on taking the aggregated query as an example to exemplarily illustrate the data query method provided by the embodiment of the present application. Correspondingly, in step 301, a query request can be obtained, and the query request is used for aggregated query. The query request can include query conditions. The query conditions can include: data objects to be queried, aggregation operators, query time ranges, etc.
[0066] The original data in the memory is the latest written. Since the query time ranges and data objects queried by different query requests may be different, there may or may not be data in the memory that meets some query requests. For the storage system, it is impossible to determine in advance whether there is data in the memory that meets the query request. Therefore, in order to improve the timeliness and accuracy of data query and prevent missing the latest data, as Figure 3 shown in step 302 and Figure 4 shown, the memory and the second persistent storage medium can be queried according to the query request.
[0067] Optionally, the query request can be semantically parsed to obtain the query conditions of the query request. Optionally, the query request can be compiled into an Abstract Syntax Tree (AST), and the statements of the query request can be error-detected during this process to ensure that the input request statements have no syntax and lexical errors. For example, detecting whether there are keyword spelling errors, whether there are redundant punctuation marks, and whether the entire statement is legal, etc.
[0068] Further, the nodes of the above abstract syntax tree can be checked in sequence, and the metadata of relevant tables and the metadata of attributes can be attached to the syntax tree, and finally a syntax tree with semantics (bound AST) is generated. Further, the access requirement content of the query request can be obtained according to the syntax tree with semantics.
[0069] Further, an execution plan can be generated according to the query conditions. Optionally, the optimizer can generate a logical operator tree (LOT) according to the semantic syntax tree. Optionally, the nodes of the semantic syntax tree can be mapped to operator nodes to obtain a logical operator tree. Each node on the logical operator tree is called a logical operator. Further, the physical operator corresponding to each logical operator can be extended to obtain a physical execution tree. Further, the physical execution tree with the smallest cost can be selected from the physical execution trees as the execution plan. Among them, the smallest cost can be the shortest path, the smallest memory consumption, the smallest calculation amount, the shortest calculation time, etc.
[0070] Further, the memory and the second persistent storage medium can be queried according to the execution plan.
[0071] In this embodiment, for the embodiment where the data satisfying the query request does not exist in the memory, the downsampled data satisfying the query request can be obtained from the second persistent storage medium; and based on the downsampled data satisfying the query request obtained from the second persistent storage medium, the query result of the query request is determined. Due to this data query method, the downsampled data satisfying the query request can be directly obtained from the downsampled data, and there is no need to perform real-time downsampling on the original data during the data query process, which helps to improve the data query efficiency.
[0072] For the embodiment where the data satisfying the query request exists in the memory, in step 303, the original data (defined as the first original data) and the downsampled data satisfying the query request can be obtained from the memory and the second persistent storage medium respectively.
[0073] Further, in step 304, the original data satisfying the query request obtained from the memory can be downsampled according to the query request to obtain downsampled data. In the embodiments of the present application, for the sake of easy description and distinction, the downsampled data satisfying the query request obtained from the second persistent storage medium is defined as the first downsampled data; the downsampled data obtained by downsampling the original data satisfying the query request obtained from the memory is defined as the second downsampled data.
[0074] Optionally, the aggregation operator and sampling time interval included in the query request can be obtained from the query request. Further, according to the sampling time interval included in the query request, the original data corresponding to each sampling time interval can be obtained from the original data that meets the query request. Further, the original data corresponding to each sampling time interval can be aggregated according to the aggregation operator included in the query request to obtain the above-mentioned second downsampled data.
[0075] Next, in step 305, the query result corresponding to the query request can be determined based on the first downsampled data and the second downsampled data.
[0076] The data query method provided in this embodiment can query the original data and downsampled data in the memory. On the one hand, real-time downsampling is performed on the original data in the memory. For the downsampled data, the downsampled data that meets the query request can be directly obtained to obtain the data query result. Since the original data in the memory is the latest original data, combined with the downsampled data query result, full-scale downsampled data query can be achieved, which can improve the timeliness and accuracy of data query and solve the disadvantage that CQ downsampling cannot query the latest downsampled data. On the other hand, for directly querying the downsampled data part, no downsampling process is required during the data query process, which helps to improve the data query efficiency compared with real-time downsampling query.
[0077] Moreover, for the case where there is original data in the memory that meets the query request, since the memory space is small and the amount of original data stored is much smaller than the original data stored in the first persistent storage medium, the real-time downsampling of the original data in the memory can be completed quickly. Compared with the method of real-time downsampling query of all original data in the above-mentioned existing solutions, the data query method provided in this application embodiment still has a high data query efficiency.
[0078] In practical applications, the data storage method may affect the data query process. Therefore, the following combines the storage process of the downsampled data and the process of writing the downsampled data into the second persistent storage medium to exemplarily illustrate the specific implementation process of the downsampling query (aggregation query).
[0079] In the embodiments of the present application, the specific implementation form of writing the downsampled data into the second persistent storage medium is not limited. Considering that the downsampled data stored in the second persistent storage medium is generally obtained by downsampling according to different downsampling rules, in order to facilitate subsequent queries and improve the subsequent data query efficiency, in the embodiments of the present application, for the downsampled data corresponding to any one of the above data units A, the target field name (Field) for characterizing the downsampling rule and the downsampling object can be determined according to the downsampling rule corresponding to the data unit A and the field name of the data unit A. In the embodiments of the present application, the specific format of the target field name (Field) is not limited. In some embodiments, the format of the target field name can be expressed as: "{raw_field}_{aggregator}_{interval}". Wherein, "raw_field" represents the column field name, that is, the field name of the data unit, and can characterize the downsampling object. "Aggregator" represents the aggregation operator; "interval" represents the sampling time interval. For example, for the downsampling rule of performing max downsampling on the CPU at a sampling time interval of 30s, the downsampling rule can be determined to represent "performing max downsampling at a sampling time interval of 30s", and the downsampling object is the CPU field. Correspondingly, the target field name can be expressed as "cpu_max_30s".
[0080] Furthermore, the target field name can be used as the field name, and the downsampled data corresponding to any one of the data units A can be used as the field value of the target field name, and the target field name and the downsampled data corresponding to the data unit A are written into the second persistent storage medium. In this way, during the downsampling query, the target field name that meets the query conditions can be determined according to the query conditions in the downsampling query request; according to the target field name that meets the query conditions, the field value corresponding to the target field name is indexed as the downsampled data that meets the query conditions. In this downsampling query process, data query can be performed according to the target field name corresponding to the downsampled data, without querying all the downsampled data, which helps to improve the data query efficiency.
[0081] Specifically, based on the above-mentioned target field names, when querying the second persistent storage medium according to the query request, the query conditions corresponding to the query request can be obtained from the query request; and according to the query conditions, a first field name that meets the field name format of the downsampled data in the second persistent storage medium (i.e., the format of the above-mentioned target field names) can be generated. Optionally, the data object to be queried, the aggregation operator, the sampling time interval, etc. can be obtained from the query conditions; further, according to the above-mentioned target field name format, based on the data object to be queried, the aggregation operator, and the sampling time interval, it can be converted into a first field name with the format of the above-mentioned target field names. For example, for the query condition of querying the maximum value (max) of the CPU within every 30s, the data object to be queried is the CPU field; the aggregation operator is the max operator; the sampling time interval is 30s. Correspondingly, the first field name converted from this query condition is "cpu_max_30s".
[0082] Further, the second persistent storage medium can be queried according to the first field name to determine the downsampled data corresponding to the first field name. Further, the first downsampled data that meets the query conditions can be obtained from the downsampled data corresponding to the first field name.
[0083] In some embodiments, such as Figure 1b and Figure 4 shown, the original data and the downsampled data can be stored in the form of files. In the embodiments of the present application, a file refers to an encoding method of information used to store information, and the specific implementation form of the file is not limited. In some embodiments, the file can be a data table, etc. Among them, the storage file of the original data is defined as the original file; the storage file of the downsampled data is defined as the downsampled file. In the embodiments of the present application, every time the original data in the memory reaches the set data volume, an operation of writing the original data in the memory into the first persistent storage medium is started to form an original file; during the process of writing the original data into the first persistent storage medium each time, an operation of performing downsampling processing on the target original data written into the first persistent storage medium and writing the downsampled data into the second persistent storage medium is started to form a downsampled file.
[0084] In the embodiments of the present application, in order to reduce the storage space occupied by the downsampled files, a hierarchical organizational structure can be used to store the downsampled files. Each level is used to store a set threshold number of downsampled files. The set threshold corresponding to each level is represented by M. Among them, M≥2, and M is an integer. The thresholds corresponding to different levels can be the same or different. In the embodiments of the present application, in order to reduce the storage space occupied by the downsampled files, such as Figure 5As shown, for any two adjacent levels, when the number of downsampled files in the lower level reaches the threshold M corresponding to the lower level, the M downsampled files are merged; the merged downsampled files are stored in the level above the lower level. For example, Figure 5 The levels of the middle-level organizational results increase successively from L0 - L5. When the number of downsampled files in the L0 level reaches the set threshold M, the M downsampled files in the L0 level can be merged; and the merged downsampled files are stored in the L1 level; for the L1 level, when the number of downsampled files in this level reaches the set threshold N, the N downsampled files in the L1 level can be merged; and the merged downsampled files are stored in the L2 level, and so on. Among them, N ≥ 2, and N is an integer. N and M can be the same or different.
[0085] Considering that there may be downsampled data with overlapping time windows in the M downsampled files, in order to further reduce the storage space occupied by the downsampled data, for the case where the M downsampled files have overlapping time windows, the aggregation operator in the downsampling rule can be used to perform an aggregation operation on the downsampling results corresponding to the overlapping time windows; and the M downsampled files after aggregation are merged into one downsampled file. Then, the merged downsampled file is stored in the upper level. Since the duplicate downsampled data in the overlapping time windows is removed during the merging process of the downsampled files, the storage of the downsampled files in the sampling level organizational structure can reduce the storage space occupied by the downsampled data.
[0086] For the embodiment of downsampled data stored in the form of files, during the aggregation query, the first downsampled data obtained from the second persistent storage medium that meets the query request may be located in one downsampled file or may be located in multiple downsampled files. Multiple means 2 or more. In this embodiment, for the embodiment where the first downsampled data is located in multiple downsampled files, the time information of the downsampled data in the multiple downsampled files can be used to determine whether there are overlapping time windows in the downsampled data in the multiple downsampled files; if the determination result is yes, the aggregation operator in the query request can be used to perform an aggregation operation on the first downsampled data corresponding to the overlapping time windows to obtain the first downsampled data. Further, based on the aggregated first downsampled data and the second downsampled data, the query result corresponding to the query request can be determined.
[0087] In practical applications, for the original data written to the first persistent storage medium, data deletion may occur. In the embodiments of the present application, in order to synchronously delete the downsampled data and the original data, when data deletion occurs in the original data of the first persistent storage medium, the deleted original data can be marked to obtain a Tombstone record. The Tombstone record is used to record the information of the deleted original data. The original data recorded in the Tombstone record can be the original data logically deleted from the first persistent storage medium or the original data physically deleted in reality.
[0088] Further, the downsampled data corresponding to the Tombstone record can be determined according to the time information of the data in the Tombstone record and the time information of the downsampled data stored in the second persistent storage medium. Optionally, for the above embodiments in which the downsampled file is stored in the form of a downsampled file, the downsampled file corresponding to the Tombstone record can be determined according to the time information of the data in the Tombstone record and the time information of the data in the downsampled file stored in the second persistent storage medium. In order to keep the downsampled data and the original data synchronously deleted, during the merging process of the downsampled file corresponding to the Tombstone record, the downsampled data corresponding to the Tombstone record can be determined from the downsampled file corresponding to the Tombstone record. Optionally, the downsampled data whose time window overlaps with the time information of the data in the Tombstone record in the downsampled file corresponding to the Tombstone record can be determined according to the time information of the data in the Tombstone record and the time information of the downsampled data in the downsampled file corresponding to the Tombstone record, as the downsampled data corresponding to the Tombstone record. Further, during the merging process of the downsampled file corresponding to the Tombstone record, the downsampled data corresponding to the Tombstone record can be deleted, so that the merged downsampled file no longer contains the downsampled data corresponding to the deleted original data, realizing the synchronous deletion of the downsampled data and the original data, and solving the defect that the above CQ downsampling method cannot synchronously delete the downsampled data when the original data is deleted.
[0089] In order to prevent the downsampled data corresponding to the deleted original data from being queried and improve the data query accuracy, in this embodiment, based on the above Tombstone record, when determining the query result of the query request during the aggregation query process, the Tombstone record used to mark the deleted original data can be obtained; and according to the time information of the data in the Tombstone record and the time information of the data in the first downsampled data, it is judged whether the first downsampled data contains the downsampled data corresponding to the Tombstone record; if the judgment result is yes, the downsampled data corresponding to the Tombstone record can be deleted from the first downsampled data; and the second downsampled data and the first downsampled data after deleting the downsampled data corresponding to the Tombstone record are determined as the query result of the query request. In this way, it can be ensured that the downsampled data corresponding to the original data marked by the Tombstone record is not queried, which helps to improve the data query accuracy and solve the defect that the above CQ downsampling method cannot synchronously delete the downsampled data when the original data is deleted.
[0090] For the embodiment in which the above first downsampled data is located in multiple downsampled files and there is an overlapping time window for the downsampled data in the multiple downsampled files, when determining the query result corresponding to the query request based on the aggregated first downsampled data and the second downsampled data, it is also possible to determine whether the aggregated first downsampled data contains the downsampled data corresponding to the tombstone record according to the time information of the data in the tombstone record and the time information of the data in the aggregated first downsampled data; if the judgment result is yes, the downsampled data corresponding to the tombstone record can be deleted from the aggregated first downsampled data; and it is determined that the second downsampled data and the aggregated first downsampled data after deleting the downsampled data corresponding to the tombstone record are the query result corresponding to the query request.
[0091] Further, the query result can be returned to the provider of the query request. In the embodiment of the present application, for the aggregated query, the reason why the aggregated query can find the downsampled data that meets the aggregated query request in the downsampled data is mainly because the downsampling rule corresponding to the downsampled data can be set by the provider of the query request. For the provider of the query request, the downsampling rule can be set according to its own query requirements; and it is pre-stored in the module, device, equipment or system that executes the data downsampling method provided by the embodiment of the present application.
[0092] It should be noted that the execution subject of each step of the method provided in the above embodiment can be the same device, or the method can also be executed by different devices as the execution subject. For example, the execution subject of steps 301 and 302 can be device A; for another example, the execution subject of step 301 can be device A, and the execution subject of step 302 can be device B; and so on.
[0093] In addition, in some of the processes described in the above embodiments and the accompanying drawings, there are multiple operations that appear in a specific order, but it should be clearly understood that these operations can be executed not in the order in which they appear in this article or in parallel. The operation numbers such as 301, 302, etc. are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in sequence or in parallel.
[0094] Correspondingly, the embodiment of the present application also provides a computer-readable storage medium storing computer instructions, and when the computer instructions are executed by one or more processors, one or more processors are caused to execute the steps in the above data downsampling method and / or data query method.
[0095] The embodiments of the present application also provide a computer program product, which includes a computer program. When the computer program is executed by a processor, the processor is caused to execute the steps in the above data downsampling method and / or data query method. In the embodiments of the present application, the specific implementation form of the computer program product is not limited. In some embodiments, the computer program product can be implemented as a query engine, a data processing system for a database, or an executor in a query engine, etc.
[0096] Figure 6 It is a schematic structural diagram of the computing system provided by the embodiments of the present application. As Figure 6 shown, the computing system includes a memory 61 and a processor 62. Among them, the memory 61 may include a memory 61a and a persistent storage medium 61b.
[0097] In this embodiment, the memory 61 and the processor 62 may be located on the same physical machine or on different physical machines. The memory 61a and the persistent storage medium 61b may belong to the same physical machine or to different physical machines. Optionally, the memory 61a and the processor 62 belong to the same physical machine. The number of persistent storage media 61b may be one or more. More than one means two or more. The multiple persistent storage media 61b may belong to the same physical machine or to different physical machines.
[0098] In this embodiment, the memory 61a and the persistent storage medium 61b are communicatively connected to the processor 62. The processor 62 can be used to: write the acquired original data into the memory 61a; when the original data in the memory 61a reaches a set data volume, write the original data in the memory 61a into the first persistent storage medium 61b1 in the persistent storage medium 61b; during the process of writing the original data into the first persistent storage medium 61b1, perform downsampling processing on the target original data written into the first persistent storage medium 61b1 according to a preset downsampling rule to obtain downsampled data; and write the downsampled data into the second persistent storage medium 61b2.
[0099] In the embodiments of the present application, the first persistent storage medium 61b1 and the second persistent storage medium 61b2 may be the same storage medium or different storage media.
[0100] In some embodiments, when the processor 62 performs downsampling processing on the target original data written into the first persistent storage medium, it is specifically used to: divide the target original data into at least one data unit according to the field name of the target original data; and perform downsampling processing on each of the at least one data unit according to a preset downsampling rule to obtain downsampled data.
[0101] Optionally, when the processor 62 performs downsampling processing on at least one data unit respectively, it is specifically configured to: obtain a sampling time interval and an aggregation operator from a preset downsampling rule; for any data unit, obtain target original data within each sampling time interval from any data unit; and perform aggregation processing on the target original data within each sampling time interval according to the aggregation operator to obtain downsampled data corresponding to any data unit.
[0102] In some other embodiments, when the processor 62 writes the downsampling processing result into the second persistent storage medium 61b2, it is specifically configured to: for the downsampled data corresponding to any data unit, determine a target field name for characterizing the downsampling rule and the downsampling object according to the downsampling rule and the field name of any data unit; and write the target field name and the downsampled data corresponding to any data unit into the second persistent storage medium 61b2 with the target field name as the field name and the downsampled data of any data unit as the field value of the target field name.
[0103] In some embodiments, the processor 62 is further configured to: store the downsampling files corresponding to the downsampled data in a hierarchical organizational structure. Correspondingly, the processor 62 is further configured to: for any two adjacent levels, in the case where the number of downsampling files in the lower level reaches the threshold M corresponding to the lower level, perform merging processing on the M downsampling files; and store the merged downsampling files in the upper level of the lower level; where M is a set threshold, M≥2, and M is an integer.
[0104] Optionally, when the processor 62 performs merging processing on the M downsampling files, it is specifically configured to: for the case where the M downsampling files have overlapping time windows, perform an aggregation operation on the downsampling processing results corresponding to the overlapping time windows according to the aggregation operator in the downsampling rule; and merge the M aggregated downsampling files into one downsampling file.
[0105] In some embodiments, the processor 62 is further configured to: for the case where there is data deletion in the original data in the first persistent storage medium 61b1, mark the deleted original data to obtain a tombstone record; determine the downsampling file corresponding to the tombstone record according to the time information of the data in the tombstone record and the time information of the data in the downsampling file; determine the downsampled data corresponding to the tombstone record from the downsampling file corresponding to the tombstone record during the merging process of the downsampling files corresponding to the tombstone record; and delete the downsampled data corresponding to the tombstone record.
[0106] In the embodiments of the present application, as Figure 6As shown, the computing system may further include: a communication component 63. The processor 62 is further configured to: obtain a query request through the communication component 63; the query request is used for an aggregation query; according to the query request, query the memory 61a and the second persistent storage medium 61b2; in the case where there is data in the memory 61a that satisfies the query request, obtain the first original data and the first downsampled data that satisfy the query request from the memory and the second persistent storage medium 61b2 respectively; according to the query request, perform downsampling processing on the first original data to obtain the second downsampled data; and, based on the first downsampled data and the second downsampled data, determine the query result of the query request.
[0107] Optionally, when determining the query result of the query request, the processor 62 is specifically configured to: obtain a tombstone record of the original data used to mark deletion; according to the time information of the data in the tombstone record and the time information of the data in the first downsampled data, determine whether the first downsampled data contains the downsampled data corresponding to the tombstone record; if the determination result is yes, delete the downsampled data corresponding to the tombstone record from the first downsampled data; and determine the second downsampled data and the first downsampled data after deleting the downsampled data corresponding to the tombstone record as the query result of the query request.
[0108] Optionally, when querying the second persistent storage medium 61b2, the processor 62 is specifically configured to: obtain, from the query request, a query condition corresponding to the query request; according to the query condition, generate a first field name that satisfies the field name format corresponding to the downsampled data in the second persistent storage medium; according to the first field name, query the second persistent storage medium 61b2 to determine the downsampled data corresponding to the first field name; obtaining the first downsampled data that satisfies the query request from the second persistent storage medium includes: obtaining the first downsampled data that satisfies the query condition from the downsampled data corresponding to the first field name.
[0109] In some embodiments, the first downsampled data is located in multiple downsampled files. Accordingly, when determining the query result of the query request, the processor 62 is specifically configured to: in the case where there are overlapping time windows for the first downsampled data in different downsampled files, perform an aggregation operation on the first downsampled data corresponding to the overlapping time windows according to the aggregation operator in the query request to obtain the aggregated first downsampled data; based on the aggregated first downsampled data and the second downsampled data, determine the query result of the query request.
[0110] In some alternative embodiments, as Figure 6 shown, the computing system may further include: a power supply component 64 and other components. Figure 6 Only some components are schematically shown, and it does not mean that the computing system must include Figure 6 all the components shown, nor does it mean that the computing system can only includeFigure 6 The components shown
[0111] It should be noted that the components included in the computing system provided in the embodiments of the present application may belong to the same physical machine or different physical machines. For the case where the included components belong to different physical machines, the different physical machines are communicatively connected. The processor 62 can control and operate other components through the communication between the physical machines.
[0112] In the computing system provided in this embodiment, during the process of writing the original data from the memory to the persistent storage medium, according to the preset downsampling rule, the target original data written to the persistent storage medium is downsampled; and the downsampled data obtained by the downsampling process is stored, realizing the pre-downsampling of the original data. In this way, during the downsampling query, the pre-downsampling result can be directly queried, without the need to perform real-time downsampling processing on the original data during the downsampling query, which helps to improve the subsequent downsampling query efficiency.
[0113] On the other hand, the data downsampling provided in the embodiment is performed during the memory flush (MemStore Flush) phase, that is, during the process of writing the data in the memory to the first persistent storage medium, the target original data written to the first persistent storage medium is downsampled. Compared with CQ downsampling, it does not need to query the inverted data and forward index of the original data to obtain the original data, and can reduce the consumption of memory and CPU resources.
[0114] For the downsampling query, in the embodiments of the present application, the original data and downsampled data in the memory can be queried. On the one hand, the original data in the memory is downsampled in real time. For the downsampled data, the downsampled data that meets the query request can be directly obtained to obtain the data query result. Since the original data in the memory is the latest original data, plus the downsampled data query result, it can realize the full-scale downsampled data query, solving the disadvantage that CQ downsampling cannot query the latest downsampled data. On the other hand, for directly querying the downsampled data part, no downsampling processing is required during the data query process, which helps to improve the data query efficiency compared with the real-time downsampling query.
[0115] In the embodiments of the present application, the memory is used to store computer programs and can be configured to store various other data to support the operations on the device where it is located. Among them, the processor can execute the computer programs stored in the memory to implement the corresponding control logic. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc.
[0116] In the embodiments of the present application, the processor may be any hardware processing device capable of executing the above method logic. Optionally, the processor may be a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), or a Microcontroller Unit (MCU); it may also be a programmable device such as a Field-Programmable Gate Array (FPGA), a Programmable Array Logic (PAL), a General Array Logic (GAL), or a Complex Programmable Logic Device (CPLD); or it may be an Advanced RISC Machines (ARM) or a System on Chip (SOC), etc., but not limited thereto.
[0117] In the embodiments of the present application, the communication component is configured to facilitate communication between the device where it is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G, 5G, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component can also be implemented based on technologies such as Near Field Communication (NFC), Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wideband (UWB), Bluetooth (BT), or other technologies.
[0118] In the embodiments of the present application, the power supply component is configured to provide power to various components of the device where it is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.
[0119] It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are different types.
[0120] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0121] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a device for realizing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.
[0122] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that realizes the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.
[0123] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.
[0124] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0125] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0126] The storage medium of a computer is a readable storage medium, also known as a readable medium. A readable storage medium includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of a computer's storage medium include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory media such as modulated data signals and carrier waves.
[0127] It should also be noted that the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0128] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A data downsampling method, characterized in that, it includes: writing the acquired original data into memory for obtaining first original data that meets the query request corresponding to the aggregation query from the memory during the aggregation query, and performing downsampling processing on the first original data according to the query request to obtain second downsampled data; when the original data in the memory reaches a set data volume, writing the original data in the memory into a first persistent storage medium; during the process of writing the original data into the first persistent storage medium, performing downsampling processing on the target original data written into the first persistent storage medium according to a preset downsampling rule to obtain downsampled data; writing the downsampled data into a second persistent storage medium for obtaining first downsampled data that meets the query request from the second persistent storage medium during the aggregation query, and determining the query result of the query request based on the first downsampled data and the second downsampled data.
2. The method according to claim 1, characterized in that, the performing downsampling processing on the target original data written into the first persistent storage medium according to a preset downsampling rule includes: dividing the target original data into at least one data unit according to the field name of the target original data; performing downsampling processing on the at least one data unit respectively according to a preset downsampling rule to obtain the downsampled data.
3. The method according to claim 2, characterized in that, the performing downsampling processing on the at least one data unit respectively according to a preset downsampling rule includes: acquiring a sampling time interval and an aggregation operator from the preset downsampling rule; for any one data unit, acquiring the target original data within each sampling time interval from the any one data unit; performing aggregation processing on the target original data within each sampling time interval according to the aggregation operator to obtain the downsampled data corresponding to the any one data unit.
4. The method according to claim 2, characterized in that, the writing the downsampled data into the second persistent storage medium includes: for the downsampled data corresponding to any one data unit, determining a target field name for characterizing the downsampling rule and the downsampling object according to the downsampling rule and the field name of the any one data unit; using the target field name as the field name and the downsampled data of the any one data unit as the field value of the target field name, writing the target field name and the downsampled data corresponding to the any one data unit into the second persistent storage medium.
5. The method according to any one of claims 1-4, characterized in that, storing the downsampled files corresponding to the downsampled data in a hierarchical organizational structure; the method further includes: for any two adjacent levels, when the number of downsampled files in the lower level reaches the threshold M corresponding to the lower level, performing merging processing on the M downsampled files; Store the merged downsampled file in the upper level of the lower level; where M is a set threshold, M , and M is an integer.
6. The method according to claim 5, characterized in that, the performing merging processing on the M downsampled files includes: In the case where there are overlapping time windows for M downsampled files, perform an aggregation operation on the downsampling processing results corresponding to the overlapping time windows according to the aggregation operator in the downsampling rule; Merge the aggregated M downsampled files into one downsampled file.
7. The method according to claim 5, wherein, further comprising: In the case where there is data deletion in the original data in the first persistent storage medium, mark the deleted original data to obtain tombstone records; According to the time information of the data in the tombstone records and the time information of the data in the downsampled files, determine the downsampled file corresponding to the tombstone records; During the merging process of the downsampled files corresponding to the tombstone records, determine the downsampled data corresponding to the tombstone records from the downsampled files corresponding to the tombstone records; Delete the downsampled data corresponding to the tombstone records.
8. The method according to any one of claims 1-4, wherein, further comprising: Obtain a query request; The query request is used for aggregation query; According to the query request, query the memory and the second persistent storage medium; In the case where there is data in the memory that satisfies the query request, obtain first original data and first downsampled data that satisfy the query request from the memory and the second persistent storage medium respectively; According to the query request, perform downsampling processing on the first original data to obtain second downsampled data; Based on the first downsampled data and the second downsampled data, determine the query result of the query request.
9. The method according to claim 8, wherein, The determining the query result of the query request based on the first downsampled data and the second downsampled data includes: Obtain tombstone records for marking deleted original data; According to the time information of the data in the tombstone records and the time information of the data in the first downsampled data, determine whether the first downsampled data contains the downsampled data corresponding to the tombstone records; If the determination result is yes, delete the downsampled data corresponding to the tombstone records from the first downsampled data; Determine the second downsampled data and the first downsampled data after deleting the downsampled data corresponding to the tombstone records as the query result of the query request.
10. The method according to claim 8, wherein, The querying the second persistent storage medium according to the query request includes: Obtain the query condition corresponding to the query request from the query request; According to the query condition, generate a first field name that satisfies the field name format corresponding to the downsampled data in the second persistent storage medium; According to the first field name, query the second persistent storage medium to determine the downsampled data corresponding to the first field name; The obtaining the first downsampled data that satisfies the query request from the second persistent storage medium includes: Obtain the first downsampled data that satisfies the query condition from the downsampled data corresponding to the first field name.
11. The method according to claim 8, wherein, The first downsampled data is located in multiple downsampled files; Determining a query result of the query request based on the first downsampled data and the second downsampled data includes: For the case where there are overlapping time windows for the first downsampled data in different downsampled files, according to the aggregation operator in the query request, performing an aggregation operation on the first downsampled data corresponding to the overlapping time windows to obtain the aggregated first downsampled data; Based on the aggregated first downsampled data and the second downsampled data, determining the query result of the query request.
12. A data query method, characterized in that, it includes: Obtaining a query request; The query request is used for aggregation query; According to the query request, querying the memory and a second persistent storage medium storing downsampled data; The memory stores the original data to be persistently stored; The downsampled data stored in the second persistent storage medium is obtained by performing downsampling processing on the target original data written into the first persistent storage medium during the process of writing the original data in the memory into the first persistent storage medium; For the case where there is data in the memory that satisfies the query request, respectively obtaining the first original data and the first downsampled data that satisfy the query request from the memory and the second persistent storage medium; According to the query request, performing downsampling processing on the first original data to obtain second downsampled data; Based on the first downsampled data and the second downsampled data, determining the query result of the query request.
13. A computing system, characterized in that, it includes: A memory and a processor; The memory includes: a memory and a persistent storage medium; The processor is communicatively connected to the memory and the persistent storage medium and is configured to execute the steps in the method according to any one of claims 1-12.
14. A computer-readable storage medium storing computer instructions, characterized in that, when the computer instructions are executed by one or more processors, causing the one or more processors to execute the steps in the method according to any one of claims 1-12.
Citation Information
Patent Citations
Data processing method, system, device and computer readable storage medium
CN112084226A
Opentsdb-based data display method and device, and medium
CN112231531A
Data downsampling method, device and system and computer readable storage medium
CN113342817A