Time series data query method, computing device and computer storage medium

By selecting query keywords with high discrimination index to query the timing database and using other keywords to filter the results, the problem of low query efficiency of timing databases is solved, and a more efficient query process is achieved.

WO2025153921A1PCT designated stage expired Publication Date: 2025-07-24CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/050208
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-15
Filing Date
2025-01-09
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

In the prior art, the query method of timing databases is relatively low in efficiency and the I/O amount is large, resulting in low query efficiency.

Method used

By determining the distinction index of query keywords, selecting the target keyword from multiple query keywords, using the target keyword to query the timing database and filtering the query results with other keywords to reduce the number of queries.

Benefits of technology

Reduces the number of queries in the timing database, improves query efficiency, and reduces I/O amount.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025050208_24072025_PF_FP_ABST
    Figure IB2025050208_24072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a time series data query method, a computing device and a computer storage medium. The time series data query method comprises: receiving a time series data query request; determining at least two query keywords in the time series data query request; on the basis of a discrimination index of the at least two query keywords, determining a target query keyword from among the at least two query keywords, the discrimination index being obtained by acquiring statistics about the number of keywords comprised in time series data stored in a time series database; querying the time series database for a first query result matched with the target query keyword; and using at least one query keyword of the at least two query keywords except the target query keyword to filter the first query result, so as to obtain a query result. The technical solution provided in the embodiments of the present disclosure reduces the number of time series database queries, so as to reduce the I / O quantity of the time series database.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Time Series Data Query Method, Computing Device, and Computer Storage Medium This disclosure claims priority to Chinese patent application number 202410057345.4, filed with the China Patent Office on January 15, 2024, the entire contents of which are incorporated herein by reference. Technical Field: Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a time series data query method, computing device, and computer storage medium. Background: Time series data refers to a collection of data recorded in chronological order, where each data point in the time series data is associated with a specific timestamp. Time series data is typically used to describe data that changes over time, such as temperature, humidity, load, flow rate, stock price, and so on. Time series data is widely used in various fields, including finance, energy, smart manufacturing, the Internet of Things, healthcare, and so on. Time series data can be stored in a time series database, which can be a data management system that provides access to time series data. To uncover the value behind time series data and analyze it, it's often necessary to query the data in a time series database. When querying time series data, the query request typically contains multiple query keywords. In related art, time series databases typically perform a data query for each query keyword, obtaining multiple query results. The final query result is then generated by taking the intersection of the multiple query results. As can be seen from the above description, related art time series data query methods suffer from high I / O (Input / Output) throughput and low efficiency. SUMMARY OF THE INVENTION Embodiments of the present disclosure provide a method, apparatus, computing device, and computer storage medium for querying time series data. In a first aspect, an embodiment of the present disclosure provides a method for querying time series data, comprising: receiving a time series data query request; determining at least two query keywords in the time series data query request; determining a target query keyword from the at least two query keywords based on a discrimination index of the at least two query keywords, wherein the discrimination index is obtained by counting the frequency of keywords contained in time series data stored in a time series database; querying the time series database for a first query result that matches the target query keyword; and filtering the first query result using at least one query keyword other than the target query keyword among the at least two query keywords to obtain a query result.In a second aspect, embodiments of the present disclosure provide a time series data query device, comprising: a request receiving module for receiving a time series data query request; a first keyword determination module for determining at least two query keywords in the time series data query request; a second keyword determination module for determining a target query keyword from the at least two query keywords based on a discrimination index of the at least two query keywords, wherein the discrimination index is obtained by counting the frequency of keywords contained in time series data stored in a time series database; a matching module for querying the time series database for a first query result that matches the target query keyword; and a query module for filtering the first query result using at least one query keyword other than the target query keyword among the at least two query keywords to obtain a query result. In a third aspect, embodiments of the present disclosure provide a computing device, comprising a processing component and a storage component; the storage component storing one or more computer instructions; the one or more computer instructions being invoked and executed by the processing component to implement the time series data query method provided in embodiments of the present disclosure. In a fourth aspect, embodiments of the present disclosure provide a computer storage medium storing a computer program. When executed by a computer, the computer program implements the time series data query method provided in the embodiments of the present disclosure. In a fifth aspect, embodiments of the present disclosure provide a computer program product comprising a computer program. When executed by a computer, the computer program can implement the time series data query method provided in the embodiments of the present disclosure. In embodiments of the present disclosure, by employing a technical solution of determining a target keyword from multiple query keywords, querying a time series database using the target keyword to obtain a first query result, and then filtering the first query result using other query keywords to obtain a query result, the number of queries to the time series database can be reduced when querying time series data, thereby reducing the I / O volume of the time series database. These and other aspects of the present disclosure will be more concise and easy to understand in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS To more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. Those skilled in the art can derive other drawings based on these drawings without inventive effort.Figure 1 is a schematic diagram of a time series diagram provided by an embodiment of the present disclosure; Figure 2 is a schematic diagram of a method for querying time series data provided by an embodiment of the present disclosure; Figure 3 is a schematic diagram of a method for querying time series data provided by an embodiment of the present disclosure; Figure 4 is a schematic diagram of a time series data query apparatus provided by an embodiment of the present disclosure; and Figure 5 is a schematic diagram of a computing device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION To help those skilled in the art better understand the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the accompanying drawings. Some processes described in the specification and claims of the present disclosure and the accompanying drawings include multiple operations that appear in a specific order. However, it should be understood that these operations may be executed out of the order in which they appear herein or in parallel. Operation numbers, such as 101 and 102, are merely used to distinguish between different operations and do not represent any specific execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the terms "first," "second," and so on, used herein are used to distinguish different messages, devices, modules, and so on, and do not represent a sequential order, nor do they limit "first" and "second" to different types. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, and displayed data, etc.) involved in this disclosure are all authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data must comply with relevant laws, regulations, and standards in the relevant region, and corresponding operation portals are provided for users to choose to authorize or deny. First, the terms used in one or more embodiments of this specification are explained. Object Tag: Used to identify time series in time series data and to indicate the specific object for which the metric item of the time series data is targeted. For example, an object tag can be a data subcategory under a specified metric. The tag key and the corresponding tag value together determine the object tag. For example, an object tag consists of a tag key (TagKey) and a corresponding tag value (TagValue). For example, "City (TagKey) = Hangzhou (TagValue)" is an object tag. Another example is "Computer Room = , IP = 172.220.110.1." The relationship between tag keys and tag values ​​is one-to-one or one-to-many.When both the tag key and tag value are the same, they are the same object tag. If the tag key is the same but the tag value is different, they are not the same object tag. For example, in time series data monitoring weather, the specified metric might be "temperature" and the object tag might be "city=Hangzhou," where "city" is the tag key and "Hangzhou" is the tag value. The monitored object in this time series data is the temperature in Hangzhou. Tag Key: Used together with the corresponding tag value to determine the object tag. A tag key can be used to indicate the type of object being monitored (and together with the corresponding tag value, defines the specific object under that object type), such as country, province, city, data center, or IP address. Tag Value: The value corresponding to the tag key. For example, if the tag key is "country," the tag value might be "China." Metric: A metric of the monitored data, such as wind speed and temperature. Metric Value: The value corresponding to the metric, such as 15 (wind speed) and 20°C (temperature). Timestamp: The time when the data point was generated. Data Point: Each metric value collected at a certain time interval (such as consecutive timestamps) for a certain indicator of an object (e.g., defined by a metric and a tag) is a data point. In other words, "one metric + N object tags (N >= 1) + one timestamp + one metric value" defines a data point. Time Series: For example, a time series, as shown in Figure 1, includes data points generated at multiple timestamps. In Figure 1, the device number (Device) and region (Region) can be tag keys, and F07A1260 and North District can be the tag values ​​for the device number and region, respectively. For example, a time series can be a description of a metric (e.g., defined by a metric and a tag) for a monitored object. "One metric + N object tag KV combinations (N >= 1)" is defined as a time series. An increase in the data value generated in a time series does not cause an increase in the time series. The following will provide a clear and complete description of the technical solutions in the embodiments of the present disclosure, in conjunction with the accompanying drawings. Obviously, the described embodiments represent only a portion of the embodiments of the present disclosure, and are not exhaustive. All other embodiments derived by those skilled in the art based on the embodiments of the present disclosure without inventive effort are within the scope of protection of the present disclosure. Figure 2 schematically illustrates a flow chart of a time series data query method provided by one embodiment of the present disclosure. As shown in Figure 2, the time series data query method may specifically include the following steps:

[0002] 201, receiving a time series data query request;

[0003] 202, determining at least two query keywords in the time series data query request;

[0004] 203. Determine a target query keyword from the at least two query keywords based on discrimination indexes of the at least two query keywords, wherein the discrimination index is obtained by counting the frequencies of the keywords contained in the time series data stored in the time series database;

[0005] 204 , searching the time series database for a first query result that matches the target query keyword;

[0006] 205. Filter the first query results using at least one query keyword other than the target query keyword among the at least two query keywords to obtain a query result. According to an embodiment of the present disclosure, a time series data query request is a request sent by a user to query time series data, such as querying data within a certain time period or sorting data in chronological order. According to an embodiment of the present disclosure, when generating a time series data query request, the user can write the keyword of the time series data they wish to search for into the query request. This keyword can serve as the query keyword, so that the time series database can match the corresponding time series data using the query keyword. According to an embodiment of the present disclosure, time series data is a collection of data arranged in chronological order. It generally involves the concept of time, such as timestamps or time intervals, and is used to describe the changes in events, behaviors, or phenomena at different time points. According to an embodiment of the present disclosure, a time series data query request can be used to request to retrieve one or more data items that meet the query criteria from a collection of time series data. In an embodiment of the present disclosure, the query criteria can, for example, include the query keyword carried in the query request. According to embodiments of the present disclosure, a piece of data can refer to data recorded at a specific moment. In the time series data shown in Figure 1, each row can represent a piece of data. For example, the second row can represent data recorded at 10:01 AM on October 24, 2020. Specifically, the time series data shown in Figure 1 can represent time series data generated by detecting device temperature. The second row of data can indicate that at 10:01 AM on October 24, 2020, the temperature of the device located in the North District, with device number F07A1260, was 12.1 degrees Celsius. According to embodiments of the present disclosure, query keywords can be used to match tag values ​​in the time series data. According to embodiments of the present disclosure, in Figure 1, both the device number and region can be tag keys. F07A1260 can be the tag value corresponding to the device number tag key, and North District can be the tag value corresponding to the region tag key. When querying using a query keyword, a time series database can match the target query keyword with the tag value in the time series data, thereby querying the time series data stored in the database to obtain one or more data items containing the same tag key as the query keyword. For example, when the target query keyword is "North District", the query keyword can be matched with the tag values ​​contained in each of the multiple data items in the time series data shown in Figure 1. Since the second and fourth rows of data in this time series data contain the "North District" tag value, the second and fourth rows of data can be used as the first query result.According to an embodiment of the present disclosure, after obtaining a first query result, the first query result can be filtered using query keywords other than the target query keyword. This allows the final query result to be obtained by screening the first query result. In an embodiment of the present disclosure, by employing a technical solution that determines a target keyword from multiple query keywords, uses the target keyword to query a time series database to obtain a first query result, and then filters the first query result using other query keywords to obtain the query result, the number of queries to the time series database can be reduced when querying time series data, thereby reducing the I / O volume of the time series database. According to an embodiment of the present disclosure, the time series data query method further includes: determining a discrimination index for each of at least two query keywords. According to an embodiment of the present disclosure, determining a target query keyword from the at least two query keywords based on their discrimination indexes can be specifically implemented by determining the query keyword with the highest discrimination index among the at least two query keywords as the target query keyword. According to an embodiment of the present disclosure, the discrimination index can be negatively correlated with the frequency of keywords contained in the time series data stored in the time series database. That is, the more times a keyword is stored in the time series database, the smaller its discrimination index; and the fewer times a keyword is stored in the time series database, the larger its discrimination index. According to an embodiment of the present disclosure, by determining the query keyword with the largest discrimination index as the target query keyword, when querying the time series database using the target query keyword, a first query result with a smaller amount of data can be obtained from the time series database. As a result, when other query keywords are used to filter the first query result, the number of query keyword matches can be reduced, thereby improving the query efficiency of time series data. According to an embodiment of the present disclosure, querying a time series database for a first query result that matches a target query keyword can be specifically implemented by: querying an inverted index table of the time series database for identification information of time series data containing the target query keyword; obtaining a tag value contained in the time series data based on the identification information; determining a tag key contained in the time series data from a forward file based on the identification information; and combining the tag key and tag value according to the identification information to generate the first query result. According to an embodiment of the present disclosure, the time series database may include an inverted index table, a forward file, and a data table. The inverted index table and forward file are data structures used to manage and store time series data. The inverted index table can record the time series data identifier corresponding to each time series data item. The forward file can record the tag key contained in each time series data item. These tag keys can be obtained based on the time series data identifier obtained from the inverted index table.According to embodiments of the present disclosure, label values ​​contained in time series data can be obtained from a TSF (Time Series File). In practical applications, various programming languages ​​or software libraries can be used to read and process TSFs, such as the pandas library in Python and the ts package in R. These tools allow time series files to be loaded into memory and analyzed and processed. According to embodiments of the present disclosure, generating a first query result by combining label keys and label values ​​according to identification information can be specifically implemented by combining at least one label key and label value corresponding to the same identification information. According to embodiments of the present disclosure, obtaining label values ​​contained in time series data based on identification information can be specifically implemented by querying a time series database to determine whether the time series data corresponding to the identification information contains a label value; if not, returning an empty query result; if so, obtaining the label value. According to an embodiment of the present disclosure, before querying the forward-ranking file, a query can first be performed to determine whether the time series data corresponding to the query keyword has a tag value written. If no data is written to the time series data corresponding to the query keyword, the query result can be directly returned as empty, without further querying the forward-ranking file, thus avoiding wasted I / O. According to an embodiment of the present disclosure, determining at least two query keywords in the time series data query request includes: determining multiple query keywords included in the time series data query request; and determining at least two query keywords whose query conditions are ANDs from the multiple query keywords. According to an embodiment of the present disclosure, the multiple query keywords can be in, for example, an AND relationship, an OR relationship, a MUST relationship, or an EXCLUSION relationship. An AND relationship indicates that all query keywords must be satisfied to match the result; an OR relationship indicates that any query keyword must be satisfied to match the result; a MUST relationship indicates that certain query keywords must be satisfied, while others are optional; and a NOT relationship indicates that certain query keywords must not be present to match the result. In embodiments of the present disclosure, the query condition "and" may be a group of at least two query keywords or multiple groups. For example, a query request includes query keywords A, B, C, and D, where query keyword A and query keyword B are in an "and" relationship, and query keyword C and query keyword D are in an "and" relationship. In this case, query keyword A and query keyword B form a group, and query keyword C and query keyword D form a group.According to an embodiment of the present disclosure, for example, in the above example, query keywords A and B form a group, and query keywords C and D form a group. The query method provided in the embodiment of the present disclosure can be executed for query keywords A and B to obtain a query result, and then the query method provided in the embodiment of the present disclosure can be executed for query keywords C and D to obtain a query result. Finally, the two query results are processed, for example, by taking an intersection or a union, to obtain a final query result. According to an embodiment of the present disclosure, the query method for time series data further includes: counting the occurrence frequency of multiple keywords contained in the time series data requested to be stored in the time series database within a preset time period; and determining a discrimination index for each keyword based on the occurrence frequency, wherein the multiple keywords include the at least two query keywords. According to an embodiment of the present disclosure, within a preset time period, a query statement can be used to retrieve the time series data stored within the time period. The query statement can, for example, use a query language similar to SQL to implement filtering and aggregation operations on the time series data. The search results can then be traversed, and the occurrence frequency of each keyword therein can be counted. This can be achieved by writing a program, for example, using a statistical library in Python or another programming language, or custom code. Specifically, each time series data point can be traversed sequentially in chronological order. For each time series data point, the number of occurrences of all keywords within it is accumulated and recorded in a statistical table. Based on the statistical results, the discriminability index of each keyword is calculated. The discriminability index can be calculated based on algorithms such as TF-IDF (Term Frequency-Inverse Document Frequency) to measure the uniqueness and importance of a keyword in the entire dataset. Specifically, the frequency of occurrence of each keyword in the entire dataset and in different time series data points can be calculated, and the corresponding formula can be used to calculate the discriminability index. According to an embodiment of the present disclosure, obtaining a first query result from a time series database based on a target query keyword can be specifically implemented as follows: determining whether the amount of time series data included in the first query result is greater than a preset threshold; if so, filtering the first query result using a query keyword other than the target query keyword among at least two query keywords to obtain a query result; if not, respectively querying the time series database based on each query keyword to obtain a sub-query result corresponding to each query keyword; and obtaining a query result by taking the intersection of multiple sub-query results.According to embodiments of the present disclosure, before executing the post-filtering operation provided by embodiments of the present disclosure, the amount of data contained in the initial query results obtained by querying using the target query keyword can be first determined. If the data amount is greater than a preset threshold, a post-filtering operation can be performed, i.e., the initial query results can be filtered using other query keywords to obtain the query results. If the amount of data contained in the initial query results is relatively small, the time series database can be queried once for each query keyword to obtain sub-query results matching each query keyword, and then the intersection of the multiple sub-query results can be taken to obtain the query results. According to embodiments of the present disclosure, by first determining the amount of data contained in the initial query results, unnecessary sub-query operations can be avoided. If the amount of data in the initial query results exceeds the preset threshold, the post-filtering operation can be directly performed, reducing the number of additional database queries and improving query efficiency. Furthermore, if the amount of data in the initial query results is relatively small, the time series database can be queried separately for each query keyword, and then the intersection of the sub-query results can be taken to obtain results that simultaneously meet multiple query conditions. This approach can reduce redundant data returns, returning only data that meets all query conditions, and improving the accuracy of query results. In the embodiments of the present disclosure, the preset threshold can be flexibly set by those skilled in the art based on actual application requirements. The embodiments of the present disclosure do not limit the specific value of the preset threshold. According to the embodiments of the present disclosure, counting the occurrence frequency of multiple keywords contained in time series data requested to be stored in a time series database within a preset time period can be specifically implemented as follows: for the time series data to be stored, determining whether the multiple keywords contained in the time series data to be stored exist in a bitmap; for keywords that already exist in the bitmap, updating the counter at the index position corresponding to the keyword, the counter is used to record the occurrence frequency of the keyword; for keywords that do not exist in the bitmap, creating an index position corresponding to the keyword in the bitmap. According to the embodiments of the present disclosure, a bitmap is a data structure that can be used to represent the membership of a set, and is represented by bits. According to the embodiments of the present disclosure, the presence or absence of a keyword can be represented by a bitmap data structure. If the keyword exists in the bitmap, it means that the keyword has been indexed. Then, when the keyword is requested to be stored in the time series database again, the counter of the index position corresponding to the keyword can be directly updated. By adding one to the counter, the frequency of occurrence of the keyword can be recorded.According to embodiments of the present disclosure, if a keyword does not exist in a bitmap, it indicates that the keyword has not been indexed and does not exist in the time series database. In this case, an index position corresponding to the keyword can be created in the bitmap, and the corresponding counter can be initialized to 1. According to embodiments of the present disclosure, the index position in the bitmap can be generated using the keyword's hash value or other mapping algorithm to ensure the uniqueness of the index position. According to embodiments of the present disclosure, by using a bitmap to record keywords, the presence of keywords in the bitmap can be quickly determined, and the counters of existing keywords can be updated. Simultaneously, index positions can be created and counters initialized for new keywords. This allows for efficient management and query of the occurrence and frequency of multiple keywords in time series data. Figure 3 schematically illustrates a method for querying time series data provided by embodiments of the present disclosure. In Figure 3, a time series data query request 301 can include multiple query keywords, such as query keyword a, query keyword b, and query keyword c. In embodiments of the present disclosure, query keyword a, query keyword b, and query keyword c can be in an AND relationship. After determining the query keywords, the discrimination index of each query keyword can be determined separately. Specifically, the discrimination index can be obtained by counting the frequency of occurrence of keywords included in the time series data requested to be written to the time series database. Specifically, the discrimination index can be negatively correlated with the frequency of occurrence: the more times a keyword appears, the smaller the discrimination index, and the fewer times it appears, the larger the discrimination index. The frequency of occurrence of each keyword can be stored in a bitmap. Based on the frequency of occurrence of each keyword stored in the bitmap, query keyword a, query keyword b, and query keyword c included in the query request can be sorted. For example, the sorting order can be from highest to lowest frequency of occurrence or from lowest to highest frequency of occurrence. In this way, a target query keyword with the highest discrimination index can be determined from multiple query keywords. After determining the target query keyword, the time series database can be queried using the target query keyword. Specifically, the target query keyword can be used to query the inverted index table to obtain identification information of the time series data including the target query keyword. Then, based on the identification information, the label value included in the time series data can be obtained from the TSF (Time Series File). If no data is hit when querying TSF, the query ends and the result returned is empty. There is no need to continue to execute subsequent query procedures.If there are hits when querying the TSF, and the data volume is greater than a preset threshold, the forward index file can be queried. The forward index file can record the tag keys contained in each time series data item. These tag keys can be obtained based on the time series data identifiers obtained from the inverted index table. Furthermore, since the TSF records tag values, after querying the TSF and obtaining the query results, the TSF query results can be filtered using other query keywords to obtain identification information and tag values ​​that include query keyword a, query keyword b, and query keyword c. The forward index file can then be queried based on this information to obtain the final query results. FIG4 schematically illustrates a block diagram of a time series data query device provided by an embodiment of the present disclosure. As shown in FIG4 , the time series data query device 400 may specifically include: a request receiving module 401, configured to receive a time series data query request; a first keyword determination module 402, configured to determine at least two query keywords in the time series data query request; a second keyword determination module 403, configured to determine a target query keyword from the at least two query keywords based on a discrimination index of the at least two query keywords, wherein the discrimination index is obtained by counting the frequency of keywords contained in the time series data stored in the time series database; a matching module 404, configured to query the time series database for a first query result that matches the target query keyword; and a query module 405, configured to filter the first query result using at least one query keyword other than the target query keyword among the at least two query keywords to obtain a query result. According to an embodiment of the present disclosure, the time series data query device further includes a discrimination determination module, configured to determine the discrimination index of each of the at least two query keywords. According to an embodiment of the present disclosure, the second keyword determination module 403 may include: a target keyword determination unit, configured to determine the query keyword with the highest discrimination index as the target query keyword. According to an embodiment of the present disclosure, the matching module 404 may include: an identification query submodule, configured to query the inverted index table of the time series database for identification information of time series data containing the target query keyword; a label value acquisition submodule, configured to acquire the label value contained in the time series data based on the identification information; a keyword query submodule, configured to determine the label key contained in the time series data from the forward index file based on the identification information; and a result query submodule, configured to combine the label key and label value according to the identification information to generate a first query result.According to an embodiment of the present disclosure, the tag value acquisition submodule includes: a tag value determination unit, configured to query a time series database to determine whether the time series data corresponding to the identification information contains a tag value; a result return unit, configured to return an empty query result if the time series data corresponding to the identification information does not contain a tag value, and to obtain the tag value if the time series data corresponding to the identification information contains a tag value. According to an embodiment of the present disclosure, the time series data query apparatus also includes: a third keyword determination module, configured to determine multiple query keywords included in a time series data query request; a fourth keyword determination module, configured to determine at least two query keywords for which the query condition is AND from the multiple query keywords. According to an embodiment of the present disclosure, the time series data query apparatus also includes: a statistics module, configured to respectively count the occurrence frequencies of multiple keywords included in the time series data requested to be stored in the time series database within a preset time period; and a discrimination determination module, configured to determine a discrimination index for each keyword based on the occurrence frequencies, wherein the multiple keywords include the at least two query keywords. According to an embodiment of the present disclosure, the query module 405 includes: a threshold determination submodule for determining whether the amount of time series data included in the initial query result is greater than a preset threshold; a filtering module for filtering the first query result using at least two query keywords, excluding the target query keyword, to obtain a query result if the amount of time series data included in the initial query result is greater than the preset threshold; a parallel query module for querying the time series database for each query keyword to obtain a subquery result corresponding to each query keyword if the amount of time series data included in the first query result is less than the preset threshold; and a result determination module for taking the intersection of multiple subquery results to obtain a query result. According to an embodiment of the present disclosure, the statistics module includes: a bitmap determination submodule for determining whether a bitmap contains multiple keywords included in the time series data to be stored; an update submodule for updating the counter of the index position corresponding to the keyword already in the bitmap, the counter being used to record the frequency of occurrence of the keyword; and a creation submodule for creating an index position corresponding to the keyword in the bitmap if the keyword does not exist in the bitmap. The time series data query device of FIG4 can execute the time series data query method of the embodiment shown in FIG2 , and its implementation principle and technical effects are not further described. The specific manner in which each module and unit performs operations in the time series data query device of the above embodiment has been described in detail in the embodiment of the method and will not be elaborated on here.In one possible design, the time series data query apparatus provided in the embodiments of the present disclosure can be implemented as a computing device. As shown in FIG5 , the computing device may include a storage component 501 and a processing component 502. The storage component 501 stores one or more computer instructions, wherein the one or more computer instructions are invoked and executed by the processing component 502 to implement the time series data query method provided in the embodiments of the present disclosure. Of course, the computing device may also include other components, such as an input / output interface and a communication component. The input / output interface provides an interface between the processing component and a peripheral interface module, which may be an output device, an input device, etc. The communication component is configured to facilitate wired or wireless communication between the computing device and other devices. The computing device may be a physical device or an elastic computing host provided by a cloud computing platform. In this case, the computing device may refer to a cloud server, and the processing component, storage component, etc. may be basic server resources rented or purchased from the cloud computing platform. When the computing device is a physical device, it may be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device. The present disclosure also provides a computer-readable storage medium storing a computer program. When executed by a computer, the computer program can implement the time series data query method provided in the present disclosure. The present disclosure also provides a computer program product, including the computer program. When executed by a computer, the computer program can implement the time series data query method provided in the present disclosure. The processing component in the above embodiments may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above method. The storage component is configured to store various types of data to support operations in the device. The storage component can be implemented by any type of volatile or non-volatile memory device or a combination of them, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.Those skilled in the art will clearly understand that, for ease and brevity of description, the specific operating processes of the systems, devices, and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be detailed here. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one location or distributed across multiple network units. Some or all of the modules can be selected based on actual needs to achieve the objectives of the present embodiment. Those skilled in the art will be able to understand and implement the present embodiment without inventive effort. Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a required general-purpose hardware platform, or alternatively, hardware. Based on this understanding, the essence of the above-mentioned technical solutions, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes instructions for enabling a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or portions thereof. Finally, it should be noted that the above-mentioned embodiments are merely illustrative of the technical solutions of the present disclosure, and are not intended to limit them. Although the present disclosure has been described in detail with reference to the above-mentioned embodiments, persons of ordinary skill in the art will understand that the technical solutions described in the above-mentioned embodiments may be modified, or some of the technical features thereof may be replaced by equivalents. Such modifications or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present disclosure.

Claims

Claims 1. A method for querying time-series data, wherein, Including: Receiving a time-series data query request; Determining at least two query keywords in the time-series data query request; Based on the discrimination index of the at least two query keywords, determining a target query keyword from the at least two query keywords, where the discrimination index is obtained by statistically counting the frequencies of the keywords included in the time-series data stored in the time-series database; Querying a first query result that matches the target query keyword from the time-series database; using at least one query keyword other than the target query keyword among the at least two query keywords to filter the first query result to obtain a query result.

2. The method according to claim 1, wherein The method further includes: respectively determining the discrimination index of each of the at least two query keywords; the determining a target query keyword from the at least two query keywords based on the discrimination index of the at least two query keywords includes: determining the query keyword with the largest discrimination index among the at least two query keywords as the target query keyword.

3. The method according to claim 1 or 2, wherein The querying a first query result that matches the target query keyword from the time-series database includes: querying the identification information of the time-series data containing the target query keyword from the inverted index table of the time-series database; based on the identification information, obtaining the tag value included in the time-series data; based on the identification information, determining the tag key included in the time-series data from the forward file; according to the identification information, combining the tag key and the tag value in a corresponding manner to generate the first query result.

4. The method according to claim 3, wherein The obtaining the tag value included in the time-series data based on the identification information includes: querying whether the time-series data corresponding to the identification information in the time-series database includes a tag value; if not, returning that the query result is empty; if so, obtaining the tag value.

5. The method according to any one of claims 1 to 4, wherein The determining at least two query keywords in the time-series data query request includes: determining a plurality of query keywords included in the time-series data query request; determining at least two query keywords whose query conditions are "AND" from the plurality of query keywords.

6. The method according to any one of claims 1 to 5, wherein The method further includes: respectively statistically counting the occurrence frequencies of the plurality of keywords included in the time-series data requested to be stored in the time-series database within a preset time period; According to the occurrence frequencies, determining the discrimination index of each keyword, where the plurality of keywords includes the at least two query keywords.

7. The method according to any one of claims 1 to 6, wherein Filtering the first query result by using at least one query keyword among the at least two query keywords other than the target query keyword, the obtained query result includes: determining whether the number of time-series data included in the first query result is greater than a preset threshold; if so, filtering the first query result by using the query keywords other than the target query keyword among the at least two query keywords to obtain the query result; if not, respectively querying from the time-series database sub-query results corresponding to each query keyword based on each query keyword; taking the intersection of the multiple sub-query results to obtain the query result.

8. The method according to claim 6, wherein The separately counting the occurrence frequencies of multiple keywords included in the time-series data requested to be stored in the time-series database within a preset time period includes: for the time-series data to be stored, determining whether there are multiple keywords included in the time-series data to be stored in the bitmap; for the keywords that already exist in the bitmap, updating the counter at the index position corresponding to the keyword, where the counter is used to record the occurrence frequency of the keyword; for the keywords that do not exist in the bitmap, creating an index position corresponding to the keyword in the bitmap.

9. A computing device, wherein, It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the time-series data query method according to any one of claims 1 to 8.

10. A computer storage medium, wherein, A computer program is stored, and when the computer program is executed by the computer, it implements the time-series data query method according to any one of claims 1 to 8.

11. A computer program product includes a computer program, and when the computer program is executed by the computer, it can implement the time-series data query method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Data query method and system based on distributed SQL (Structured Query Language)

    CN114817293A

  • Search engine with suggestion tool and method of using same

    US20060248078A1

  • Text search of database with one-pass indexing including filtering

    US20190034523A1

  • Pipelined search query, leveraging reference values of an inverted index to determine a set of event data and performing further queries on the event data

    US20200034363A1

  • Keywords associated with document categories

    US7996393B1