Data query method and equipment

By estimating the query time and outputting recommended information, the problem of long response time in Ad Hoc data queries is solved, the query success rate and user experience are improved, and the query efficiency and system response speed are optimized.

CN120653830APending Publication Date: 2025-09-16HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410302720.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

During the Ad Hoc data query process, the query task response time is too long or even times out because the amount of data to be queried is too large due to the query conditions specified by the user, which affects the user experience and system reliability.

Method used

By estimating the query time and outputting recommended information, it guides users to adjust the query conditions to optimize the query process, including recommending query time periods and data ranges, thereby improving query efficiency and success rate.

Benefits of technology

It improves the success rate of Ad Hoc queries, improves system response speed and user experience, optimizes query efficiency, and gives users greater customization options to adapt to query needs in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653830A_ABST
    Figure CN120653830A_ABST
Patent Text Reader

Abstract

The invention provides a data query method and equipment, and relates to the technical field of computers. The problem that the data query task fails due to the fact that the query condition specified by the user corresponds to the situation that the data size of the data needing to be queried is too large in the data query process in an Ad Hoc mode is solved. Receiving a first query condition, wherein the first query condition is used for querying first data of the log library; on the basis of the first query condition, estimating query time required for querying the first data in the log library; under the condition that the query time consumption is larger than a preset threshold value, recommendation information is output, the recommendation information is used for recommending the user to query second data in the log library by using a second query condition, and the second data is part of data in the first data; the query time consumption of the second data queried based on the second query condition is less than the query time consumption of the first data queried based on the first query condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method and device for data query. Background Art

[0002] With the rapid development of information technology and the accelerated pace of enterprise digital transformation, cloud-native service architectures have become a core trend in modern information technology environments. In this context, the volume of log data generated by various systems and services is growing exponentially, encompassing multiple dimensions, including internal cloud-native service operation records, system operation status logs, security audit information, and internet user behavior traces. These scenarios are referred to as pan-logging scenarios. With the advancement of technology, data management requirements in pan-logging scenarios have become more important than ever. These requirements not only affect system stability and security but are also key to improving business insights and decision-making efficiency.

[0003] In order to achieve more flexible data query and immediate business decision support, existing technologies can use ad hoc query (Ad Hoc) to meet the data management needs in the pan-log scenario. Among them, ad hoc query allows users to specify query conditions arbitrarily. However, in actual application, if the query conditions specified by the user correspond to the amount of data to be queried that is too large, the data to be queried involves high-dimensional fields, and the length of the fields to be queried is too long, it will cause the response time to be too long or even timeout when executing the query task, thereby affecting the user experience and causing users to question the processing capabilities and reliability of the system. Summary of the Invention

[0004] The embodiments of the present application provide a method and device for data query, which solves the problem of data query task failure during data query in an Ad Hoc manner due to the fact that the amount of data to be queried corresponding to the query conditions specified by the user is too large.

[0005] To achieve the above objectives, the present invention provides the following technical solutions:

[0006] In a first aspect, a data query method is provided, which may include:

[0007] First, a first query condition is received, where the first query condition is used to query first data in a log library; based on the first query condition, a query time required to query the first data in the log library is estimated; if the query time is greater than a preset threshold, recommendation information is output, where the recommendation information is used to recommend that the user use a second query condition to query second data in the log library, where the second data is part of the first data, and the query time for querying the second data based on the second query condition is less than the query time for querying the first data based on the first query condition.

[0008] This technical solution estimates the time required to search the log library for data matching the query criteria and, based on the estimated time, adaptively outputs recommended information, guiding users to adjust their query criteria to obtain the desired data. This significantly improves the success rate of AdHoc queries and enhances system response speed, query efficiency, and user experience.

[0009] In conjunction with the first aspect, in one possible implementation, the above-mentioned estimation of the query time required to query the first data in the log library based on the first query condition specifically includes: estimating the query time required to query the first data in the log library based on statistical information and the first query condition; wherein the log library stores data in the form of segments, and the statistical information includes: the identification of all segments included in the log library, the identification of all fields of the data included in each segment, the total count of the data corresponding to each field, and the size of the storage space occupied by the data corresponding to each field. In this way, targeted retrieval based on statistical information and estimation of query time can optimize query performance and improve query efficiency.

[0010] In combination with the first aspect, in a possible implementation, the first query condition includes: the identifier of the field to be queried of the first data and the time period to be queried. The above-mentioned estimation of the query time required for querying the first data in the log library based on the statistical information and the first query condition specifically includes: obtaining the number of rows of the first data in the log library that conform to the time period to be queried; according to the identifier of the field to be queried, obtaining the target total count and target size from the statistical information, the target total count being the total count of the data corresponding to the field to be queried, and the target size being the size of the storage space occupied by the data corresponding to the field to be queried; estimating the query time based on the number of rows of the first data, the target total count and the target size. In this way, the time consumption of the query task is estimated based on the number of rows of the first data, the target total count and the target size, which can accurately estimate the time consumption required for the query and improve the query efficiency.

[0011] In conjunction with the first aspect, one possible implementation includes, before estimating the query time required to query the first data in the log repository based on the first query condition, further comprising: performing segment-based statistics on the data stored in the log repository to obtain statistical information. Thus, performing segment-based statistics on the log repository helps refine data management, improve query speed, and reduce query costs.

[0012] In conjunction with the first aspect, the recommendation information may further include one or more of the following: a recommended query time period and a recommended search range for data in the log library. Furthermore, when the recommendation information is a recommended query time period, the recommendation information may also include an estimated query time using the second query condition. This provides users with more options and greater customization to meet query needs in different scenarios, thereby enhancing the user experience.

[0013] In combination with the first aspect, in the case where the recommendation information includes a recommended query time period, before outputting the recommendation information, it also includes: based on a preset rule, adjusting the time period to be queried in the first query condition once or multiple times to obtain a second query condition; determining the recommendation information based on the second query condition; wherein, in any adjustment process, the query condition adjusted last time is adjusted again according to the preset rule to obtain the query condition adjusted this time; based on the query condition adjusted this time, estimating the query time required to query the data corresponding to the query condition adjusted this time in the log library, and when the query time is less than the preset threshold, using the query condition adjusted this time as the second query condition. In this way, in the case where the recommendation information output by the present application includes a recommended query time period, by guiding the user to obtain the expected results by only adjusting the query time range, the query efficiency can be optimized without affecting other query conditions.

[0014] In addition, when the recommendation information includes: a recommended range for searching data in the log library, the query system can calculate the number of rows of data in the log library that the query system can process within the expected time period based on the query time expected by the user. In this way, when the recommendation information output by the present application includes a recommended range for searching data in the log library, the user can determine the second query condition based on the number of rows of data in the log library that the query system can process. This can give users a higher degree of customization to adapt to the time range query requirements in different scenarios, significantly improving user satisfaction.

[0015] In conjunction with the first aspect, before estimating the query duration required to query the log repository for the first data based on the first query condition, the method further includes: executing a query task in the log repository based on the first query condition; and determining if the duration of executing the query task exceeds a preset threshold. In this way, by pre-executing the query task based on the user's query condition, system response delays or failures caused by excessive query time can be prevented while still meeting user needs.

[0016] In a second aspect, a device is provided that implements the electronic device behavior described in the method of the first aspect. This functionality can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned functionality, such as a communication unit or module, a processing unit or module, and a storage unit or module.

[0017] In a third aspect, an electronic device is provided, comprising: a processor; a memory; and a computer program, wherein the computer program is stored in the memory, and when the computer program is executed by the processor, the electronic device executes the method described in the first aspect above.

[0018] In a fourth aspect, a computer-readable storage medium is provided, which includes a computer program. When the computer program runs on an electronic device, the electronic device can execute the method described in the first aspect.

[0019] In a fifth aspect, a computer program product comprising instructions is provided, which, when executed on an electronic device, enables the electronic device to execute the method described in the first aspect above.

[0020] In a sixth aspect, an embodiment of the present application provides a chip, the chip including a processor, the processor being used to call a computer program in a memory to execute the method described in the first aspect.

[0021] It can be understood that the beneficial effects that can be achieved by the methods described in the first and second aspects, the device described in the second aspect, the electronic device described in the third aspect, the computer-readable storage medium described in the fourth aspect, the computer program product described in the fifth aspect, and the chip described in the sixth aspect can be referred to the beneficial effects in the first aspect and any possible implementation method thereof, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 A schematic diagram of a data query method provided for related technologies;

[0023] Figure 2 A schematic diagram of a query system provided in an embodiment of the present application;

[0024] Figure 3 A flowchart of a data query method provided in an embodiment of the present application;

[0025] Figure 4 A schematic diagram of a data query method provided in an embodiment of the present application;

[0026] Figure 5 A schematic diagram of another data query method provided in an embodiment of the present application;

[0027] Figure 6 A schematic diagram of another data query method provided in an embodiment of the present application;

[0028] Figure 7 A schematic diagram of another data query method provided in an embodiment of the present application;

[0029] Figure 8 A schematic diagram of the composition of a data query device provided in an embodiment of the present application;

[0030] Figure 9 A schematic diagram of the hardware structure of a data query device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0032] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0033] In the following description, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "multiple" means two or more.

[0034] In the description of this application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship, for example, A / B can represent A or B; "and / or" in this application is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural.

[0035] It is understood that some optional features in the embodiments of the present application may, in certain scenarios, be implemented independently of other features, such as the solution on which they are currently based, to solve corresponding technical problems and achieve corresponding effects. They may also be combined with other features in certain scenarios as needed. Accordingly, the devices provided in the embodiments of the present application may also implement these features or functions accordingly, which will not be described in detail here.

[0036] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0037] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0038] 1) Ad Hoc Query: This is a database query method that allows users to flexibly construct query commands based on their needs without pre-programming or predefined structures. This means users can flexibly select query criteria based on their needs. The log service system can then generate corresponding statistical results based on the user's selected query criteria. This query method gives users greater freedom and flexibility. Compared to ordinary application queries, users' queries are not limited to pre-defined fixed query templates. Users can customize query criteria according to their own needs to obtain the required data.

[0039] With the trend of enterprise digital transformation and cloud-native services, data management requirements for general logging scenarios, such as cloud-native service and system logs, security audits, and internet user behavior analysis, are becoming increasingly important. In response to this trend, numerous cloud-native logging service systems have emerged to collect log data from hosts and cloud services. By analyzing and processing this massive amount of log data, these systems maximize the availability and performance of cloud services and applications, providing real-time, efficient, and secure log processing capabilities. This allows users to quickly and efficiently conduct real-time decision analysis, manage device operations and maintenance, and analyze user business trends.

[0040] In the related art, users can use the Ad Hoc method to implement data management by specifying query conditions by themselves. However, in actual application, when users select query conditions according to their own needs, they often fail to take into account the actual amount of data queried and the ability of electronic devices to process data, which can easily lead to situations where the amount of data to be queried is too large, the data to be queried involves high-dimensional fields, the length of the fields to be queried is too large, etc. These situations may cause the log service system to have a response time that is too long or even exceeds the time threshold when performing data query tasks. For queries that exceed the time threshold, the user will generally be returned the following: Figure 1 This timeout prompt can cause confusion and frustration for users, even leading them to question the system's processing capabilities and reliability, impacting the user experience. Furthermore, after seeing the prompt, users may retry, which will still result in query failure and waste system resources.

[0041] Therefore, it is necessary to find a data query method that can guide users to improve the success rate of AdHoc queries while ensuring the smooth completion of data query tasks.

[0042] To address the aforementioned issues, embodiments of the present application provide a data query method that can be applied to data queries in a log repository. Specifically, a user can select a query condition, such as a first query condition, based on their query needs. Based on the first query condition, the method can then estimate the query time required to search the log repository for the first data corresponding to the first query condition. If the query time exceeds a preset threshold, recommended information is output. By estimating the query time required to search the log repository for data corresponding to the query condition and adaptively outputting recommended information based on the estimated result, the user is guided to obtain the desired data by adjusting the query condition. This not only significantly improves the success rate of AdHoc queries, but also enhances the system's response speed, query efficiency, and user experience. Furthermore, the recommended information output by the present application can include a recommended query time period. By guiding the user to adjust only the query time range to obtain the desired results, query efficiency can be optimized without affecting other query conditions. Furthermore, the recommended information output by the present application can include a recommended search range for data in the log repository. This gives users greater customization options to meet the time range query needs of different scenarios, significantly improving user satisfaction.

[0043] The data query method provided in the embodiments of the present application can be applied to a query system. The query system can be a log query system (or log service system) or a data warehouse query system. The embodiments of the present application do not impose any restrictions on the specific form of the query system. For example, the query system can be a terminal device or a network device. The terminal device can be referred to as a terminal, terminal equipment, access terminal, mobile station, remote station, remote terminal, mobile device or user terminal, wireless communication device, etc. The terminal device can be a mobile phone, tablet computer, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc. The query system can be a server, etc. The server can be a physical or logical server, or two or more physical or logical servers sharing different responsibilities and cooperating to implement the various functions of the server.

[0044] In this embodiment, the query system that executes the above solution can be a server in a server cluster (composed of multiple servers communicating with each other), or a chip in the server, or a system on a chip in the server, or can be implemented by a virtual machine (VM) deployed on a physical machine. The query system that executes the above solution can also be a server cluster composed of multiple servers communicating with each other. This embodiment of the application is not limited to this.

[0045] Figure 2 2 is a schematic diagram of a query system 20 provided in an embodiment of the present application. The query system 20 may include a column storage counter 201, a query evaluator 202, and a front-end director 203. Figure 2 The query system 20 shown is only a feasible example. In other feasible embodiments, the query system 20 may also include only Figure 2 Some of the components shown may also include other components.

[0046] In some embodiments, the column storage statistician 201, the query evaluator 202 and the front-end director 203 establish communication connections in sequence and can communicate through an agreed protocol.

[0047] The column storage counter 201 can be used to count the data in the log library in segments to obtain statistical information. The statistical information may include: the identifiers of all segments included in the log library, the identifiers of all fields of data included in each segment, the total count of data corresponding to each field, and the size of the storage space occupied by the data corresponding to each field.

[0048] The query evaluator 202 can be used to receive query conditions input by the user based on his or her own query needs. It can also be used to obtain statistical information from the column storage statistician 201. Based on the received query conditions and the obtained statistical information, the query evaluator 202 can estimate the query time required to query the data required by the user in the log library. Finally, when it is determined that the data the user wants to query cannot be returned based on the query time, new query conditions are determined based on the query conditions input by the user, and this is used as recommendation information. Among them, the recommendation information is used to recommend the user to use the new query conditions to query the corresponding data in the log library, and the data corresponding to the new query conditions are part of the data the user wants to query. The query time required to query the corresponding data based on the new query conditions is less than the query time required to query the corresponding data based on the query conditions initially input by the user.

[0049] After determining the recommended information, query evaluator 202 may input the recommended information to front-end director 203. Front-end director 203 may be configured to display the recommended information for the user to view, thereby enabling the user to determine new query conditions based on the recommended information and thereby obtain desired results using the new query conditions. It is understood that the new query conditions may improve the success rate of AdHoc queries.

[0050] It should be noted that the functions of the above-mentioned devices can be implemented by the same server or a chip in a server, or a system on a chip in a server, or a VM deployed on a physical machine. In this case, it can be understood that the solution of this application is applied to one server. The functions of the above-mentioned devices can also be implemented by multiple different servers respectively. In this case, it can be understood that the solution of this application is applied to a server cluster. For example, the functions of the column storage statistician 201, the query evaluator 202, and the front-end director 203 are respectively implemented by three different servers. These three servers can be understood as a server cluster, and the solution of this application can be implemented through the interaction between these three servers in the server cluster.

[0051] Figure 3 This is a flow chart of a query method provided in an embodiment of the present application. Figure 3 As shown, the method may include:

[0052] S301: Receive a first query condition.

[0053] The first query condition can be used to query first data in the log library. The first query condition can include: an identifier of a field to be queried in the first data and a time period to be queried. The first data can be partial data in the log library. It is understood that the first data can specifically be data in the log library that meets the first query condition.

[0054] In an embodiment of the present application, when a user needs to search for required data in a log library, such as first data, the user can enter query conditions in the query interface of the corresponding query engine, such as the first query conditions described above, including the identifier of the field to be queried of the first data and the time period to be queried. In this way, an electronic device equipped with the query engine, such as a computer, can receive the first query conditions entered by the user and send the first query conditions to the query system. Accordingly, the query system can receive the first query conditions from the electronic device.

[0055] In a log library, data of the same type can be aggregated and stored in columns. Each column represents a specific attribute or event element. These columns are typically referred to as fields in database terminology. Each field has a corresponding field identifier, often referred to as an identifier, which clarifies the specific meaning of the data associated with that field. For example, the field name can be used as the field identifier.

[0056] The identifier of the field to be queried may refer to the identifier of the field corresponding to the data that the user specifies when performing a data query and wishes to obtain from the log library. The identifier of the field may specifically be the name of the field, etc. For example, take the name of the field as an example. In a log query scenario related to e-commerce, the log library may contain fields such as the "user level" field, the "user location information" field, and the "order status" field, and of course, the data corresponding to these fields are also stored. When the user wants to obtain data related to the user's location information from the log library, the identifier of the field to be queried may be the "user location information" field. For another example, in a log query scenario related to a school, the log library may contain fields such as the "name" field, the "gender" field, the "class" field, and the "grade" field, and the data corresponding to these fields. When the user wants to obtain data related to the class from the log library, the identifier of the field to be queried is the "class" field.

[0057] The time period to be queried can refer to the time interval corresponding to the data that the user specifies when performing a data query and wishes to obtain from the log library. It is understandable that the log library will store the corresponding data in the log library according to the time when the log occurred. For example, a user wishes to query the data recorded in the log library between 10:20 on February 21, 2023 and 12:20 on February 23, 2023. In this way, the time period to be queried is "10:20 on February 21, 2023 to 12:20 on February 23, 2023".

[0058] It is understood that users can write query conditions in various languages, and common languages ​​may include Structured Query Language (SQL).

[0059] For example, if a user enters a first query condition in SQL on a query interface, an electronic device, such as a computer, can receive the first query condition written in SQL and send it to a query system. The query system can then receive the first query condition written in SQL from the computer. After receiving the first query condition written in SQL, the query system can parse it to obtain the identifier of the field to be queried and the time period to be queried.

[0060] For example, suppose a user wishes to query the names of users in the log library between 12:00 AM and 12:00 PM on February 28, 2024. The user enters the first query condition "request_method:get select name, count(*) group by name[time:2024 / 02 / 28 00:00~2024 / 02 / 28 12:00]" using SQL in the query interface. An electronic device, such as a computer, can receive the first query condition entered by the user and send it to the query system. After receiving the first query condition, the query system can parse it to obtain the following information contained in the first query condition: the identifier of the query field is name, i.e., the "name" field; and the query time period is "2024 / 02 / 28 00:00~2024 / 02 / 28 12:00," i.e., from 12:00 AM to 12:00 PM on February 28, 2024.

[0061] After receiving the first query condition, the query time required to search the log library for the first data corresponding to the first query condition can be estimated. As an example, the following S302 can be specifically performed.

[0062] S302: Estimate the query time required to query the first data in the log library based on the statistical information and the first query condition.

[0063] The statistical information may be information obtained by analyzing the data in the log library. The log library may store data in segments. The statistical information may specifically include: the identifiers of all segments included in the log library, the identifiers of all fields of data included in each segment, the total count of data corresponding to each field, and the amount of storage space occupied by the data corresponding to each field.

[0064] In some embodiments, before estimating the query time required to query the first data in the log repository, the query system may perform statistics on the data in the log repository to obtain statistical information of the log repository.

[0065] As described above, the log library stores data in segments. Specifically, Figure 4As shown in (a), when a new document is indexed, it is placed in a buffer. When the buffer reaches a certain size or meets other conditions, the search engine (Lucene) can perform an operation to write the data in the cache to the main memory or disk, thereby merging the documents in the buffer into a new segment, which can also be understood as storing the data in the log library. When a new segment is created, the corresponding buffer will be cleared and ready to receive the next batch of new documents. The number of documents contained in each segment can be the same or different. For example, a segment can contain the data of one document, and a segment can also contain the data of two or more documents. In addition, based on the previous introduction to the log library data storage rules, the document data in the log library is generally stored in the form of columns to aggregate the same type of data, and such columns are called fields and have corresponding identifiers.

[0066] After storing data in segments, the query system can count the data in each segment in units of segments. For example, the query system can periodically count the data in each segment. That is, the data in each segment is counted once at fixed intervals. The query system can also count the data in all segments for the first time, and then re-count the data in a segment when the data in a segment is updated. The statistics of the data in the segment can be based on the segment as a unit, and for each field in all the fields included in the segment, the total count of the corresponding data and the size of the storage space occupied by the corresponding data can be counted.

[0067] For example, if the log library includes two segments, segment0 and segment1, the query system can first count the data in segment0 and segment1 in the log library when obtaining segment0 and segment1 for the first time. Then, when the data in the documents in segment0 and / or segment1 changes / updates, the data of the changed / updated segments is re-counted. In this way, the total count of the data corresponding to each field in all the fields included in segment0 and the size of the storage space occupied by the corresponding data can be obtained. The total count of the data corresponding to each field in all the fields included in segment1 and the size of the storage space occupied by the corresponding data can also be obtained.

[0068] After statistics are collected for the data in each segment in the log library, the query system can generate statistical information based on the statistical results of each segment. The statistical information may include: the identifier of each segment, the identifiers of all fields of the data included in each segment, the total count of data corresponding to each field, the size of the storage space occupied by the data corresponding to each field, etc. The statistical information can be presented in the form of a table. In addition, the query system can also present the statistical information of the log library in other forms, such as statistical charts, etc., which is not limited in this application.

[0069] For example, Figure 4 As shown in (b), the query system generates statistical information based on the statistical results of segment0 and segment1, and presents it in the form of a table. It can be seen that the statistical information can include: the identifiers of segment0 and segment1, the fields (field) corresponding to the identifiers of segment0 and segment1, the location (ip) field and the status (status) field, the total count (count) of the data corresponding to the ip field, the size (size) of the storage space occupied by the data corresponding to the ip field, the total count (count) of the data corresponding to the status field, and the size (size) of the storage space occupied by the data corresponding to the status field. Based on this statistical information, it can be understood that in segment0, the total count of the data corresponding to the ip field is 200 million (M), the size of the storage space occupied is 1000M, and the total count of the data corresponding to the status field is 200M, the size of the space occupied is 250M. In segment 1, the total count of data corresponding to the ip field is 400M, and the space occupied is 4000M. The total count of data corresponding to the status field is 400M, and the space occupied is 500M.

[0070] In this way, after receiving the first query condition, the query system can estimate the query time required to obtain the first data corresponding to the first query condition based on the above statistical information.

[0071] It is understandable that in the process of user searching the log library, the query time is generally the time from the user inputting the query conditions to the query system returning the query results, which can be called the end-to-end (E2E) time. In this scenario, the E2E time mainly includes: indexing time, aggregation time, merging time and network time. Since Lucene itself does not directly support complex aggregation operations, the slow query speed of Lucence is usually because aggregation takes a long time. In other words, the aggregation time is much greater than the indexing time, merging time and network time. Therefore, in the embodiment of the present application, the indexing time, merging time and network time can be ignored, and the aggregation time can be used as the query time.

[0072] The aggregation time can be estimated through the following steps.

[0073] S51: Obtain the number of rows of first data that matches the time period to be queried in the log library.

[0074] In some embodiments, the query system can obtain the number of rows of first data in the log library that match the query time period in the received first query condition. Specifically, the number of rows of first data in each segment in the log library that match the query time period is obtained.

[0075] For example, the query system can filter each segment in the log repository based on the query time period in the first query condition to obtain the number of rows of data in each segment that meet the query time period. It is understandable that the query system can find the number of rows of all data whose occurrence time falls within the query time period in the massive log records, thereby filtering out the number of rows of data in the log repository that meet the query time period in the first query condition.

[0076] For example, if a user wants to query the addresses of customers from February 21, 2023 to February 23, 2023 in the log library, and the first query condition entered by the user in the query interface using SQL is "request_method:get select remote_addr,count(*)group by remote_addr[time:2023 / 02 / 21~2023 / 02 / 23]", the log library contains approximately 100 million rows of log records, and these records are distributed in segment 0 and segment 1. Figure 5As shown in (a), it can be understood that the query time period included in the first query condition is "[time:2023 / 02 / 21~2023 / 02 / 23]", that is, February 21, 2023 to February 23, 2023. The query system can filter segment0 and obtain a total of 20 million rows of data that meet the query time period "February 21, 2023 to February 23, 2023", that is, the number of rows that meet the condition in segment0 is 20 million. The query system can also filter segment1 and obtain a total of 40 million rows of data that meet the query time period "February 21, 2023 to February 23, 2023", that is, the number of rows that meet the condition in segment1 is 40 million.

[0077] S52. According to the identifier of the field to be queried, obtain the total target count and target size from the statistical information.

[0078] The target total count refers to the total count of data corresponding to the query field, and the target size refers to the storage space occupied by the data corresponding to the query field. Specifically, the total count and storage space occupied by the data corresponding to the query field in each segment are obtained.

[0079] In some embodiments, the query system may first parse the identifier of the field to be grouped, i.e., the identifier of the field to be queried, based on the first query condition. Then, based on the identifier of the field to be queried, the query system may obtain the total count and the amount of storage space occupied by the data corresponding to the field to be queried from the statistical information for each segment.

[0080] For example, continuing the above example, take the first query condition as "request_method:get select remote_addr,count(*)group by remote_addr[time:2023 / 02 / 21~2023 / 02 / 23]". It is understandable that the query system can parse the obtained field that needs to be grouped based on the first query condition, that is, the identifier of the field to be queried is the address field (that is, the addr field). Combined with Figure 5In (a), the statistical information obtained from the log library statistics is: in segment0, the total count of the data corresponding to the addr field is 200M, and the size of the storage space occupied is 1000M; the total count of the data corresponding to the status field is 200M, and the size of the space occupied is 250M. In segment1, the total count of the data corresponding to the addr field is 400M, and the size of the space occupied is 4000M; the total count of the data corresponding to the status field is 400M, and the size of the space occupied is 500M. The query system can obtain the total count of the data corresponding to the addr field and the size of the storage space occupied from the statistical information for each segment, that is, for each segment, obtain the target total count and target size. For example, the query system can obtain the target total count and target size based on the statistics. Figure 5 The statistical information shown in (a) shows that the total count of data corresponding to the addr field in segment0 is 200M, and the storage space occupied is 1000M. The query system can also be based on Figure 5 According to the statistical information shown in (a), the total count of the data corresponding to the addr field in segment1 is 400M, and the space occupied is 4000M.

[0081] S53: Estimate the query time consumption according to the number of rows of the first data, the total target count, and the target size.

[0082] In some embodiments, the query system may first calculate the average size of the field to be queried based on the target total count and target size corresponding to each segment.

[0083] For example, for a segment, the average size of the field to be queried can be calculated using the following formula:

[0084]

[0085] Among them, size(fi) is the average size of the field to be queried in the segment, size is the total count of the data corresponding to the field to be queried in the segment, that is, the above-mentioned target total count, and count is the size of the storage space occupied by the data corresponding to the field to be queried in the segment, that is, the above-mentioned target size.

[0086] For example, continuing the above example, combined with Figure 5 In (a), taking the addr field as an example, the average size of the addr field in segment 0 is 1000M / 200M=5M. The average size of the addr field in segment 1 is 4000M / 400M=10M.

[0087] After obtaining the size of the field to be queried for each segment, the query system can determine the size of the first data that meets the first query condition in the segment based on the number of rows of first data in each segment and the average size of the field to be queried. In some embodiments, the total number of first data that meets the first query condition in each segment can be obtained by multiplying the average size of the field to be queried for each segment by the number of rows of first data in the segment. Thereafter, the total number of first data that meets the first query condition in all segments in the log library is added together to obtain the size of the first data that meets the first query condition in the log library.

[0088] Finally, the aggregation time can be estimated based on the size of the data that meets the first query condition in the log library and the speed at which the query system processes data. The aggregation time is the query time.

[0089] For example, the following formula can be used to calculate the expected time for aggregation, i.e., the query time:

[0090]

[0091] Where f(x) is the query duration required to find the first data item in the log repository, n is the total number of segments in the log repository, and size(fi) is the average size of the queried field in the i-th segment. row(i) refers to the number of rows of the first data item in the i-th segment. IOPS represents the number of read and write operations performed by the query system per second, and BRSize represents the amount of data that can be read per unit time.

[0092] For example, continuing the above example, let's assume that the number of read and write operations performed by the query system per second is 2500 times, the size of the amount that can be read per unit time of the query system is 4KB / s, and the field to be queried is the addr field. Figure 5 (a) in the example. Expect Time = [20M * (1000M / 200M) + 40M * (4000M / 400M)] / (2500 * 4KB / s) = [100M + 400M] / (10M / s) = 50s. Thus, the calculated aggregation time (i.e., query time) is 50 seconds.

[0093] After obtaining the query duration, the query system may determine whether the query duration is greater than a preset threshold. If it is determined that the query duration is greater than the preset threshold, the following S303 is executed.

[0094] In addition, in some embodiments, after receiving the first query condition, the query system may first execute a query task in the log library based on the received first query condition before estimating the query time required to query the first data in the log library to determine whether the query task can be successfully executed. Then, the query system may perform the above-mentioned operation of estimating the query time required to query the first data in the log library if the first query task fails to execute. When the first query task is successfully executed, the query system may return the query result to the electronic device to complete the first query task. In this application, if the time taken by the query system to execute the first query task is greater than a preset threshold, that is, the query times out, it can be considered that the first query task has failed to execute; if the time taken by the query system to execute the first query task is less than or equal to the preset threshold, it can be considered that the first query task has been successfully executed.

[0095] S303: Adjust the first query condition according to a preset rule to obtain recommended information.

[0096] The recommendation information includes one or more of the following information: a recommended query time period, and a recommended range for searching data in the log library.

[0097] In some embodiments, when the recommendation information includes a recommended query time period, the query system may, based on a preset rule, adjust the query time period in the first query condition one or more times to obtain a second query condition. The time taken by the query system to execute the second query task based on the second query condition is less than or equal to a preset threshold. The query system may then determine the recommended information based on the second query condition.

[0098] For example, after estimating the query time required to query the first data in the log library, the query system can compare the query time with a preset threshold. If the query time exceeds the preset threshold, the query time period in the first query condition is adjusted according to a preset rule to obtain an adjusted query condition.

[0099] Based on the adjusted query conditions, the query system can re-estimate the query time required to retrieve data corresponding to the adjusted query conditions in the log repository. If the query time is still greater than a preset threshold, the query system can continue to adjust the query time period in the adjusted query conditions according to the preset rules to obtain a second adjusted query condition. Based on the second adjusted query conditions, the query time required to retrieve data corresponding to the second adjusted query conditions in the log repository is estimated. This adjustment cycle continues until the query time is less than or equal to the preset threshold, at which point the query conditions obtained from the last adjustment are used as the second query conditions. It will be understood that the query system can execute a second query task in the log repository based on the second query conditions and obtain a result indicating that the second query task was successfully executed. In other words, during any adjustment process, the query system can further adjust the query time period in the previously adjusted query conditions according to the preset rules to obtain a newly adjusted query condition. The query system can then estimate the query time required to retrieve data corresponding to the newly adjusted query conditions in the log repository based on the newly adjusted query conditions, and if the newly adjusted query time is less than or equal to the preset threshold, the newly adjusted query conditions are used as the second query conditions. It is understandable that, in order for the query system to successfully execute the AdHoc query task, the query system may use the second query condition as recommendation information, recommending that the user use the second query condition to query the log library for second data. The second data is part of the first data. The query system consumes less time to query the second data based on the second query condition than it does to query the first data based on the first query condition.

[0100] For example, continuing with the above example, let's take the preset rule of performing a binary search on the time period to be queried as an example. Combined with (a) in 5, when the field to be queried of the first query task is the addr field, and the time period to be queried is from February 21, 2023 to February 23, 2023, the query time is calculated to be 50 seconds, which is greater than the preset threshold of 30S, and the first query task times out. Perform a binary search on the time period to be queried, that is, select 1 / 2 of the time period to be queried, and obtain a new time period to be queried from February 22, 2023 to February 23, 2023. Figure 5As shown in (b), based on the new query time period from February 22, 2023 to February 23, 2023, the query system can filter segment0 and obtain a total of 10 million rows of data that meet the query time period "February 22, 2023 to February 23, 2023"; and filter segment1 to obtain a total of 20 million rows of data that meet the query time period. In this way, the aggregation time (i.e., query time) can be calculated by the above formula to be 25 seconds, which is less than the preset threshold of 30S. In this way, a new query condition (i.e., the second query condition) can be determined based on the new query time period, and recommended information can be determined based on the second query condition.

[0101] In addition, the preset rule may also be selecting 1 / 3 of the time period to be queried, selecting 1 / 4 of the time period to be queried, etc., and this application does not limit this.

[0102] When the recommendation information includes a recommended search range for data in the log library, the query system can calculate the number of rows of data in the log library that the query system can process within the expected time period based on the user's expected query time, and determine the recommendation information, i.e., the recommended search range for data in the log library, based on the number of rows that can be processed within the expected time period. This allows the user to re-determine new query conditions, i.e., the second query conditions, based on the recommended search range for data in the log library.

[0103] For example, continuing with the above example, let's assume the log repository contains approximately 100 million rows of log records, and the user expects an execution time of 20 seconds. According to the formula: [20s / 50s]*100 million = 40 million, we can calculate that the query system can process 40 million rows of data in the log repository in 20 seconds, which serves as the recommended information. It's understandable that based on this recommendation, the user can determine that the query time period corresponding to the 40 million rows of data in the log repository is from 10:20 on February 21, 2023, to 12:20 on February 23, 2023. In this case, the user can set the query time period "10:20 on February 21, 2023, to 12:20 on February 23, 2023" as the second query condition.

[0104] S304: Output recommendation information.

[0105] Wherein, in the case where the recommendation information is a recommended query time period, the recommendation information may further include: an estimated time consumption when querying using the second query condition.

[0106] In some embodiments, after obtaining the recommendation information, the query system may send the recommendation information to an electronic device. Accordingly, the electronic device, such as a computer, may receive the recommendation information sent by the query system and display it to the user for reference and modification of the query conditions.

[0107] For example, the electronic device may display the recommendation information in the form of a list or menu, etc., which is not limited in this application.

[0108] For example, the electronic device may also provide the user with recommendation information in a single-choice or multiple-choice selection format. This application does not limit this.

[0109] For example, the recommendation information displayed by the electronic device may include a recommended query time period. Alternatively, it may include a recommended search range for data in the log library. Alternatively, it may include a recommended query time period and the estimated query duration for the recommended query time period. This application does not limit the content of the recommendation information displayed by the electronic device.

[0110] For example, combining the above Figure 5 For example, the recommendation information obtained by the query system may include: the recommended query time period is: February 22, 2023 to February 23, 2023, the estimated query time using the recommended query time period is 25 seconds, and the recommended search range for data in the log library is no more than 40 million rows. Figure 6 As shown, the electronic device may display the following recommended information: Recommended information 1: Narrow the query time range. The time in the SQL input is 2023 / 02 / 22 to 2023 / 02 / 23. Based on the recommended query time range, the estimated execution time is 25 seconds. Recommended information 2: It is recommended to select a time period with no more than 40 million rows.

[0111] The above process is described below with reference to specific examples.

[0112] Combine Figure 7 For example, in an e-commerce related log query scenario, in order to count the regional distribution of product sales. Figure 7As shown in (a), according to the first query condition input by the user, the query system can receive the first query condition as request_method:get select remote_addr,count(*)group by remote_addr[time:2021 / 07 / 20 00:00~2021 / 07 / 22 00:00]. By parsing the first query condition, the query system can obtain: the identifier of the field to be queried of the first data is the addr field, and the time period to be queried is from 0:00 on July 20, 2021 to 0:00 on July 22, 2021. According to the first query condition, the query system can determine that the number of rows of the first data in segment0 that meet the time period to be queried is 20 million rows (i.e., 20M rows), and based on the identifier of the addr field, determine that the target total count of the data corresponding to the addr field is 200M and the target size is 1000M. It is determined that the number of rows of the first data in segment1 that meet the time period to be queried is 25 million rows, and based on the identifier of the addr field, the target total count of the data corresponding to the addr field is 400M and the target size is 4000M. It is determined that the number of rows of the first data in segment2 that meet the time period to be queried is 25 million rows, and based on the identifier of the addr field, the target total count of the data corresponding to the addr field is 400M and the target size is 4000M. In this way, it can be estimated that the query time is 55S, which is greater than the preset time threshold of 30S. Afterwards, the query system can cut the time period to be queried in half and re-estimate the query time. For example, when it is calculated that the query time in the new query time period is 20S, the recommended query time period is obtained as the recommended solution 1. In addition, the query system can also calculate the recommended number of rows of data in the search log library as 40 million rows as the recommended solution 2 based on the user's expectation to obtain the query results within 20S. Figure 7 As shown in (b) of FIG, two recommended solutions are displayed to the user, and a "Go to Help" option may also be displayed to the user for selection. It is understandable that no matter whether the user selects Recommendation 1 or Recommendation 2, the electronic device can return the query result within 20 seconds.

[0113] It should be noted that the above embodiments utilize the methods implemented in this application to query data within a query system, i.e., the query system is used as the execution entity for the purpose of illustration. In other embodiments, the query system may include a column-stored statistician, a query evaluator, and a front-end director, and the above implementation process may also be implemented by the column-stored statistician, the query evaluator, and the front-end director in coordination.

[0114] In some other embodiments of the present application, the data query method provided in the embodiment of the present application can also be applied to Figure 2In the query system shown in FIG. , the query system may include a storage statistician 201 , a query evaluator 202 , and a front-end director 203 .

[0115] For example, the storage statistics unit 201 may be used to execute S301 , the query evaluator 202 may be used to execute S302 and S303 , and the front-end director 203 may be used to execute S304 .

[0116] It should be noted that the specific implementation can refer to the description of the corresponding content in S301-S306. The only difference is that the devices performing the corresponding operations are different, and this embodiment will not be described in detail here.

[0117] The technical solution provided by the embodiment of the present application estimates the query time required to search for data corresponding to the query conditions in the log library, and adaptively outputs recommendation information based on the estimated results, thereby guiding users to obtain the data they want to query by adjusting the query conditions. In this way, not only the success rate of AdHoc queries is greatly improved, but also the response speed, query efficiency and user experience of the system are improved. In addition, the recommendation information output by the present application may include a recommended query time period, so that by guiding the user to obtain the expected results by only adjusting the query time range, the query efficiency can be optimized without affecting other query conditions. Moreover, the recommendation information output by the present application may include a recommended range for searching data in the log library, so that users can be given a higher degree of customization options to adapt to the time range query requirements in different scenarios, which significantly improves user satisfaction.

[0118] The above mainly introduces the scheme of the embodiment of the present application from the perspective of method. It is understandable that, in order to realize the above functions, the data query device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0119] The embodiment of the present application can divide the data query device into functional units according to the above method example. For example, each functional unit can be divided corresponding to each function, or two or more functions can be integrated into one processing unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. It should be noted that the division of units in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation.

[0120] The embodiment of the present application provides a data query device 80, which can be applied to the above electronic device. Figure 8 As shown, the data query device 80 includes: an acquisition unit 801 and a processing unit 802. Optionally, the data query device 80 further includes a display unit 803. The display unit 803 is used to display the data of the data query device 80.

[0121] The acquisition unit 801 is configured to receive a first query condition, where the first query condition is used to query first data in a log library.

[0122] Processing unit 802 is used to estimate the query time required to query the first data in the log library based on the first query condition; when the query time is greater than a preset threshold, output recommendation information, and the recommendation information is used to recommend the user to use the second query condition to query the second data in the log library, the second data is part of the first data, and the query time of the second data queried based on the second query condition is less than the query time of the first data queried based on the first query condition.

[0123] In one achievable method, the processing unit 802 is further used to estimate the query time required to query the first data in the log library based on the statistical information and the first query condition; wherein the log library stores data in the form of segments, and the statistical information includes: the identification of all segments included in the log library, the identification of all fields of the data included in each segment, the total count of the data corresponding to each field, and the size of the storage space occupied by the data corresponding to each field.

[0124] In one achievable method, the processing unit 802 is further configured to calculate the number of high-authority pixels included in the coding unit; wherein the high-authority pixels refer to pixels corresponding to high-authority objects; based on the number of high-authority pixels included in the coding unit, the maximum value of the high-authority pixels is calculated; the proportion of high-authority pixels in the coding unit is calculated based on the maximum value; and the authority of the coding unit is determined based on a preset threshold and the proportion of high-authority pixels in the coding unit.

[0125] In one achievable method, the first query condition includes: an identifier of the field to be queried of the first data and the time period to be queried, and the processing unit 802 is further used to obtain the number of rows of the first data in the log library that matches the time period to be queried; according to the identifier of the field to be queried, the target total count and target size are obtained from the statistical information, the target total count is the total count of the data corresponding to the field to be queried, and the target size is the size of the storage space occupied by the data corresponding to the field to be queried; the query time is estimated based on the number of rows of the first data, the target total count and the target size.

[0126] In one practicable manner, before estimating the query time required to query the first data in the log library based on the first query condition, the processing unit 802 is further configured to collect statistics of the data stored in the log library in segments to obtain statistical information.

[0127] In one achievable manner, the display unit 803 displays the recommended information including one or more of the following information: a recommended query time period, and a recommended range for searching data in the log library.

[0128] In one achievable manner, when the recommendation information includes a recommended query time period, the processing unit 802 is further used to adjust the time period to be queried in the first query condition one or more times based on preset rules before outputting the recommendation information to obtain a second query condition; determine the recommendation information based on the second query condition; wherein, during any adjustment process, the query condition after the last adjustment is adjusted again according to the preset rules to obtain the query condition after this adjustment; based on the query condition after this adjustment, estimate the query time required to query the data corresponding to the query condition after this adjustment in the log library; when the query time this time is less than the preset threshold, use the query condition after this adjustment as the second query condition.

[0129] In one achievable method, before estimating the query time required to query the first data in the log library based on the first query condition, the processing unit 802 is also used to execute a query task in the log library based on the first query condition; and determine that the duration of executing the query task is greater than a preset threshold.

[0130] Figure 8 The units in the can also be called modules, for example, the acquisition unit can be called an acquisition module, and the processing unit can be called a processing module. Figure 8 In the illustrated embodiment, the names of the various units may not be the names shown in the figure. For example, the acquisition unit may also be called a communication unit.

[0131] Figure 8If the various units in the embodiment of the present application are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling an electronic device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The storage medium for storing computer software products includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0132] The present application also provides a hardware structure diagram of a data query device, see Figure 9 The data query device includes a processor 901 and a transceiver 902 , and optionally, further includes a memory 903 connected to the processor 901 .

[0133] The processor 901 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application. The processor 901 may also include multiple CPUs, and the processor 901 may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. The processor here may refer to one or more devices, circuits, or processing cores for processing data (e.g., computer program instructions).

[0134] The processor 901, memory 903, and transceiver 902 are connected via a bus. Optionally, the transceiver 902 may include a transmitter and a receiver. The device in the transceiver 902 that implements the receiving function can be considered a receiver, and the receiver is used to perform the receiving steps in the embodiments of the present application. The device in the transceiver 902 that implements the transmitting function can be considered a transmitter, and the transmitter is used to perform the transmitting steps in the embodiments of the present application.

[0135] In the first possible implementation, see Figure 9, the encoding device also includes a memory 903. The memory 903 can be a ROM or other types of static storage devices that can store static information and instructions, a RAM or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, an optical disc storage (including a compressed optical disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of an instruction or data structure and can be accessed by a computer, and the present embodiment does not impose any restrictions on this. The memory 903 can be independent or integrated with the processor 901. Among them, the memory 903 can contain computer program code. The processor 901 is used to execute the computer program code stored in the memory 903, thereby realizing the method provided by the embodiment of the present application.

[0136] An embodiment of the present application also provides a computer-readable storage medium, comprising instructions, which, when executed on a computer, enables the computer to execute any of the above methods.

[0137] An embodiment of the present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute any of the above methods.

[0138] An embodiment of the present application further provides a chip, including: a processor and an interface, wherein the processor is coupled to a memory via the interface, and when the processor executes a computer program or instruction in the memory, any one of the methods provided in the above embodiments is executed.

[0139] An embodiment of the present application also provides a query system, including: the terminal device and access network device in the above embodiment.

[0140] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more media integrated therewith. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state drive (SSD)).

[0141] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprise" does not exclude other components or steps, and "a" or "an" does not exclude multiple components or steps. A single processor or other unit may implement several functions listed in the claims. The fact that certain measures are recorded in different dependent claims does not mean that these measures cannot be combined to produce good results.

[0142] Although the present application has been described in conjunction with features and embodiments thereof, it is apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present application. Accordingly, this specification and drawings are merely illustrative of the present application as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art may make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, the present application is intended to include such modifications and variations as fall within the scope of the claims of the present application and their equivalents.

[0143] The above is merely an embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application shall be included in the scope of protection of the present application. Therefore, the scope of protection of the present application shall be based on the scope of protection of the claims.

Claims

1. A data query method, characterized in that: include: Receive a first query condition, where the first query condition is used to query first data in a log library; estimating, based on the first query condition, a query time required to query the first data in the log library; When the query time is greater than a preset threshold, recommendation information is output, where the recommendation information is used to recommend that the user use a second query condition to query second data in the log library, where the second data is part of the first data, and the query time for querying the second data based on the second query condition is less than the query time for querying the first data based on the first query condition.

2. The method according to claim 1, characterized in that The estimating, based on the first query condition, the query time required to query the first data in the log library includes: estimating a query time required to query the first data in the log library based on the statistical information and the first query condition; In which, the log library stores data in the form of segments, and the statistical information includes: the identification of all segments included in the log library, the identification of all fields of the data included in each segment, the total count of data corresponding to each field, and the size of the storage space occupied by the data corresponding to each field.

3. The method according to claim 2, characterized in that The first query condition includes: an identifier of a field to be queried and a time period to be queried of the first data; The estimating, based on the statistical information and the first query condition, the query time required to query the first data in the log library includes: Obtain the number of rows of the first data in the log library that matches the time period to be queried; According to the identifier of the field to be queried, obtaining a target total count and a target size from the statistical information, wherein the target total count is the total count of data corresponding to the field to be queried, and the target size is the size of storage space occupied by the data corresponding to the field to be queried; The query duration is estimated according to the number of rows of the first data, the target total count, and the target size.

4. The method according to claim 2 or 3, characterized in that Before estimating the query time required to query the first data in the log library based on the first query condition, the method further includes: The data stored in the log library is counted in segments to obtain the statistical information.

5. The method according to any one of claims 1 to 4, characterized in that The recommendation information includes one or more of the following information: a recommended query time period, and a recommended range for searching data in the log library.

6. The method according to claim 5, characterized in that In the case where the recommendation information includes the recommended query time period, Before outputting the recommendation information, the method further includes: Based on a preset rule, adjusting the time period to be queried in the first query condition once or multiple times to obtain the second query condition; determining the recommended information based on the second query condition; Among them, during any adjustment process, the query conditions adjusted last time are adjusted again according to the preset rules to obtain the query conditions adjusted this time; based on the query conditions adjusted this time, the query time required to query the data corresponding to the query conditions adjusted this time in the log library is estimated. If the query time is less than the preset threshold, the query conditions adjusted this time are used as the second query conditions.

7. The method according to any one of claims 1 to 6, characterized in that Before estimating the query time required to query the first data in the log library based on the first query condition, the method further includes: executing a query task in the log library based on the first query condition; It is determined that the duration of executing the query task is greater than a preset threshold.

8. An electronic device, characterized in that: include: processor; The processor is connected to a memory, the memory is used to store computer-executable instructions, and the processor executes the computer-executable instructions stored in the memory, so that the electronic device implements the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that The method comprises instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 7.