Data query method and related apparatus
By using the aggregate query method and approximate estimation of partially accurate query results in the data query interface, combined with a single maximum query delay and data block sampling, the problem of rapid query and cost reduction in massive data is solved, and fast response and efficient query are achieved.
Patent Information
- Application Number
- PCT/CN2024/094293
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-05-20
- Publication Date
- 2025-05-08
AI Technical Summary
It is a difficult problem to quickly find target data in massive data and reduce query costs, especially in the context of the ever-expanding data volume.
By providing a data query interface, the query tasks entered by the user are obtained, and the aggregate query method is used to perform approximate estimation using some accurate query results to display the first estimated query results. At the same time, the query process is optimized by setting a single maximum query delay and sampling of data blocks.
It realizes rapid response to users, saves overhead, and provides available data in a short time, while improving user experience.
Smart Images

Figure CN2024094293_08052025_PF_FP_ABST
Abstract
Description
Data query method and related device
[0001] This application claims priority to Chinese patent application No. 202311426946.X, filed on October 30, 2023, with invention name “A data query method, device and computing device cluster”, and claims priority to Chinese patent application No. 202311861516.0, filed on December 29, 2023, with invention name “Data query method and related device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of cloud computing, and in particular to a data query method and related devices. Background Art
[0003] With the continuous development of information technology, various industries inevitably use computer equipment in their work processes, and the related data generated must also be stored accordingly. The amount of data involved by various users in various industries is constantly expanding with business expansion. In real life, when inspectors and operations personnel are looking for target data, they sometimes need to query massive amounts of data covering a period of several months or even a year. Faced with such a large amount of data, how to quickly find the target data and minimize query costs are urgent challenges.
[0004] Summary of the Invention
[0005] This application provides a data query method and related devices that can quickly find target data and minimize query costs. The technical solution is as follows:
[0006] In a first aspect, a data query method is provided, the method comprising: providing a data query interface, the data query interface being used to obtain a query task input by a user; obtaining an aggregate query task from the data query interface, the aggregate query task being used to query a total amount of data that meets a query condition, the query condition including a query term and a query time range; and displaying a first estimated query result of the aggregate query task in the data query interface, wherein the first estimated query result comprises a first total amount of data that meets the query condition obtained by approximate estimation based on a first portion of precise query results, the first portion of precise query results comprising an amount of data that meets the query condition queried from M data blocks, the M data blocks being sampled from N data blocks that match the aggregate query task, M and N being integers, and M being less than N.
[0007] For aggregated query tasks, approximate estimates are made based on partially accurate query results to obtain query results that meet the query criteria, and the estimated query results are then displayed in the data query interface. Compared to complete and accurate query results, approximate estimates based on partially accurate query results not only provide faster response times and facilitate quicker understanding of query result trends, but also reduce overhead. In other words, by using a sample-to-whole approach, an approximate estimate of the total amount of data that meets the query criteria based on the first portion of accurate query results can be quickly obtained, allowing users to obtain usable data in a shorter period of time.
[0008] Optionally, before displaying the first estimated query result of the aggregate query task in the data query interface, the method also includes: displaying the second estimated query result of the aggregate query task in the data query interface, wherein the second estimated query result includes a second total amount of data that meets the query conditions obtained by approximate estimation based on the second part of the precise query result, and the second part of the precise query result includes the amount of data that meets the query conditions queried from P data blocks, and the P data blocks are sampled from N data blocks matching the aggregate query task, and the above-mentioned M data blocks include the P data blocks; displaying the first estimated query result of the aggregate query task in the data query interface includes: replacing the second estimated query result with the first estimated query result in the data query interface.
[0009] For aggregated query tasks, multiple rounds of precise queries can be performed to obtain complete precise query results. The first set of precise query results consists of the results from the first i rounds of precise queries, while the second set consists of the results from the first i-1 rounds of precise queries. Each round adds new precise query results, and an approximate estimate of the amount of data that meets the query criteria is made based on all the precise query results obtained so far, and the result is displayed on the data query interface. This gradually increases the number of precise query results, making the approximate estimate more accurate.
[0010] Optionally, before replacing the second estimated query result with the first estimated query result in the data query interface, the method also includes: determining a newly sampled data block; based on the query conditions, querying the newly sampled data block to determine a new query result, the new query result including the amount of data that meets the query conditions queried from the newly sampled data block; determining the first part of the precise query result based on the new query result and the second part of the precise query result.
[0011] Each data block contains multiple data contents, and each data content corresponds to a data generation time. Therefore, when querying the newly sampled data blocks in the i-th round, all data contents in each data block can be traversed. When a certain data content contains the query term and the data generation time of the data content is within the query time range, the data content is determined to be data that meets the query conditions. In this way, the amount of data that meets the query conditions in the newly sampled data blocks in the i-th round can be obtained. After obtaining the amount of data that meets the query conditions in the newly sampled data blocks in the i-th round, the amount of data that meets the query conditions in the newly sampled data blocks in each of the previous i rounds is added together to obtain the amount of data that meets the query conditions in the M data blocks. Among them, the sum of the amount of data that meets the query conditions in the newly sampled data blocks in the previous i-1 rounds is the amount of data that meets the query conditions in the P data blocks.
[0012] Optionally, the method further includes displaying a complete and precise query result of the aggregate query task in a data query interface, where the complete and precise query result includes a total amount of data that meets the query conditions and is queried from the N data blocks.
[0013] For aggregate query tasks, each round makes an approximate estimate of the total amount of data that meets the query conditions based on the partially accurate query results. Through multiple iterations, a complete and accurate query result can be obtained, allowing the data query interface to progressively display increasingly accurate results to users.
[0014] Optionally, the first estimated query result further includes a first data statistical graph, where the first data statistical graph is used to indicate the distribution of the first total data volume in different time intervals.
[0015] By displaying the first data statistics chart in the data query interface, the temporal distribution of the amount of data that meets the query conditions can be intuitively displayed to the user, allowing the user to have a further understanding of the amount of data that meets the query conditions, giving the user a comfortable experience.
[0016] Optionally, the method further includes: displaying the query progress of the aggregate query task in the data query interface.
[0017] Optionally, the method further includes: determining N data blocks matching the aggregate query task from the target database based on the query condition and the index directory of the target database.
[0018] By using the index directory, N data blocks that match the aggregate query task can be efficiently filtered out. The filtered N data blocks can then be scanned to obtain the data content that meets the query conditions without traversing all the data in the target database. This can improve the efficiency of data query and reduce the consumption of cloud platform processing resources.
[0019] In some cases, determining the first portion of precise query results may take a relatively long time. To quickly respond to users, after receiving an aggregate query task, you can first perform a rough query on the target database based on the query conditions to obtain a rough query result. This rough query result includes the amount of data that meets the query conditions obtained through the rough query of the target database. In this way, before displaying the first estimated query result, you can first display the rough query result in the data query interface.
[0020] Because rough queries are very fast, rough query results can be quickly obtained and displayed on the data query interface, allowing users to quickly see the query results and quickly gain a general understanding of the amount of data that meets the query conditions. Moreover, in the above-mentioned aggregate query, rough queries and multiple rounds of precise queries are designed, and precise query results are obtained through multiple iterations. This allows the data query interface to quickly display rough query results to users while also gradually presenting more precise results to users, further enabling users to obtain usable data in a shorter time.
[0021] In a second aspect, another data query method is provided, which includes: providing a data query interface, which is used to obtain a query task input by a user; obtaining a content query task from the data query interface, which is used to query target data content that meets the query conditions, and the query conditions include query terms and query time range; based on the query conditions and the single query result return conditions, performing a first query operation, and displaying the first data content obtained by the first query operation in the data query interface, wherein the single query result return conditions include the single maximum query data volume and the single maximum query delay, the first query operation is used to query a part of the target data content, the query delay of the first query operation is not greater than the single maximum query delay and the data volume of the first data content is not greater than the single maximum query data volume.
[0022] The maximum amount of data in a single query returned in the single query result condition refers to the maximum amount of data content that meets the query conditions when each query operation is performed.
[0023] The maximum single query latency in the single query result return conditions refers to the maximum query time for each query operation. If the query operation time reaches the maximum single query latency, the query is terminated.
[0024] When the target database has a large amount of data, finding the maximum number of data that meets the query criteria can take a long time. If the number of retrieved data reaches the maximum number of data, or if the query results are not fed back to the user until all data has been traversed, the user experience can be significantly reduced. However, the maximum single query delay allows the response time of each query operation to be controlled within a certain period of time. This allows the cloud platform to provide users with timely feedback on the current query status of content query tasks, allowing users to be informed of the current query results in a timely manner, reducing the impact of excessive data volume on the user experience.
[0025] Optionally, based on the query conditions and the single query result return conditions, a first query operation is performed, and the first data content obtained through the first query operation is displayed in the data query interface, including: based on the query conditions and the single query result return conditions, a first query operation is performed; when any one of the multiple target conditions is met first, the first query operation is terminated, and the first data content obtained through the first query operation is displayed in the data query interface, the multiple target conditions including: the query delay of the first query operation reaches the single maximum query delay, the amount of data retrieved by the first query operation reaches the single maximum query data amount, and the first query operation has completed the query of the target data content.
[0026] For content query tasks, multiple query operations are often required to find all target data content. In this case, the first query operation can be any query operation. Of course, the cloud platform may also be able to find all target data content with only one query operation. In this case, this query operation is considered the first query operation.
[0027] Optionally, based on the query conditions and the single query result return conditions, a first query operation is performed, and the first data content obtained by the first query operation is displayed in the data query interface, including: determining the estimated total query delay of the content query task; based on the query conditions, the estimated total query delay and the single maximum query delay, the content query task is divided into multiple ordered subtasks, wherein each subtask is used to query a part of the target data content, and the estimated query delay of each subtask does not exceed the single maximum query delay; the first query operation is performed by executing at least one subtask; when any one of the multiple target conditions is met first, the first query operation is terminated, and the first data content obtained by the first query operation is displayed in the data query interface, the multiple target conditions include: the query delay of the first query operation reaches the single maximum query delay, the data volume of the first data content reaches the single maximum query data volume, and the target subtask is completed, wherein the target subtask is the last subtask among the multiple subtasks, or, when the target subtask is completed, the sum of the query delay of the first query operation and the estimated query delay of the next subtask of the target subtask is greater than the single maximum query delay.
[0028] Optionally, after the first query operation is completed, the method also includes: recording the pause position of the content query task; displaying a target query option in the data query interface, where the target query option is used to indicate to continue querying the data content; and performing a second query operation based on the pause position after receiving a user trigger operation on the target query option.
[0029] Optionally, after the first query operation is completed, the method further includes: determining the query progress of the content query task; and displaying the query progress of the content query task in the data query interface. This allows the user to promptly obtain the completion status of the current content query task, allowing the user to terminate the task early as needed, thereby saving time and reducing costs.
[0030] Optionally, the method also includes: obtaining a context query task triggered by the user regarding the benchmark data content, the context query task being used to query the target context data content that meets the context query conditions, the context query conditions including the previous time range, the next time range and the query statement, and the benchmark data content being any data content displayed on the data query interface; performing a first previous query operation and a first next query operation based on the context query conditions and the single query result return conditions, and displaying the first previous data content obtained by the first previous query operation and the first next data content obtained by the first next query operation in the data query interface, the first previous query operation and the first next query operation being used to query a part of the target context data content, the query delays of the first previous query operation and the first next query operation are both no greater than the single maximum query delay, and the data volumes of the first previous data content and the first next data content are both no greater than the single maximum query data volume.
[0031] After performing the above content query task, the user may trigger a context query task on a certain data content displayed in the data query interface. In this case, the method further includes the above content.
[0032] Optionally, the maximum single query data volume and / or the maximum single query delay are input by the user in the data query interface.
[0033] For content query tasks and context query tasks, by setting the maximum single query delay, the time it takes for each query operation to return results can be controlled within a certain period of time. By setting the conditions for returning single query results, the cloud platform can promptly return some ordered results in various situations, improving the user experience. The target query option provided by the data query interface allows users to initiate queries for more data after obtaining the current results, ensuring that users can obtain the data content they need. In addition, the data query interface displays the current task query progress, which helps users understand the current query status of the content query task and helps users determine whether to end the content query task based on their needs, thereby saving time and reducing expenses.
[0034] In a third aspect, a data query device is provided, wherein the data query device has the function of implementing the data query method described in the first aspect. The data query device includes at least one module, and the at least one module is used to implement the data query method described in the first aspect.
[0035] In a fourth aspect, another data query device is provided, which has the function of implementing the data query method in the second aspect. The data query device includes at least one module, which is used to implement the data query method provided in the second aspect.
[0036] In a fifth aspect, a computer device cluster is provided, wherein the computer device cluster includes at least one computer device, each computer device includes a processor and a memory, and the processor of the at least one computer device is used to execute instructions stored in the memory of the at least one computer device to implement the method described in the first aspect and / or the second aspect above.
[0037] Optionally, the computer device cluster may further include a communication bus, which is used to establish a connection between the processor and the memory of the at least one computer device.
[0038] In a sixth aspect, a computer-readable storage medium is provided, wherein the storage medium stores instructions. When the instructions are executed by a computing device cluster, the computing device cluster executes the steps of the data query method described in the first aspect and / or the second aspect.
[0039] In a seventh aspect, a computer program product comprising instructions is provided. When the instructions are executed by a computer device cluster, the computer device cluster executes the steps of the data query method described in the first and / or second aspects. Alternatively, a computer program is provided. When the computer program is executed on a computer, the computer executes the steps of the data query method described in the first and / or second aspects.
[0040] The technical effects obtained by the above-mentioned third and fourth aspects are similar to the technical effects obtained by the corresponding technical means in the first and second aspects, respectively. The technical effects obtained by the above-mentioned fifth, sixth and seventh aspects are similar to the technical effects obtained by the corresponding technical means in the first and / or second aspects, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] FIG1 is a flow chart of a data query method provided in an embodiment of the present application;
[0042] FIG2 is a schematic diagram of a data query interface including aggregated query task execution results provided by an embodiment of the present application;
[0043] FIG3 is a schematic diagram of an approximate estimation method provided in an embodiment of the present application;
[0044] FIG4 is a schematic diagram of another data query interface including the execution results of an aggregated query task provided in an embodiment of the present application;
[0045] FIG5 is a flow chart of another data query method provided in an embodiment of the present application;
[0046] FIG6 is a schematic diagram of a data query interface including content query task execution results provided by an embodiment of the present application;
[0047] 7 is a schematic diagram of another data query interface including content query task execution results provided by an embodiment of the present application;
[0048] FIG8 is a schematic diagram of a data query interface including a context query task execution result provided by an embodiment of the present application;
[0049] FIG9 is a schematic structural diagram of a data query device provided in an embodiment of the present application;
[0050] FIG10 is a schematic structural diagram of another data query device provided in an embodiment of the present application;
[0051] FIG11 is a schematic structural diagram of a computer device provided in an embodiment of the present application;
[0052] FIG12 is a schematic diagram of a computer device cluster provided in an embodiment of the present application;
[0053] FIG13 is a schematic diagram of one or more computer devices in a computer device cluster provided by an embodiment of the present application being connected via a network. DETAILED DESCRIPTION
[0054] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0055] Before explaining in detail the data query method provided in the embodiment of the present application, the application scenarios involved in the embodiment of the present application are first introduced.
[0056] With the continuous development of information technology, various industries inevitably use computer equipment in their work processes, and the related data generated is also stored accordingly. The amount of data involved by various users in various industries is constantly expanding with business expansion. In real life, when inspectors and operations personnel are searching for target data, they sometimes need to query massive amounts of data covering a period of several months or even a year. For example, with log data, inspectors may discover errors in a log during routine inspections. Since the inspectors may not know the time period when the error occurred, they may need to traverse log data over a longer period of time to obtain detailed data related to the log where the error occurred, thereby facilitating analysis of the root cause of the error. Alternatively, sometimes log data needs to be used to count the number of visits to a certain Internet Protocol (IP) address over a period of time. Since such logs are typically high-frequency logs (i.e., logs that appear frequently), it may be necessary to traverse a considerable amount of log data to obtain the number of visits to the IP address over a period of time. Faced with such a large amount of data, how to quickly find the target data and minimize query costs is a difficult problem that needs to be solved.
[0057] Related technology provides two data query methods. The first method is to query data by setting a maximum query data volume. That is, in the process of querying data, if the amount of data that meets the query conditions reaches the maximum query data volume, the query result is returned. If the amount of data that meets the query conditions does not reach the maximum query data volume, the traversal will continue until the maximum query data volume is reached and the query result is returned, or until the amount of data that meets the query conditions does not reach the maximum query data volume but all data has been traversed. The second method is to query data by setting a maximum query delay. That is, in the process of querying data, if the query delay reaches the maximum query delay, the query result is returned. If the query delay does not reach the maximum query delay, the traversal will continue.
[0058] The first approach can result in inaccurate query results and often takes a long time to return, resulting in a poor user experience. Furthermore, in scenarios where data is strongly time-ordered, such as contextual queries, a significant amount of data may need to be traversed to return the maximum query data volume. These products typically calculate overhead based on the amount of data traversed. However, in many cases, users can find the required data item or items by only partially meeting the query criteria. This approach can also implicitly increase user query costs. The second approach controls query time by setting a maximum query latency. When a query times out, the query is terminated. However, in scenarios with massive amounts of data, this approach can fail to obtain query results.
[0059] Based on this, an embodiment of the present application provides a data query method. In the method provided in the embodiment of the present application, the task that requires statistics on non-detailed data content such as IP address visits is called an aggregate query task, and the task that requires querying detailed content data such as log content is called a content query task. For aggregate query tasks, the embodiment of the present application performs approximate estimation based on partial precise query results to obtain query results that meet the query conditions, and then displays the estimated query results in the data query interface. Compared with complete precise query results, approximate estimation through partial precise query results can not only respond to users quickly and facilitate users to quickly understand the trend of query results, but also save overhead. For content query tasks, query operations are performed by setting a single query result return condition, and the query delay of the query operation is not greater than the single maximum query delay included in the single query result return condition, and the data volume of the data content obtained by the query operation is not greater than the single maximum query data volume included in the single query result return condition; at the same time, the pause position is also recorded and a target query option is provided, which is used to indicate the continuation of the query data content. In this way, after performing a query operation, the data content obtained through the query operation can be displayed in the data query interface in a timely manner, so that the user can end the task early after obtaining the desired query results based on the data content obtained by the query operation, saving time and expenses, and allowing the user to initiate multiple queries through the target query option to ensure that the data content required by the user can be obtained.
[0060] The execution entity of the data query method provided in the embodiment of the present application can be a cloud platform or an electronic device with data query function. The cloud platform can be a server cluster or distributed system composed of multiple physical servers, etc., or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or it can be a cloud computing service center.
[0061] Next, the data query method provided in the embodiment of the present application is explained in detail. For ease of description, the data query method provided in the embodiment of the present application is introduced with the cloud platform as the execution entity. It should also be noted that the data query method provided in the embodiment of the present application does not limit the type of data being queried, and can be logs, images, documents, etc.
[0062] FIG1 is a flow chart of a data query method provided by an embodiment of the present application. Referring to FIG1 , the method includes the following steps.
[0063] Step 101: providing a data query interface, where the data query interface is used to obtain a query task input by a user.
[0064] The data query interface may include multiple different areas, each of which is used to trigger different query tasks. Of course, the data query interface may also include multiple independent sub-interfaces, each of which is used to trigger different query tasks. This embodiment of the present application does not limit this.
[0065] Each query task has a corresponding target database. The target database is one or more databases that the user currently needs to query. These databases are typically stored on a cloud platform. The present embodiment does not limit the type of data stored in these databases. The data stored in the target database is stored in a data block as the smallest operation unit. Each data block has a corresponding time range, which is the range of the time when the data stored in the data block was generated.
[0066] In some embodiments, the target database is selected by the user, and the embodiments of the present application do not limit the method by which the user selects the target database. For example, the data query interface provided by the cloud platform includes a database selection area, and the user can select the target database through the database selection area. Of course, the database selection area can also be an independent sub-interface in the data query interface, or an interface independent of the data query interface. Alternatively, the cloud platform provides a command input window, and the user can enter relevant commands in the command input window to select the target database.
[0067] Step 102: Acquire an aggregate query task from the data query interface. The aggregate query task is used to query the total amount of data that meets the query conditions. The query conditions include query terms and query time range.
[0068] An aggregation query task counts the amount of data in the target database that meets the query criteria. For example, you might want to count the amount of data in the target database that contains the character "life," or count the number of times user A visited website B on November 1, 2023.
[0069] The data query interface may include multiple different areas, which are used to obtain different query parameters in the query conditions of the aggregate query task (such as query terms and query time range, etc.). Of course, the data query interface may also include multiple independent sub-interfaces, which are used to obtain different query parameters in the query conditions of the aggregate query task. The embodiments of the present application do not limit this.
[0070] For different application scenarios and the type of data queried by the aggregate query task, the query terms may be different. For example, if you currently need to query the reason why a customer failed to successfully pay for a takeout order, and the type of data queried by the aggregate query task is log, the query term can be the IP address of the electronic device used by the customer to pay for the takeout order, so as to facilitate the search for the log content within the time period when the customer paid for the takeout order, and analyze the reasons for the unsuccessful payment. For another example, if you currently need to count the number of errors reported on a service platform on a given day, and the type of data queried by the aggregate query task is log, the query term can be error, so as to facilitate the determination of the number of logs of all errors reported on the service platform on that day.
[0071] The query term can be a single query term or a combination of multiple query terms. For example, to query the number of errors reported by a server that day, the query term could be a combination of the server's IP address and "error." The query term can be entered by the user in the data query interface, selected from multiple query terms provided in the data query interface, or obtained through other means.
[0072] The query time range is determined by the time information entered or selected by the user in the data query interface. The time information entered or selected by the user in the data query interface can be a relative time length or an absolute time range, which is not limited in this embodiment of the present application.
[0073] The relative time length is the length of time from the start time of the query operation. The query time range is the time range from the relative time to the start time of the query operation. The relative time is the time before the start time of the query operation and the time from the start time of the query operation is the relative time length. For example, if the relative time length is 1 hour and the start time of the query operation is 20:30 on November 23, 2023, then the relative time is 19:30 on November 23, 2023, and the query time range is from 19:30 on November 23, 2023 to 20:30 on November 23, 2023. The absolute time range is the time range from the first absolute time to the second absolute time, where the first absolute time and the second absolute time are both times entered or selected by the user.
[0074] The relative time length can be input by the user in the data query interface, or can be selected by the user from multiple relative time lengths provided in the data query interface. For example, the data query interface provides a relative time length menu that includes multiple relative time lengths, and the user can select a relative time length based on the time range required for the aggregate query task.
[0075] The first absolute time and the second absolute time may be input by the user in the data query interface, or may be selected by the user from absolute times provided in the data query interface. For example, the data query interface may further include an absolute time input box, into which the user may directly input the first absolute time and the second absolute time. Alternatively, the data query interface may further include an absolute time calendar page, through which the user may select the first absolute time and the second absolute time.
[0076] Of course, the above is just an example. In actual applications, the data query interface can also provide other types of operation areas for users to input or select relative time length, first absolute time and second absolute time. The embodiment of the present application does not limit this.
[0077] After the cloud platform obtains the aggregate query task through the data query interface, the aggregate query task can be executed immediately, or the aggregate query task can be executed within the first time period after obtaining the aggregate query task, if it is determined that the user has not modified the query conditions. Of course, the aggregate query task can also be executed after receiving the query instruction issued by the user. Among them, the size of the first time period can be set according to the actual application scenario. The first time period can be pre-set by the cloud platform or set by the user. The embodiment of the present application does not limit the size and setting method of the first time period. There are many ways for users to issue query instructions. For example, when the cloud platform provides an instruction input window, users can issue query instructions through the instruction input window. When the data query interface also includes a query analysis button, users can complete the issuance of query instructions by clicking the query analysis button. The embodiment of the present application does not limit the way users issue query instructions.
[0078] For example, the data query interface is shown in Figure 2, where area A instructs the user to enter a query word, area B instructs the user to select a relative time length, area C is a query analysis button for issuing query instructions, and area D is used to display query results.
[0079] Step 103: Display the first estimated query result of the aggregate query task in the data query interface, wherein the first estimated query result includes a first total amount of data that meets the query conditions obtained by approximate estimation based on the first part of the precise query results, and the first part of the precise query results includes the amount of data that meets the query conditions queried from M data blocks, and the M data blocks are sampled from N data blocks that match the aggregate query task, M and N are integers, and M is less than N.
[0080] In some embodiments, the first estimated query result of the aggregate query task may be determined according to the following steps (1)-(3).
[0081] (1) M data blocks are sampled from N data blocks that match the aggregation query task.
[0082] For aggregate query tasks, the cloud platform can obtain complete and precise query results through multiple rounds of precise queries. In this case, the first partial precise query result can be the results obtained from the first i rounds of precise queries. Of course, the cloud platform does not need to obtain complete and precise query results, but only needs to approximate a query result through a single round of precise queries. In this case, the first partial precise query result is the result obtained from that round of precise queries. The implementation method for sampling M data blocks from N data blocks varies in different scenarios, so each will be described below.
[0083] In the first case, the cloud platform obtains complete and accurate query results through multiple rounds of accurate query. The first part of the accurate query results is the results of the first i rounds of accurate query. In this case, the newly sampled data block in round i is determined from the N data blocks. The newly sampled data block in round i and P data blocks are combined to form the M data blocks. The P data blocks are the confluence of the newly sampled data blocks in each of the previous i-1 rounds.
[0084] In some embodiments, any sampling algorithm can be used to determine the newly sampled data blocks in the i-th round from the N data blocks. The sampling algorithm can be uniform sampling, exponential sampling, etc. In practical applications, a suitable sampling algorithm can be selected according to needs, and the embodiments of the present application do not limit this. For example, if you do not want to affect the overall performance, you can use uniform sampling to make the number of newly sampled data blocks in each round the same, but this may take slightly longer to obtain complete and accurate query results; if you do not mind slightly affecting the overall performance, you can use exponential sampling to make the number of newly sampled data blocks in each round grow explosively, so that you can quickly obtain complete and accurate query results.
[0085] In some embodiments, the sampling algorithm can be selected by the user. For example, when the data query interface provides an algorithm selection menu, the user can select a sampling algorithm through the algorithm selection menu, and the cloud platform will determine the newly sampled data blocks in the i-th round based on the sampling algorithm selected by the user.
[0086] In the second case, the cloud platform uses a round of precise query to approximate the query results. The first part of the precise query results is the result obtained from this round of precise query. In this case, a portion of the N data blocks is sampled and determined as the M data blocks.
[0087] In some embodiments, a portion of the data blocks can be sampled from the N data blocks using any sampling algorithm, such as randomly selecting a portion of the data blocks, which is not limited in this embodiment of the present application.
[0088] In some embodiments, before sampling M data blocks from N data blocks matching the aggregate query task, the N data blocks may be first determined from the target database.
[0089] Since the target database typically contains multiple data blocks, each data block contains multiple data contents, each data content corresponds to a data generation time, and the target database's index directory typically records the content summary of each data block and the range of data generation time, it is possible to determine N data blocks that match the aggregate query task from the target database based on the query conditions and the target database's index directory. The data content stored in each of these N data blocks matches the query term, and the data generation time of the data content stored in each data block matches the query time range. In other words, the content summary of the data content stored in these N data blocks matches the query term, and the range of data generation time of the data content stored in these N data blocks at least overlaps with the query time range.
[0090] The index directory is a structure that sorts the values of one or more columns in a database (databases are usually stored in table form). Each database can establish its own corresponding index directory, which can be used to quickly access specific information in the database. For example, a query task needs to find the data content with an identity identification code (ID) of 44 in database D. If there is no index directory, the entire database must be traversed until the row with ID equal to 44 is found. However, if there is an index directory and it is established for the ID column, the index directory can be directly searched for 44 to determine the location of this row. In other words, if this row is found, the data content with ID 44 can be obtained.
[0091] By using the index directory, N data blocks that match the aggregate query task can be efficiently filtered out. The filtered N data blocks can then be scanned to obtain the data content that meets the query conditions without traversing all the data in the target database. This can improve the efficiency of data query and reduce the consumption of cloud platform processing resources.
[0092] (2) Determine a first portion of accurate query results, where the first portion of accurate query results includes the amount of data found in the M data blocks that meets the query condition.
[0093] Based on the description above, for aggregate query tasks, the cloud platform can obtain a complete and precise query result through multiple rounds of precise queries. In this case, the first part of the precise query result is the result obtained by the first i rounds of precise queries in these multiple rounds. Of course, the cloud platform does not need to obtain a complete and precise query result, but only needs to approximate a query result through a single round of precise queries. In this case, the first part of the precise query result is the result obtained by this round of precise queries. The implementation method for determining the first part of the precise query result varies in different situations, so we will explain each of these methods below.
[0094] In the first scenario, the cloud platform obtains complete and precise query results through multiple rounds of precise queries. The first part of the precise query results is the results from the first i rounds of precise queries. Based on the query conditions, the newly sampled data blocks from round i are queried to determine the new query results. The new query results include the amount of data that meets the query conditions found in the newly sampled data blocks. The first part of the precise query results is determined based on the new query results and the second part of the precise query results. The second part of the precise query results includes the amount of data that meets the query conditions found in the P data blocks.
[0095] As can be seen from the above description, each data block contains multiple data contents, and each data content corresponds to a data generation time. Therefore, when querying the newly sampled data blocks in the i-th round, all data contents in each data block can be traversed. When a certain data content contains the query term and the data generation time of the data content is within the query time range, the data content is determined to be data that meets the query conditions. In this way, the amount of data that meets the query conditions in the newly sampled data blocks in the i-th round can be obtained. After obtaining the amount of data that meets the query conditions in the newly sampled data blocks in the i-th round, the amount of data that meets the query conditions in the newly sampled data blocks in each of the previous i rounds is added together to obtain the amount of data that meets the query conditions in the M data blocks. Among them, the sum of the amount of data that meets the query conditions in the newly sampled data blocks in the previous i-1 rounds is the amount of data that meets the query conditions in the P data blocks.
[0096] Because each data content in a data block corresponds to a data generation time, when a query is performed on a newly sampled data block based on the query criteria, the data generation time of the data content in the newly sampled data block that meets the query criteria can also be obtained. Thus, the newly added query result also includes the data generation time of the data content found in the newly sampled data block that meets the query criteria. Similarly, the second part of the precise query result also includes the data generation time of the data content found in the P data blocks that meet the query criteria. Therefore, the first part of the precise query result can also include the data generation time of the data content found in the M data blocks that meet the query criteria.
[0097] In the second case, the cloud platform uses a round of precise query to approximate the query result. The first part of the precise query result is the result obtained by this round of precise query. In this case, based on the query conditions, the M data blocks are queried to determine the first part of the precise query result.
[0098] In some embodiments, all data contents in the M data blocks are traversed. When a certain data content contains a query word and the data generation time of the data content is within the query time range, the data content is determined to be data that meets the query conditions. In this way, the amount of all data that meet the query conditions in the M data blocks can be obtained.
[0099] Because each data content in a data block corresponds to a data generation time, when querying the M data blocks based on the query criteria, the data generation time of the data content in the M data blocks that meets the query criteria can also be obtained. In this way, the first part of the precise query results can also include the data generation time of the data content that meets the query criteria found in the M data blocks.
[0100] (3) Based on the first part of the precise query results, an approximate estimate is made to obtain a first estimated query result.
[0101] Based on the above description, for aggregate query tasks, the cloud platform can obtain complete and precise query results through multiple rounds of precise queries. Thus, the first partial precise query results are the results obtained from the first i rounds of precise queries in these multiple rounds. Of course, the cloud platform does not need to obtain complete and precise query results, and can simply use one round of precise queries to approximate a query result. In this case, the first partial precise query results include the amount of data matching the query criteria found in the M data blocks. Therefore, in the first case, when the i-th round is the last round, the amount of data matching the query criteria found in the first i rounds is the total amount of data matching the query criteria in the N data blocks, i.e., the first total data amount. If the i-th round is not the last round, an approximate estimate can be made based on the first partial precise query results to obtain the first total data amount matching the query criteria. In the second case, an approximate estimate can be made directly based on the first partial precise query results to obtain the first total data amount matching the query criteria.
[0102] The implementation process of approximate estimation based on the precise query results of the first part to obtain the first total data volume that meets the query conditions includes: dividing the data volume that meets the query conditions queried from the M data blocks by M to obtain a first estimated average value; multiplying the first estimated average value by N to obtain the total data volume that meets the query conditions in the N data blocks, that is, the first total data volume.
[0103] Of course, the first total data volume can also be approximately estimated by other methods, and this embodiment of the present application does not limit this.
[0104] The above approximate estimation process is explained below with reference to Figure 3. Referring to Figure 3, assume that there are three data sets in the target database, and each data set has 9 data blocks that match the aggregation query task. There are a total of 27 data blocks that match the aggregation query task in the three data sets. The data generation time ranges of the three data sets are the same, but for any data set, the data generation time ranges of each data block in the data set do not overlap. Assume that in the current aggregation query task, 3 data blocks that match the query task have been taken out of the three data sets in a streaming reading manner and scanned, that is, 9 data blocks have been extracted from the 27 data blocks that match the aggregation query task and scanned, and the scan results have been aggregated and cached. Assuming that the amount of data that meets the query conditions in each scanned data block is 1000, the amount of data that meets the query conditions in the first part of the precise query results is 9000, so the first estimated average value is 9000 / 9=1000, so it can be estimated that the total amount of data that meets the query conditions in the 27 data blocks that match the aggregate query task is 1000×27=27000, that is, the first total data amount is approximately estimated to be 27000.
[0105] This method uses the precise results from scanning a subset of data blocks to approximate the total amount of data that meets the query criteria. Because only a subset of data blocks is scanned, this method significantly reduces query time compared to scanning all matching data blocks for the aggregate query, allowing users to quickly gain a rough understanding of the amount of data that meets the query criteria.
[0106] Based on the above description, the first part of the precise query results also includes the data generation time of the data content in the M data blocks that meets the query conditions. Therefore, in some embodiments, a first data statistical graph can also be generated based on the data generation time of the data content in the M data blocks that meets the query conditions, as well as the amount of data in the M data blocks that meets the query conditions. That is, in addition to the first total data volume, the first estimated query result also includes a first data statistical graph, which is used to indicate the distribution of the first total data volume in different time intervals. The first data statistical graph includes, but is not limited to, at least one of the following: an interval histogram, a broken line statistical graph, a pie chart, etc.
[0107] Among them, each time interval in the first data statistics chart can be obtained by dividing the query time range through a time interval division mechanism, and the time interval division mechanism can be pre-set by the cloud platform or set by the user. For example, the time interval division mechanism pre-set by the cloud platform is: dividing the query time range into 10 intervals on average; or, the data query interface also includes a time interval division mechanism setting area, and the time interval division mechanism set by the user through this area is: dividing the query time range into intervals of 3 minutes. Of course, these two cases are only examples, and the time interval division mechanism can be set according to actual application requirements. In addition, it should be noted that the length of each time interval obtained by division is usually not less than the length of the time range corresponding to each data block.
[0108] When generating the first data statistics graph, for the M data blocks that have been traversed, the amount of data that meets the query conditions in these data blocks and the data generation time of each data content that meets the query conditions can be obtained. Therefore, based on the data generation time of the data content that meets the query conditions in these data blocks, the amount of data in different time intervals can be directly counted; and for the N data blocks that have not been traversed that match the aggregate query task, the amount of data that meets the query conditions in each data block can be estimated. At the same time, the time range corresponding to each data block can be obtained from the metadata corresponding to each data block. This time range is the range of the data generation time of the data content in each data block. Therefore, the amount of data in different time intervals in the data blocks that have not been traversed can also be counted. Among them, the metadata of each data block contains basic information about the data block, such as the range of the data generation time of the data content in the data block, the type and amount of data contained in the data block, etc.
[0109] After obtaining the first estimated query result in the above manner, the first estimated query result can be displayed in the data query result. Furthermore, if the first estimated query result includes a first data statistical graph, displaying the first data statistical graph in the data query interface can intuitively display to the user the temporal distribution of the amount of data that meets the query criteria, allowing the user to have a deeper understanding of the amount of data that meets the query criteria and providing a comfortable user experience.
[0110] In the case where the first portion of the precise query results is the result obtained from the first i rounds of precise queries in multiple rounds of precise queries, before displaying the first estimated query result in the data query interface, the second estimated query result of the aggregate query task may also be displayed in the data query interface, wherein the second estimated query result includes a second total amount of data that meets the query conditions, obtained by approximating the second portion of the precise query results, and the second portion of the precise query result includes the amount of data that meets the query conditions queried from P data blocks, where the P data blocks are sampled from N data blocks that match the aggregate query task, and the M data blocks include P data blocks. In this case, when displaying the first estimated query result of the aggregate query task in the data query interface, the second estimated query result displayed in the data query interface may be replaced with the first estimated query result.
[0111] Among them, the process of determining the precise query results of the second part, and the process of making an approximate estimate based on the precise query results of the second part to obtain the second total data volume that meets the query conditions, are similar to the above-mentioned process of determining the precise query results of the first part, and the process of making an approximate estimate based on the precise query results of the first part to obtain the first total data volume that meets the query conditions, and will not be repeated here.
[0112] Similarly, in some embodiments, in addition to the second total data volume, the second estimated query result also includes a second data statistical graph indicating the distribution of the second total data volume over different time intervals. The method for determining the second data statistical graph is similar to the method for determining the first data statistical graph, and is not further described here.
[0113] To summarize, for aggregated query tasks, after determining the N data blocks that match the query, multiple rounds of sampling can be performed on these N data blocks. Each round generates new precise query results. Each round approximates the amount of data that meets the query criteria based on all the precise query results obtained so far, and displays the results on the data query interface. This gradually increases the number of precise query results, making the approximate estimate more accurate.
[0114] In some embodiments, when displaying the first estimated query result, the query progress of the aggregate query task can also be displayed in the data query interface. For example, the query progress of the aggregate query task is obtained by dividing the above M by N.
[0115] By displaying the query progress of the aggregate query task on the data query interface, users can clearly understand the query status of the current task, so that users can control the query task in a timely manner and end the task in advance when the required results are obtained, avoiding excessive overhead.
[0116] In some embodiments, when the first estimated query result of the aggregate query task is displayed in the data query interface, the first query result description can also be displayed in the data query interface. The first query result description can be "The current result is an estimated result" or the like, to prompt the user that the query result displayed on the current data query interface is an estimated result, not an accurate result.
[0117] Continuing with the above example, please refer to Figure 2. The query result shown in Figure 2 is the first estimated query result. At this time, the N data blocks matching the aggregate query task have not been traversed. The query result is an approximate estimate. The progress of the query progress bar is approximately 10%. The first total data volume is 9886605. The first data statistical chart corresponding to the first total data volume is the interval histogram in Figure 2. The interval histogram uses 3 minutes as an interval interval to show the distribution of the first total data volume over time.
[0118] In some embodiments, the data query interface also provides a pause button and a continue button. In this case, if the user, after seeing the query results of a certain round, determines that he no longer needs more accurate query results based on his needs, he can trigger the pause task operation through the pause button. The cloud platform responds to the pause task operation triggered by the user and does not perform the next round of query; if the user, after seeing the query results of a certain round, determines that he still needs more accurate query results based on his needs, he can trigger the continue task operation through the continue button. The cloud platform responds to the continue task operation triggered by the user and performs the next round of query. Of course, in other embodiments, after displaying the query results of a certain round in the data query interface, the cloud platform can directly perform the next round of query.
[0119] In some embodiments, the complete and precise query results of the aggregate query task can also be displayed in the data query interface. The complete and precise query results include the total amount of data that meets the query criteria queried from the above N data blocks. In this case, all N data blocks matching the aggregate query task have been traversed.
[0120] In some embodiments, when displaying complete and accurate query results, the data query interface also displays a second query result description, which may be "The current result is an accurate result" or the like, to prompt the user that the query result currently displayed on the data query interface is an accurate query result.
[0121] Similarly, in some embodiments, in addition to displaying complete and accurate query results, a third data statistical graph may also be displayed in the data query interface. This third data statistical graph is used to indicate the distribution of the total amount of data in the N data blocks that meets the query criteria over different time intervals. The method for determining this third data statistical graph is similar to the method for determining the first and second data statistical graphs described above and will not be further described here.
[0122] Continuing with the above example, please refer to Figure 4. The query result shown in Figure 4 is a complete and accurate query result. The current task query progress is 100%. The N data blocks matching the aggregate query task have all been traversed. The total amount of data that meets the query conditions is 10458243. The third data statistical chart corresponding to this total number is the interval histogram in Figure 4.
[0123] In some cases, determining the first portion of precise query results may take a relatively long time. To quickly respond to users, after receiving an aggregate query task, you can first perform a rough query on the target database based on the query conditions to obtain a rough query result. This rough query result includes the amount of data that meets the query conditions obtained through the rough query of the target database. In this way, before displaying the first estimated query result, you can first display the rough query result in the data query interface.
[0124] When performing a rough query on the target database, it can be implemented through the HyperLogLog algorithm, or through the pre-aggregation algorithm, or through other similar algorithms. As long as the rough query results can be obtained quickly, the embodiments of the present application do not limit this.
[0125] Optionally, the rough query result may further include a rough data statistics graph, where the rough data statistics graph is used to indicate the distribution of the amount of data meeting the query conditions obtained by performing a rough query on the target database in different time intervals.
[0126] Since the rough query is very fast, the rough query result can be quickly obtained and displayed on the data query interface, so that the user can quickly see the query result and quickly have a general understanding of the number of data that meet the query conditions.
[0127] In some embodiments, the data query interface also provides a pause button and a continue button. In this case, if the user only needs to understand the rough query results, after seeing the rough query results, the pause task operation can be triggered by the pause button. The cloud platform responds to the pause task operation triggered by the user and does not proceed to subsequent steps; if the user sees the rough query results and judges that a more accurate query result is still needed based on the current application scenario, the continuation task operation can be triggered by the continue button. The cloud platform responds to the continue task operation triggered by the user and proceeds to subsequent steps.
[0128] Of course, in other embodiments, after displaying the rough query results in the data query interface, the cloud platform can directly proceed to the subsequent steps, or after the duration of displaying the rough query results in the data query interface reaches a second duration, the cloud platform automatically performs multiple rounds of precise queries on the target database. The second duration can be set according to the actual application scenario, and the second duration can be pre-set by the cloud platform or set by the user. The embodiments of this application do not limit the size and setting method of the second duration.
[0129] In some embodiments, the data query interface further displays a third query result description, such as "the current result is a rough query result", etc., to remind the user that the query result currently displayed on the data query interface is a rough query result.
[0130] Since aggregate query tasks do not require high accuracy in query results but require high efficiency in feeding query results back to users, for aggregate query tasks, a sample-to-whole approach is used. Each round of the query approximates the total amount of data that meets the query criteria based on partially accurate query results. Through multiple iterations, complete and accurate query results are obtained, allowing the data query interface to progressively display increasingly accurate results to users. Since each round only requires traversing a portion of the data blocks, the response time for each query round is very short, allowing users to obtain usable data in a relatively short period of time. Furthermore, the current task query progress is displayed on the data query interface, making it easier for users to understand the completion status of the current query task. Furthermore, users can trigger the task pause operation in various ways, allowing users to terminate the query task promptly after a certain round of query based on the data accuracy requirements of the actual application, thereby saving query time and reducing query costs. Furthermore, in the above-mentioned aggregate query, rough queries and multiple rounds of precise queries are designed, and precise query results are obtained through multiple iterations, so that the data query interface can not only quickly display rough query results to users, but also gradually display more precise results to users, further enabling users to obtain available data in a shorter time.
[0131] FIG5 is a flow chart of another data query method provided in an embodiment of the present application. Referring to FIG5 , the method includes the following steps.
[0132] Step 501: providing a data query interface, where the data query interface is used to obtain a query task input by a user.
[0133] This step is the same as the content in the above step 101. Please refer to the above description for detailed implementation content, which will not be repeated here.
[0134] Step 502: Obtain a content query task from the data query interface. The content query task is used to query target data content that meets query conditions. The query conditions include query terms and query time range.
[0135] Content query tasks are tasks that require detailed data, such as log content. For example, if you need to find all error logs in the target database and analyze the causes of the errors, this query task is a content query task. The error log contents need to be displayed in the data query interface to facilitate error cause analysis.
[0136] It should be noted that the data that the content query task needs to query is strongly time-ordered, that is, the data that the content query task needs to query is sorted in sequence according to the chronological order of data generation time. For multiple data with the same data generation time, they are sorted and scanned according to the storage order of the multiple data.
[0137] The target data content refers to all data content in the target database that meets the query conditions.
[0138] The method for obtaining the query conditions of the content query task is similar to the method for obtaining the query conditions of the above-mentioned aggregation query task. Please refer to the above description for detailed implementation content, which will not be repeated here.
[0139] After the cloud platform obtains the content query task through the data query interface, the content query task can be executed immediately, or it can be executed within the third time period after obtaining the content query task, if it is determined that the user has not modified the query conditions. Of course, the content query task can also be executed when a query instruction triggered by the user is received. Among them, the size of the third time period can be set according to the actual application scenario. The third time period can be pre-set by the cloud platform or set by the user. The embodiment of the present application does not limit the size and setting method of the third time period. There are many ways for users to trigger query instructions. For example, when the cloud platform provides an instruction input window, users can issue query instructions through the instruction input window. When the data query interface also includes a query analysis button, users can complete the issuance of query instructions by clicking the query analysis button. The embodiment of the present application does not limit the way users issue query instructions.
[0140] Step 503: Based on the query conditions and the single query result return conditions, a first query operation is performed, and the first data content obtained through the first query operation is displayed in the data query interface, wherein the single query result return conditions include the single maximum query data volume and the single maximum query delay. The first query operation is used to query a part of the target data content, the query delay of the first query operation is not greater than the single maximum query delay and the data volume of the first data content is not greater than the single maximum query data volume.
[0141] In content query tasks, due to the existence of conditions for returning single query results, only one query operation can often only retrieve part of the data content that meets the query conditions, that is, part of the target data content. If all data content that meets the query conditions is to be queried, multiple query operations are often required.
[0142] The maximum single query data volume in the single query result return criteria refers to the maximum number of data contents that meet the query criteria during each query operation. As described above, content query tasks require sequential data. Therefore, each query operation requires traversing the data sequentially, either forward or backward. When the number of data that meets the query criteria reaches the maximum single query data volume, the data scanning stops.
[0143] The maximum single query latency in the single query result return conditions refers to the maximum query time for each query operation. If the query operation time reaches the maximum single query latency, the query is terminated.
[0144] When the target database has a large amount of data, finding the maximum number of data that meets the query criteria can take a long time. If the number of retrieved data reaches the maximum number of data, or if the query results are not fed back to the user until all data has been traversed, the user experience can be significantly reduced. However, the maximum single query delay allows the response time of each query operation to be controlled within a certain period of time. This allows the cloud platform to provide users with timely feedback on the current query status of content query tasks, allowing users to be informed of the current query results in a timely manner, reducing the impact of excessive data volume on the user experience.
[0145] Based on the above description, the cloud platform often finds all target data content through multiple query operations. In this case, the first query operation can be any query operation. If the first query operation is the first query operation, the first query operation starts from the earliest time or the latest time of the query time range of the content query task, and queries the data content forward or backward in chronological order; if the first query operation is not the first query operation, the first query operation starts from the pause position of the last query operation. Of course, the cloud platform may also find all target data content through only one query operation. In this case, the query operation is regarded as the first query operation, and the first query operation also starts from the earliest time or the latest time of the query time range of the content query task, and queries the data content forward or backward in chronological order.
[0146] In some embodiments, a first query operation is performed based on the query conditions and the single query result return conditions, and the implementation process of displaying the first data content obtained through the first query operation in the data query interface includes: performing the first query operation based on the query conditions and the single query result return conditions; when any one of multiple target conditions is met first, the first query operation is terminated, and the first data content obtained through the first query operation is displayed in the data query interface, and the multiple target conditions include: the query delay of the first query operation reaches the single maximum query delay, the amount of data retrieved by the first query operation reaches the single maximum query data amount, and the first query operation has completed the query of the target data content.
[0147] That is to say, if the first query operation has not completed the query of the target data content, the amount of data found by the first query operation has not reached the maximum single query data amount, but the query delay of the first query operation has reached the maximum single query delay, or the first query operation has not completed the query of the target data content, the query delay of the first query operation has not reached the maximum single query delay, but the amount of data found by the first query operation has reached the maximum single query data amount, then the first query operation is terminated; if the query delay of the first query operation has not reached the maximum single query delay, the amount of data found by the first query operation has not reached the maximum single query data amount, but the first query operation has found the last data that needs to be scanned in chronological order, then it means that the target data content has been completely checked, and therefore, the first query operation is also terminated.
[0148] In some embodiments, when the cloud platform receives a first query request triggered by a user, the cloud platform performs a first query operation. The user may trigger the query request by pressing the Enter key, clicking a query button if the data query interface includes a query button, or other methods, which are not limited in the embodiments of the present application.
[0149] In other embodiments, when it is detected that the query condition has not been updated within a fourth time period, the first query operation is performed. The fourth time period may be pre-set by the cloud platform or set by the user. The size of the fourth time period may be set according to actual application requirements. The embodiments of the present application do not limit the setting method and size of the fourth time period.
[0150] In some embodiments, if the first query operation is the first query operation, then when performing the first query operation, Q data blocks in the target database that match the content query task are first determined. In this case, each query operation scans within these Q data blocks, rather than scanning all data blocks within the query time range. This reduces the amount of data scanned individually by the content query task, shortens the query time, and enables faster return of query results. The method for determining these Q data blocks is similar to the method for determining the N data blocks that match the aggregate query task in step 103 above, and will not be further described here.
[0151] In other embodiments, the content query task may be split into multiple subtasks. In this case, based on the query conditions and the single query result return conditions, a first query operation is performed, and the first data content obtained by the first query operation is displayed in the data query interface, including the following steps (1)-(4):
[0152] (1) Determine the estimated total query latency of the content query task.
[0153] In some embodiments, based on metadata of data blocks in the target database within the query time range, the amount of data within the query time range, ie, the amount of pre-queried data, is determined, and the estimated total query latency is determined based on the amount of pre-queried data.
[0154] Among them, the estimated total query delay refers to the time required to scan the pre-checked data volume one by one. When determining the estimated total query delay, the average time required to scan one data can be multiplied by the pre-checked data volume to obtain the estimated total query delay. Among them, the average time required to scan one data can be estimated by the cloud platform based on the currently available processing resources. Of course, the estimated total query delay can also be determined by other means, and the embodiments of the present application are not limited to this.
[0155] (2) Based on the query conditions, the estimated total query delay, and the single maximum query delay, the content query task is divided into multiple ordered subtasks, where each subtask is used to query a part of the target data content, and the estimated query delay of each subtask does not exceed the single maximum query delay.
[0156] In some embodiments, when splitting a content query task, all data blocks within the query time range are split in chronological order according to the task splitting mechanism to obtain groups of data blocks. The data contained in each group of data blocks is the data that needs to be queried for each subtask, thereby obtaining multiple subtasks.
[0157] The task splitting mechanism can be pre-set by the cloud platform or set by the user. For example, the task splitting mechanism pre-set by the cloud platform is: the estimated query delay of each subtask is half of the single maximum query delay. At this time, assuming that when splitting the content query task, the single total query delay is divided by half of the single maximum query delay to obtain the number n of multiple subtasks, and the data blocks within the query time range are evenly split into n groups in chronological order, and one group of data blocks corresponds to one subtask. Of course, this situation is only an example, and the task splitting mechanism can be set according to needs in actual applications.
[0158] It should be noted that the data that each subtask needs to query is still strongly ordered by time. That is to say, the data that each subtask needs to query is sorted in sequence according to the chronological order of data generation time. For multiple data with the same data generation time, they are sorted and scanned according to the storage order of the multiple data.
[0159] (3) Performing the first query operation by executing at least one subtask.
[0160] Based on the above description, after splitting the content query task, we can obtain multiple ordered subtasks, each corresponding to a portion of the data block. Therefore, if the first query operation is the first query operation, the scan starts from the data block corresponding to the first subtask; if the first query operation is not the first query operation, the scan starts from the pause point of the previous query operation.
[0161] During the scanning process of the data content in the data block, if the scanned data content contains the query term and the data generation time of the data content is within the query time range, the data content is determined to meet the query conditions. In this manner, the data block corresponding to the at least one subtask is queried to obtain the first data content obtained by the first query operation.
[0162] (4) When any one of the multiple target conditions is satisfied first, the first query operation is terminated, and the first data content obtained by the first query operation is displayed in the data query interface, wherein the multiple target conditions include: the query delay of the first query operation reaches the single maximum query delay, the data volume of the first data content reaches the single maximum query data volume, and the target subtask is completed, wherein the target subtask is the last subtask among the multiple subtasks, or, when the target subtask is completed, the sum of the query delay of the first query operation and the estimated query delay of the next subtask of the target subtask is greater than the single maximum query delay.
[0163] That is, when the jth subtask among the multiple subtasks is executed through the first query operation, the first query operation is terminated when any of the following three situations occur:
[0164] In the first case, if the j-th subtask is not completed, but the query delay of the first query operation has reached the single maximum query delay, or if the j-th subtask is not completed, the query delay of the first query operation has not reached the single maximum query delay, but the amount of data retrieved by the first query operation has reached the single maximum query data amount, then the first query operation is terminated.
[0165] In the second case, if the jth subtask has been completed, the query delay of the first query operation has not reached the single maximum query delay, the amount of data retrieved by the first query operation has not reached the single maximum query data amount, but the jth subtask is the last subtask, then the first query operation is terminated.
[0166] In this case, all target data contents have been checked.
[0167] In the third case, if the jth subtask has been completed, but the query delay of the first query operation has not reached the single maximum query delay, the amount of data retrieved by the first query operation has not reached the single maximum query data amount, and the jth subtask is not the last subtask, then determine whether the sum of the query delay of the first query operation and the estimated query delay of the next subtask is greater than the single maximum query delay when the jth subtask is completed. If the sum of the query delay of the first query operation and the estimated query delay of the j+1th subtask is greater than the single maximum query delay, then terminate the first query operation; otherwise, execute the j+1th subtask. During the execution of the j+1th subtask, the above logic is also used for judgment, and the first query operation is terminated when certain conditions are met.
[0168] The estimated query latency for the next subtask can be determined by dividing the number of data blocks corresponding to the next subtask by the number of data blocks corresponding to the completed j-th subtask, and then multiplying the result by the query latency for completing the j-th subtask to obtain the estimated query latency for the next subtask. Of course, the estimated query latency for the next subtask can also be determined by other methods, which are not limited in the present embodiment.
[0169] In some embodiments, after a first query operation is completed, the pause position of the content query task is recorded; a target query option is displayed in the data query interface, the target query option being used to instruct to continue querying data content; and after receiving a user trigger operation for the target query option, a second query operation is performed based on the pause position. The second query operation is the next query operation after the first query operation.
[0170] It should be noted that after the second query operation is completed, the data content obtained through the second query operation is also displayed in the data query interface, but the data content does not overwrite the first data content obtained by the first query operation.
[0171] That is to say, the user can determine whether to continue with the next query operation based on the query results of a certain query operation displayed on the data query interface. If the next query operation is required, the user can initiate the next query operation by triggering the target query option. If the next query operation is not required, the user can not trigger the target query option, thereby ending the content query task in advance, saving time and expenses.
[0172] In some embodiments, after the first query operation is completed, the query progress of the content query task is determined and displayed in the data query interface. This allows the user to obtain timely information on the completion status of the current content query task, allowing the user to terminate the task early as needed, saving time and reducing costs.
[0173] The query progress of a content query task can be determined by dividing the amount of data in the target database within the scanned time range by the total amount of data in the target database within the query time range, and using the result as the query progress of the content query task. Of course, the query progress of a content query task can also be determined by other methods, which are not limited in this embodiment of the present application.
[0174] For example, please refer to Figure 6. Suppose the query term of a content query task is @version: do you wanna build a snowman, and the query time range is 1 year (hourly time) (i.e., from 0:00 on January 1 of this year to the start time of the task). Figure 6 is the data query interface after a query operation is completed. It can be seen from Figure 6 that the query progress of the current task is 7.14%. Since no data content that meets the query conditions has been queried at present, no data content appears in the data query interface, and "No table data" and "Query progress 7.14%" appear. After the user obtains the current query result, he wants to query more data content, so he clicks the target query option, that is, the underlined "Query more" in Figure 6. The cloud platform responds to the operation and continues to query more data content. Afterwards, the data query interface after the next query operation is completed is shown in Figure 7. It can be seen from Figure 7 that the query progress of the current content query task is 56.04%, and a data content that meets the query conditions is queried. The row number of the data content is 1, and the data generation time is 14:20:40 on September 19, 2023.
[0175] It should be noted that before the first query operation is performed, the content query task can be split into multiple subtasks. After each subsequent query operation is completed, the starting query position of the next query operation is the pause position of the previous query operation. Of course, before the first query operation, the content query task can be split into multiple subtasks. After each subsequent query operation is completed, the remaining part of the content query task can be split again to obtain multiple subtasks, and then execution can be started from the first subtask. The subsequent steps are still the same as the above steps (3) and (4).
[0176] After querying and displaying the data content through the above steps, the user may trigger a context query task regarding a certain data content. In this case, the method further includes the following steps.
[0177] Step 504: Obtain a context query task about the benchmark data content. The context query task is used to query the target context data content that meets the context query conditions. The context query conditions include the previous time range, the following time range and the query statement. The benchmark data content is any data content displayed on the data query interface.
[0178] After finding data content that meets the query criteria through the above content query task, the user can select a data content from the data content displayed on the data query interface as a reference data content, thereby triggering a context query task related to the reference data content. Of course, the reference data content can also be any data content displayed on the data query interface, not necessarily the data content found by the above content query task.
[0179] In some embodiments, a user can trigger a contextual query task about the baseline data content in various ways. For example, referring to FIG7 , a user clicks the view details area before a piece of data content, i.e., area E in FIG7 , to select the data content in row 1 in FIG7 as the baseline data content, thereby triggering a contextual query task about the baseline data content.
[0180] The definition of a query statement is similar to the query terms described above and will not be further elaborated here. Furthermore, a query statement is typically a statement within the data content being queried. If the baseline data content is a data content queried in the content query task described above, the query statement for the context query task is typically different from the query terms for the content query task.
[0181] The context query task includes the previous query task and the following query task. The previous query task is used to query the data content whose data generation time is before the data generation time of the benchmark data content and meets the context query conditions. The following query task is used to query the data content whose data generation time is after the data generation time of the benchmark data content and meets the context query conditions.
[0182] Regardless of whether the benchmark data content is the data content retrieved by the aforementioned content query task or other data content, when determining the above-context time range based on the data generation time of the benchmark data content, a first time range may be used as the above-context time range. This first time range is the time range from the earliest data generation time in the target database to the data generation time of the benchmark data content. If the benchmark data content is the data content retrieved by the aforementioned content query task, a second time range may also be used as the above-context time range. This second time range is the time range from the earliest time of the query time range in the aforementioned query condition to the data generation time of the benchmark data content.
[0183] Similarly, regardless of whether the benchmark data content is the data content retrieved by the aforementioned content query task or other data content, when determining the subsequent time range based on the data generation time of the benchmark data content, a third time range may be used as the subsequent time range. This third time range is the time range from the data generation time of the benchmark data content to the latest data generation time in the target database. If the benchmark data content is the data content retrieved by the aforementioned content query task, a fourth time range may also be used as the subsequent time range. This fourth time range is the time range from the data generation time of the benchmark data content to the latest time of the query time range in the aforementioned query condition.
[0184] Step 505: Based on the context query conditions and the single query result return conditions, perform the first previous query operation and the first next query operation, and display the first previous data content obtained through the first previous query operation and the first next data content obtained through the first next query operation in the data query interface. The first previous query operation and the first next query operation are both used to query a part of the target context data content. The query delays of the first previous query operation and the first next query operation are not greater than the single maximum query delay, and the data volume of the first previous data content and the first next data content is not greater than the single maximum query data volume.
[0185] The implementation process for both the previous and next query tasks is similar to the content query task described above. Similarly, the process for performing the first previous query and the first next query are similar to the first query described above and will not be described in detail here. Since the implementation process for the previous query task is similar to the next query task, we will briefly explain the previous query task as an example.
[0186] In some embodiments, a first context query operation is performed based on the query statement, the context time range, and the single query result return condition, and the first context data content obtained through the first context query operation is displayed in the data query interface.
[0187] As an example, a first above-text query operation is performed based on the query statement, the above-text time range and the single query result return conditions; when any one of the multiple target conditions is met first, the first above-text query operation is terminated, and the first above-text data content obtained by the first above-text query operation is displayed in the data query interface. The multiple target conditions include: the query delay of the first above-text query operation reaches the single maximum query delay, the amount of data retrieved by the first above-text query operation reaches the single maximum query data amount, and the first above-text query operation has completed the query of the target above-text content in the target context data content.
[0188] As another example, the above query task can be split into multiple above subtasks. In this case, based on the query statement, the above time range, and the single query result return condition, a first above query operation is performed, and the first above data content obtained by the first above query operation is displayed in the data query interface, including the following steps (1)-(4):
[0189] (1) Determine the estimated total query latency of the query task.
[0190] The method for determining the estimated total query delay is similar to the method for determining the estimated total query delay in step 503 above, and will not be repeated here.
[0191] (2) Based on the query statement, the context time range, the estimated total context query delay, and the single maximum query delay, the context query task is divided into multiple ordered context subtasks, where each context subtask is used to query a part of the target context data content, and the estimated query delay of each context subtask does not exceed the single maximum query delay.
[0192] The target context data content refers to data content in the target context data content, the data generation time of which is earlier than the data generation time of the reference data content.
[0193] Similarly, the method of cutting multiple ordered subtasks can refer to the method of cutting multiple ordered subtasks in the above step 503.
[0194] (3) Performing the first context query operation by executing at least one context subtask.
[0195] (4) When any one of the multiple target conditions is satisfied first, the first context query operation is terminated, and the first context data content obtained by the first context query operation is displayed in the data query interface, wherein the multiple target conditions include: the query delay of the first context query operation reaches the single maximum query delay, the data volume of the first context data content reaches the single maximum query data volume, and the target context subtask is completed, wherein the target context subtask is the last subtask among the multiple context subtasks, or, when the target context subtask is completed, the sum of the query delay of the first context query operation and the estimated context query delay of the next context subtask of the target context subtask is greater than the single maximum query delay.
[0196] For the description of the case where the first preceding text query operation is ended, please refer to the description of the case where the first query operation is ended in the above step 503, which will not be repeated here.
[0197] Similarly, before the first above query operation is performed, the above query task can be split into multiple above subtasks. After each above query operation is completed, the starting query position of the next above query operation is the pause position of the previous above query operation. Of course, before the first above query operation is performed, the above query task can be split into multiple above subtasks. After each above query operation is completed, the remaining part of the above query task can be split again to obtain multiple above subtasks, and then execution can be started from the first above subtask. The subsequent steps are still the same as steps (3) and (4) above.
[0198] In some embodiments, after a first context query operation is completed, the pause position of the context query task is recorded; a target context query option is displayed in the data query interface, the target context query option being used to instruct to continue querying the context data content; and after receiving a user trigger operation on the target context query option, a second context query operation is performed based on the pause position. The second context query operation is the next context query operation after the first context query operation.
[0199] That is to say, the user can determine whether to continue the next previous query operation based on the current query results of the previous query task displayed on the data query interface. If the next previous query operation is required, the user can initiate the next previous query operation by triggering the target previous query option. If the next previous query operation is not required, the user can not trigger the target previous query option, thereby ending the previous query task early and saving time and expenses.
[0200] Next, we will use the query task above as an example to illustrate:
[0201] In some embodiments, when a user-triggered first context query request is received, a first context query operation is performed. The user may trigger the first context query request by pressing the Enter key, clicking a query button when the data query interface includes a query button, or other methods, which are not limited in the embodiments of the present application.
[0202] In other embodiments, when it is detected that the query condition has not been updated within the fifth time period, the first above-mentioned query operation is performed. The fifth time period can be pre-set by the cloud platform or set by the user. The size of the fifth time period can be set according to actual application requirements. The embodiment of the present application does not limit the setting method and size of the fifth time period.
[0203] In some embodiments, after the first context query operation is completed, the query progress of the context query task is determined and displayed in the data query interface. This allows the user to obtain the completion status of the current context query task in a timely manner, allowing the user to terminate the task early as needed, saving time and reducing costs.
[0204] The query progress of the query task can be determined by dividing the amount of data scanned in the target database within the time range by the total amount of data in the target database within the time range. The result is the query progress of the query task. Of course, the query progress of the query task can also be determined by other methods, which are not limited in the present embodiment.
[0205] In some embodiments, after the first previous query operation is completed, the data generation time of the last data currently scanned in the previous query task is recorded, and the data generation time is the time position currently scanned by the above query task; the time position currently scanned by the above query task is displayed in the data query interface.
[0206] Through the current scanned time position, users can know the time range that has been scanned. Before starting a task, users can often roughly understand the time range of the data to be obtained. If users are sure that there is no content to be queried within the unscanned time range, they can terminate the query in advance to avoid the overhead caused by meaningless scanning.
[0207] For example, please refer to Figure 7, which shows the current data query interface for a content query task. Assume that the data content in row 1 shown in Figure 7 is selected as the baseline data content, and a context query is performed on the baseline data content. Clicking area E in Figure 7 triggers the context query task for the baseline data content. Next, please refer to Figure 8, where area 1 in Figure 8 indicates that the user selects a context, i.e., enters the query statement for the triggered context query task. The query statement for the context query task is "Source_host:xxx", the previous time range is from the earliest data generation time in the target database to the data generation time of the baseline data content, and the next time range is from the data generation time of the baseline data content to the latest data generation time in the target database. After the user enters the query statement and selects the previous and next time ranges, they click the "OK" button to trigger the context query request. The cloud platform will execute both the previous and next query tasks simultaneously. The "Earlier" and "Update" buttons in Figure 8 represent the target previous and next query options, respectively. Clicking the "Earlier" and "Update" buttons allows the cloud platform to continue querying more previous and next data content based on the paused position, respectively.
[0208] Figure 8 shows the data query interface of the triggered context query task after a query operation has completed. As shown in Figure 8 , the query progress of the current context query task is 1.65%, and the current scanned time position is 6:29:57:510 on September 19, 2023. The query progress of the current context query task is 2.86%, and the current scanned time position is 8:44:22:070 on September 20, 2023. The data content retrieved by the context query task is displayed above the baseline data content, as shown in Figure 8 , with row number -1 and data generated at 14:17:18:550 on September 19, 2023. The data content retrieved by the context query task is displayed below the baseline data content, as shown in Figure 8 , with row number 1 and data generated at 14:29:49:580 on September 19, 2023.
[0209] In some embodiments, when the query task is a content query task and / or a context query task, the single maximum query data volume and / or the single maximum query delay are input by the user in the data query interface. If the user wants to see the query results of each time as quickly as possible, a smaller single maximum query data volume and a shorter single maximum query delay can be set. The size of the single maximum query data volume and the single maximum query delay can be adjusted according to actual needs, and the embodiments of the present application do not limit the setting method and size of the two. For example, as shown in area F in Figure 6, area G in Figure 7, and area H in Figure 8, the data query interface also provides a single maximum query delay input area, and the user can enter the single maximum query delay through this area.
[0210] In some embodiments, when the query task is a content query task and / or a context query task, the multiple target conditions also include the amount of data scanned by the query operation reaching a maximum single scan number, which refers to the maximum amount of data traversed during each query operation. This maximum single scan number can be pre-set by the cloud platform or set by the user.
[0211] Since related data query products are often charged based on the amount of data scanned by the cloud platform, allowing users to set the maximum number of single scans themselves can enable users to control query costs from the bottom level and thus save expenses.
[0212] In content query tasks and context query tasks, by setting a single maximum query delay, the time it takes for each query operation to return results can be controlled within a certain period of time. By setting the conditions for returning single query results, the cloud platform can promptly return partially ordered results in various situations, improving the user experience. The target query options, target previous query options, and target next query options provided by the data query interface allow users to initiate queries for more data after obtaining the current results, ensuring that users can obtain the data content they need. Since related data query products are often billed based on the amount of data scanned by the cloud platform, setting a single maximum scan number allows users to control query costs from the bottom layer, thereby saving expenses. In addition, displaying the current task query progress and the current scanned time position through the data query interface helps users understand the current query status of the context query task and helps users determine whether to end the context query task based on their needs, thereby saving time and reducing expenses.
[0213] FIG9 is a schematic diagram of the structure of a data query device provided by an embodiment of the present application. Referring to FIG9 , the device includes: an interface providing module 901 , a task acquiring module 902 , and a first result displaying module 903 .
[0214] The interface providing module 901 is used to provide a data query interface, which is used to obtain the query task input by the user;
[0215] A task acquisition module 902 is used to acquire an aggregate query task from a data query interface. The aggregate query task is used to query the total amount of data that meets the query conditions. The query conditions include query terms and query time range.
[0216] The first result display module 903 is used to display the first estimated query result of the aggregate query task in the data query interface, wherein the first estimated query result includes a first total data volume that meets the query conditions obtained by approximate estimation based on the first part of the precise query results, and the first part of the precise query results includes the amount of data that meets the query conditions queried from M data blocks, and the M data blocks are sampled from N data blocks that match the aggregate query task, M and N are integers, and M is less than N.
[0217] Optionally, the device further comprises:
[0218] a second result display module, configured to display a second estimated query result of the aggregate query task in the data query interface, wherein the second estimated query result includes a second total amount of data that meets the query condition obtained by approximating the second portion of the precise query result, the second portion of the precise query result including the amount of data that meets the query condition queried from P data blocks, the P data blocks being sampled from N data blocks that match the aggregate query task, and the M data blocks including the P data blocks;
[0219] In this case, the first result display module is specifically used to:
[0220] In the data query interface, the second estimated query result is replaced with the first estimated query result.
[0221] Optionally, the device further comprises:
[0222] A first determining module, configured to determine a newly added sampled data block;
[0223] A query module, configured to query the newly sampled data blocks based on the query conditions and determine a new query result, the new query result including the amount of data found in the newly sampled data blocks that meets the query conditions;
[0224] The second determining module is configured to determine the first part of the precise query results according to the newly added query results and the second part of the precise query results.
[0225] Optionally, the device further comprises:
[0226] The complete result display module is used to display the complete and accurate query results of the aggregate query task in the data query interface. The complete and accurate query results include the total amount of data that meets the query conditions queried from the above N data blocks.
[0227] Optionally, the first estimated query result further includes a first data statistical graph, where the first data statistical graph is used to indicate the distribution of the first total data volume in different time intervals.
[0228] Optionally, the device further comprises:
[0229] The progress display module is used to display the query progress of the aggregate query task in the data query interface.
[0230] Optionally, the device further comprises:
[0231] The third determining module is configured to determine N data blocks matching the aggregate query task from the target database based on the query condition and the index directory of the target database.
[0232] Each of the above modules can be implemented by software or hardware. For example, the following describes the implementation of the interface providing module using the interface providing module as an example. Similarly, the implementation of the other modules mentioned above can refer to the implementation of the interface providing module.
[0233] As an example of a software functional unit, the interface providing module may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the interface providing module may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Generally, a region may include multiple AZs.
[0234] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0235] As an example of a hardware functional unit, the interface providing module may include at least one computing device, such as a server. Alternatively, the interface providing module may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0236] The multiple computing devices included in the interface provision module can be distributed in the same region or in different regions. The multiple computing devices included in the interface provision module can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the interface module can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0237] In an embodiment of the present application, for aggregated query tasks, a sample-based estimation method is used to approximate the total amount of data that meets the query criteria in each round based on partially accurate query results. Through multiple iterations, complete and accurate query results can be obtained, allowing the data query interface to progressively display increasingly accurate results to the user. Furthermore, since each round only requires traversing a portion of the data blocks, the response time for each query round is very short, allowing users to obtain usable data in a relatively short period of time. Furthermore, the data query interface displays the current task query progress, making it easier for users to understand the completion status of the current query task.
[0238] It should be noted that the data query device provided in the above embodiment only uses the division of the above functional modules as an example to illustrate when performing data query. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the data query device provided in the above embodiment and the data query method embodiment corresponding to Figure 1 are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0239] FIG10 is a schematic diagram of the structure of another data query device provided in an embodiment of the present application. Referring to FIG10 , the device includes: an interface providing module 1001 , a first task acquiring module 1002 , and a first query module 1003 .
[0240] The interface providing module 1001 is used to provide a data query interface, which is used to obtain the query task input by the user;
[0241] A first task acquisition module 1002 is used to acquire a content query task from a data query interface, wherein the content query task is used to query target data content that meets a query condition, wherein the query condition includes a query term and a query time range;
[0242] The first query module 1003 is used to perform a first query operation based on the query conditions and the single query result return conditions, and display the first data content obtained through the first query operation in the data query interface, wherein the single query result return conditions include the single maximum query data volume and the single maximum query delay. The first query operation is used to query a part of the target data content. The query delay of the first query operation is not greater than the single maximum query delay and the data volume of the first data content is not greater than the single maximum query data volume.
[0243] Optionally, the first query module 1003 includes:
[0244] A query submodule, configured to perform a first query operation based on a query condition and a single query result return condition;
[0245] The first content display submodule is used to end the first query operation when any one of multiple target conditions is met first, and display the first data content obtained through the first query operation in the data query interface. The multiple target conditions include: the query delay of the first query operation reaches the single maximum query delay, the amount of data retrieved by the first query operation reaches the single maximum query data amount, and the first query operation has completed the query of the target data content.
[0246] Optionally, the first query module 1003 includes:
[0247] The latency determination submodule is used to determine the estimated total query latency of the content query task;
[0248] A splitting submodule is used to split the content query task into multiple ordered subtasks based on the query conditions, the estimated total query latency, and the single maximum query latency, wherein each subtask is used to query a portion of the target data content, and the estimated query latency of each subtask does not exceed the single maximum query latency;
[0249] an execution submodule, configured to perform a first query operation by executing at least one subtask;
[0250] The second content display sub-module is used to end the first query operation when any one of multiple target conditions is met first, and display the first data content obtained through the first query operation in the data query interface. The multiple target conditions include: the query delay of the first query operation reaches the single maximum query delay, the data volume of the first data content reaches the single maximum query data volume, and the target subtask is completed, wherein the target subtask is the last subtask among the multiple subtasks, or, when the target subtask is completed, the sum of the query delay of the first query operation and the estimated query delay of the next subtask of the target subtask is greater than the single maximum query delay.
[0251] Optionally, the device further comprises:
[0252] A recording module, used to record the pause position of the content query task;
[0253] A display module is used to display a target query option in the data query interface, where the target query option is used to indicate to continue querying the data content;
[0254] The second query module is configured to perform a second query operation based on the pause position after receiving a trigger operation on the target query option by the user.
[0255] Optionally, the device further comprises:
[0256] A progress determination module, used to determine the query progress of the content query task;
[0257] The progress display module is used to display the query progress of the content query task in the data query interface.
[0258] Optionally, the device further comprises:
[0259] A second task acquisition module is configured to acquire a context query task triggered by a user regarding a reference data content. The context query task is configured to query target context data content that meets context query conditions. The context query conditions include a preceding time range, a following time range, and a query statement. The reference data content is any data content displayed on the data query interface.
[0260] A context query module is used to perform a first previous query operation and a first next query operation based on context query conditions and single query result return conditions, and display the first previous data content obtained through the first previous query operation and the first next data content obtained through the first next query operation in a data query interface. The first previous query operation and the first next query operation are both used to query a part of the target context data content. The query delay of the first previous query operation and the first next query operation is not greater than the single maximum query delay, and the data volume of the first previous data content and the first next data content is not greater than the single maximum query data volume.
[0261] Optionally, the above-mentioned maximum single query data volume and / or maximum single query delay are input by the user in the data query interface.
[0262] Each of the above modules can be implemented by software or hardware. For example, the following describes the implementation of the interface providing module using the interface providing module as an example. Similarly, the implementation of the other modules mentioned above can refer to the implementation of the interface providing module.
[0263] As an example of a software functional unit, the interface providing module may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the interface providing module may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.
[0264] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0265] As an example of a hardware functional unit, the interface providing module may include at least one computing device, such as a server. Alternatively, the interface providing module may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0266] The multiple computing devices included in the interface provision module can be distributed in the same region or in different regions. The multiple computing devices included in the interface provision module can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the interface module can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0267] In the embodiment of the present application, for content query tasks and context query tasks, by setting a single maximum query delay, the time for each query operation to return results can be controlled within a certain period of time. By setting the conditions for returning single query results, the cloud platform can return partial ordered results in a timely manner under various circumstances, thereby improving the user experience. The target query option provided by the data query interface allows users to initiate queries for more data after obtaining the current results, ensuring that users can obtain the required data content. In addition, displaying the current task query progress through the data query interface helps users understand the current query status of the content query task, and helps users determine whether to end the content query task according to their needs, so as to save time and reduce expenses.
[0268] It should be noted that the data query device provided in the above embodiment only uses the division of the above functional modules as an example to illustrate when performing data query. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the data query device provided in the above embodiment and the data query method embodiment corresponding to Figure 5 are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0269] The present application also provides a computing device 1100. As shown in FIG11 , computing device 1100 includes a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. Processor 1104, memory 1106, and communication interface 1108 communicate with each other via bus 1102. Computing device 1100 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1100.
[0270] Bus 1102 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG11 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 1102 may include a path for transmitting information between various components of computing device 1100 (e.g., memory 1106, processor 1104, and communication interface 1108).
[0271] The processor 1104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0272] The memory 1106 may include volatile memory, such as random access memory (RAM). The processor 1104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0273] The memory 1106 stores executable program code, and the processor 1104 executes the executable program code to respectively implement the functions of each module in the at least one data query device, thereby implementing the data query method corresponding to Figure 1 and / or Figure 5. In other words, the memory 1106 stores instructions for executing the data query method corresponding to Figure 1 and / or Figure 5.
[0274] The communication interface 1108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1100 and other devices or a communication network.
[0275] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0276] As shown in Figure 12, the computing device cluster includes at least one computing device 1100. The memory 1106 in one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the data query method corresponding to Figure 1 and / or Figure 5.
[0277] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the computing device cluster may also store partial instructions for executing the data query method corresponding to Figure 1 and / or Figure 5. In other words, the combination of one or more computing devices 1100 can jointly execute instructions for executing the data query method corresponding to Figure 1 and / or Figure 5.
[0278] It should be noted that the instructions stored in the memory 1106 in different computing devices 1100 in the computing device cluster can implement the functions of one or more modules in the modules of the at least one data query device mentioned above.
[0279] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network or a local area network, etc. FIG13 shows a possible implementation. As shown in FIG13 , two computing devices 1100A and 1100B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 1106 in the computing device 1100A stores instructions for executing the functions of a portion of the modules in the aforementioned at least one data query device. At the same time, the memory 1106 in the computing device 1100B stores instructions for executing the functions of another portion of the modules in the aforementioned at least one data query device.
[0280] It should be understood that the functionality of the computing device 1100A shown in FIG13 may also be implemented by multiple computing devices 1100. Similarly, the functionality of the computing device 1100B may also be implemented by multiple computing devices 1100.
[0281] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the data query method corresponding to Figure 1 and / or Figure 5.
[0282] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the data query method corresponding to Figure 1 and / or Figure 5.
[0283] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), or a semiconductor medium (e.g., a solid state disk (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiments of the present application may be a non-volatile storage medium, in other words, a non-transient storage medium.
[0284] It should be understood that the "plurality" mentioned herein refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in order to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.
[0285] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.
[0286] The above description is an embodiment provided for this application and is not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.
Claims
1. A data query method, characterized in that: The method comprises: Providing a data query interface, wherein the data query interface is used to obtain a query task input by a user; Acquire an aggregate query task from the data query interface, wherein the aggregate query task is used to query the total amount of data that meets the query conditions, wherein the query conditions include a query term and a query time range; A first estimated query result of the aggregate query task is displayed in the data query interface, wherein the first estimated query result includes a first total data volume that meets the query conditions obtained by approximate estimation based on the first part of the precise query results, and the first part of the precise query results includes the amount of data that meets the query conditions queried from M data blocks, and the M data blocks are sampled from N data blocks that match the aggregate query task, M and N are integers, and M is less than N.
2. The method according to claim 1, characterized in that Before displaying the first estimated query result of the aggregate query task in the data query interface, the method further includes: Displaying a second estimated query result of the aggregate query task in the data query interface, wherein the second estimated query result includes a second total amount of data that meets the query condition and is obtained by approximate estimation based on the second part of the precise query result, and the second part of the precise query result includes the amount of data that meets the query condition and is queried from P data blocks, wherein the P data blocks are sampled from N data blocks that match the aggregate query task, and the M data blocks include the P data blocks; The displaying the first estimated query result of the aggregate query task in the data query interface includes: The second estimated query result is replaced with the first estimated query result in the data query interface.
3. The method according to claim 2, characterized in that Before replacing the second estimated query result with the first estimated query result in the data query interface, the method further includes: Determine the data block for new sampling; Based on the query condition, query the newly added sampled data block to determine a newly added query result, wherein the newly added query result includes the amount of data found in the newly added sampled data block that meets the query condition; The first part of accurate query results is determined according to the newly added query results and the second part of accurate query results.
4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: The complete and accurate query result of the aggregate query task is displayed in the data query interface, and the complete and accurate query result includes the total amount of data that meets the query conditions queried from the N data blocks.
5. The method according to any one of claims 1 to 4, characterized in that The first estimated query result also includes a first data statistical graph, and the first data statistical graph is used to indicate the distribution of the first total data volume in different time intervals.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: The query progress of the aggregate query task is displayed in the data query interface.
7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: Based on the query condition and the index directory of the target database, the N data blocks matching the aggregate query task are determined from the target database.
8. A data query method, characterized in that: The method comprises: Providing a data query interface, wherein the data query interface is used to obtain a query task input by a user; Acquire a content query task from the data query interface, wherein the content query task is used to query target data content that meets a query condition, wherein the query condition includes a query term and a query time range; Based on the query condition and the single query result return condition, a first query operation is performed, and the first data content obtained by the first query operation is displayed in the data query interface, wherein the single query result return condition includes a single maximum query data volume and a single maximum query result. The first query operation is used to query a part of the target data content. The query delay of the first query operation is not greater than the single maximum query delay and the data volume of the first data content is not greater than the single maximum query data volume.
9. The method according to claim 8, characterized in that The performing a first query operation based on the query condition and the single query result return condition, and displaying the first data content obtained by the first query operation in the data query interface, includes: Based on the query condition and the single query result return condition, perform the first query operation; When any one of multiple target conditions is met first, the first query operation is terminated, and the first data content obtained through the first query operation is displayed in the data query interface. The multiple target conditions include: the query delay of the first query operation reaches the single maximum query delay, the amount of data retrieved by the first query operation reaches the single maximum query data amount, and the first query operation has completed the query of the target data content.
10. The method according to claim 8, characterized in that The performing a first query operation based on the query condition and the single query result return condition, and displaying the first data content obtained by the first query operation in the data query interface, includes: Determining an estimated total query latency for the content query task; Based on the query condition, the estimated total query delay and the single maximum query delay, the content query task is divided into a plurality of ordered subtasks, wherein each subtask is used to query a portion of the target data content, and the estimated query delay of each subtask does not exceed the single maximum query delay; Performing the first query operation by executing at least one subtask; When any one of multiple target conditions is met first, the first query operation is terminated, and the first data content obtained through the first query operation is displayed in the data query interface, and the multiple target conditions include: the query delay of the first query operation reaches the single maximum query delay, the data volume of the first data content reaches the single maximum query data volume, and the target subtask is completed, wherein the target subtask is the last subtask among the multiple subtasks, or, when the target subtask is completed, the sum of the query delay of the first query operation and the estimated query delay of the next subtask of the target subtask is greater than the single maximum query delay.
11. The method according to any one of claims 8 to 10, characterized in that After the first query operation is completed, the method further includes: Recording the pause position of the content query task; Displaying a target query option in the data query interface, wherein the target query option is used to indicate to continue querying data content; After receiving a trigger operation of the user on the target query option, a second query operation is performed based on the pause position.
12. The method according to any one of claims 8 to 11, characterized in that After the first query operation is completed, the method further includes: Determining the query progress of the content query task; The query progress of the content query task is displayed in the data query interface.
13. The method according to any one of claims 8 to 12, characterized in that The method further comprises: Acquire a context query task about the benchmark data content triggered by the user, wherein the context query task is used to query the target context data content that meets the context query condition, wherein the context query condition includes a previous time range, a following time range, and a query statement, and the benchmark data content is any data content displayed on the data query interface; Based on the context query condition and the single query result return condition, a first preceding context query operation and a first following context query operation are performed, and the first preceding context data content obtained through the first preceding context query operation and the first following context data content obtained through the first following context query operation are displayed in the data query interface, the first preceding context query operation and the first following context query operation are both used to query a part of the target context data content, the query delays of the first preceding context query operation and the first following context query operation are both no greater than the single maximum query delay, and the data volumes of the first preceding context data content and the first following context data content are both no greater than the single maximum query data volume.
14. The method according to any one of claims 8 to 13, characterized in that The single maximum query data volume and / or the single maximum query delay are input by the user in the data query interface.
15. A data query device, characterized in that: The device comprises: An interface providing module, used to provide a data query interface, wherein the data query interface is used to obtain a query task input by a user; A task acquisition module, used to acquire an aggregate query task from the data query interface, wherein the aggregate query task is used to query the total amount of data that meets the query conditions, wherein the query conditions include query terms and query time range; A first result display module is used to display a first estimated query result of the aggregate query task in the data query interface, wherein the first estimated query result includes a first total data volume that meets the query conditions obtained by approximate estimation based on the first part of the precise query results, and the first part of the precise query results includes the amount of data that meets the query conditions queried from M data blocks, and the M data blocks are sampled from N data blocks that match the aggregate query task, M and N are integers, and M is less than N.
16. The device according to claim 15, characterized in that The device also includes: A second result display module is used to display a second estimated query result of the aggregate query task in the data query interface, wherein the second estimated query result includes a second total amount of data that meets the query condition obtained by approximate estimation based on the second part of the precise query result, and the second part of the precise query result includes the amount of data that meets the query condition queried from P data blocks, wherein the P data blocks are sampled from N data blocks that match the aggregate query task, and the M data blocks include the P data blocks; The first result display module is specifically used for: The second estimated query result is replaced with the first estimated query result in the data query interface.
17. The device according to claim 16, characterized in that The device also includes: A first determination module, used to determine a newly added sampled data block; A query module, configured to query the newly sampled data block based on the query condition and determine a newly added query result, wherein the newly added query result includes the amount of data found in the newly sampled data block that meets the query condition; The second determining module is used to determine the first part of accurate query results according to the newly added query results and the second part of accurate query results.
18. The device according to any one of claims 15 to 17, characterized in that The device also includes: A complete result display module is used to display the complete and accurate query results of the aggregate query task in the data query interface, and the complete and accurate query results include the total amount of data that meets the query conditions queried from the N data blocks.
19. The device according to any one of claims 15 to 18, characterized in that The first estimated query result also includes a first data statistical graph, and the first data statistical graph is used to indicate the distribution of the first total data volume in different time intervals.
20. The device according to any one of claims 15 to 19, characterized in that The device also includes: A progress display module is used to display the query progress of the aggregate query task in the data query interface.
21. The device according to any one of claims 15 to 20, characterized in that The device also includes: The third determination module is used to determine the N data blocks matching the aggregate query task from the target database based on the query condition and the index directory of the target database.
22. A data query device, characterized in that: The device comprises: An interface providing module, used to provide a data query interface, wherein the data query interface is used to obtain a query task input by a user; A first task acquisition module, used to acquire a content query task from the data query interface, wherein the content query task is used to query target data content that meets a query condition, wherein the query condition includes a query term and a query time range; The first query module is used to perform a first query operation based on the query condition and the single query result return condition, and display the first data content obtained by the first query operation in the data query interface, wherein the single query result return condition includes a single query result return condition. The maximum query data volume and the single maximum query delay, the first query operation is used to query a part of the target data content, the query delay of the first query operation is not greater than the single maximum query delay and the data volume of the first data content is not greater than the single maximum query data volume.
23. The device according to claim 22, characterized in that The first query module includes: A query submodule, configured to perform the first query operation based on the query condition and the single query result return condition; The first content display submodule is used to end the first query operation and display the first data content obtained by the first query operation in the data query interface when any one of multiple target conditions is met first. The multiple target conditions include: the query delay of the first query operation reaches the single maximum query delay, the amount of data retrieved by the first query operation reaches the single maximum query data amount, and the first query operation has completed the query of the target data content.
24. The device according to claim 22, characterized in that The first query module includes: A delay determination submodule, used to determine the estimated total query delay of the content query task; A cutting submodule, configured to cut the content query task into a plurality of ordered subtasks based on the query condition, the estimated total query delay and the single maximum query delay, wherein each subtask is used to query a portion of the target data content, and the estimated query delay of each subtask does not exceed the single maximum query delay; An execution submodule, configured to perform the first query operation by executing at least one subtask; The second content display submodule is used to end the first query operation and display the first data content obtained by the first query operation in the data query interface when any one of multiple target conditions is met first, and the multiple target conditions include: the query delay of the first query operation reaches the single maximum query delay, the data volume of the first data content reaches the single maximum query data volume, and the target subtask is completed, wherein the target subtask is the last subtask among the multiple subtasks, or, when the target subtask is completed, the sum of the query delay of the first query operation and the estimated query delay of the next subtask of the target subtask is greater than the single maximum query delay.
25. The device according to any one of claims 22 to 24, characterized in that The device also includes: A recording module, used for recording the pause position of the content query task; A display module, used to display a target query option in the data query interface, wherein the target query option is used to indicate to continue to query the data content; The second query module is used to perform a second query operation based on the pause position after receiving a trigger operation of the user on the target query option.
26. The device according to any one of claims 22 to 25, characterized in that The device also includes: A progress determination module, used to determine the query progress of the content query task; The progress display module is used to display the query progress of the content query task in the data query interface.
27. The device according to any one of claims 22 to 26, characterized in that The device also includes: A second task acquisition module is used to acquire a context query task about the benchmark data content triggered by the user, wherein the context query task is used to query the target context data content that meets the context query condition, wherein the context query condition includes a previous time range, a following time range, and a query statement, and the benchmark data content is any data content displayed on the data query interface; A context query module is used to perform a first preceding query operation and a first following query operation based on the context query condition and the single query result return condition, and display the first preceding data content obtained through the first preceding query operation and the first following data content obtained through the first following query operation in the data query interface, wherein the first preceding query operation and the first following data content are both used to query a part of the target context data content, the query delays of the first preceding query operation and the first following query operation are both no greater than the single maximum query delay, and the data volumes of the first preceding data content and the first following data content are both no greater than the single maximum query data volume.
28. The device according to any one of claims 22 to 27, characterized in that The single maximum query data volume and / or the single maximum query delay are input by the user in the data query interface.
29. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 7.
30. A computer-readable storage medium, characterized in that: The storage medium stores instructions. When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 7.
31. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 7.
32. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 8 to 14.
33. A computer-readable storage medium, characterized in that: The storage medium stores instructions. When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 8 to 14.
34. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 8 to 14.
Citation Information
Patent Citations
Data query method and related device
CN119917530A
Data query job submission management
CN107077490A
Sampling query method and device
CN108491262A
Information search method, search request control method and system
CN110827108A
Data processing method and system thereof
CN112308637A