Query task arrangement method and device
By splitting the multi-cycle query task into single-cycle tasks and executing it in parallel, the problems of inefficient query efficiency and waste of resources in the existing technology are solved, and efficient query processing and optimized user experience is achieved.
Patent Information
- Application Number
- CN202510189083.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-24
AI Technical Summary
When facing a single or multiple multi-cycle query tasks, existing query task orchestration methods will lead to inefficient query and waste of computing resources, and will also affect the user query experience.
By splitting the multi-cycle query task into multiple single-cycle query tasks, and determining the number of allocable threads based on the complexity of the query task and database performance, the single-cycle query task is executed in parallel in the database, and finally the single-cycle query results are aggregated to obtain the multi-cycle query results.
This method improves query efficiency and thread utilization by reducing the data size and operation complexity of single-cycle query tasks, avoids waste of computing resources, and reduces the impact on other query tasks, thereby improving user query experience.
Smart Images

Figure CN120196660A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of query technologies, and in particular, to a query task scheduling method and apparatus. Background Art
[0002] Currently, users can configure analysis services on their own on a behavior analysis system and generate query tasks, and obtain analysis data from a database based on the query tasks.
[0003] However, when a query task spans multiple cycles, the data volume involved in the query task will become larger and the operations involved will become more complex. According to the existing query task scheduling method, on the one hand, for a single multi-cycle query task, since it only occupies one query thread, the large data volume and complex operations will cause a greater computing pressure on this query thread, while other threads are idle, resulting in both low query efficiency and waste of computing resources; on the other hand, for multiple multi-cycle query tasks, since they occupy multiple query threads, the large data volume and complex operations will simultaneously cause a greater computing pressure on multiple query threads, which will also result in low query efficiency, and since all computing resources are occupied during the query period, other query tasks cannot be carried out.
[0004] In summary, the existing query task scheduling method will result in both low query efficiency and waste of computing resources when facing a single multi-cycle query task, and will result in low query efficiency and affect the execution of other query tasks when facing multiple multi-cycle query tasks, thus affecting the user query experience. Summary of the Invention
[0005] Embodiments of this application provide a query task scheduling method and apparatus, which are used to solve the technical problem that the existing query task scheduling method will result in both low query efficiency and waste of computing resources when facing a single multi-cycle query task, and will result in low query efficiency and affect the execution of other query tasks when facing multiple multi-cycle query tasks, thus affecting the user query experience.
[0006] In a first aspect, embodiments of this application provide a query task scheduling method, including: Splitting a multi-cycle query task into multiple single-cycle query tasks; Determining the number of assignable threads for the multiple single-cycle query tasks based on the complexity of the multi-cycle query task and the performance of the database to be queried; Based on the number of assignable threads, parallelly executing each single-cycle query task in the database to be queried to obtain single-cycle query results for each single-cycle query task; Aggregating the single-cycle query results to obtain a query result for the multi-cycle query task.
[0007] In one embodiment, the complexity of the multi-cycle query task is determined based on the following method: Obtain the scores of multiple complexity metrics of the multi-cycle query task; Perform weighted summation on the scores of the multiple complexity metrics to obtain the complexity score of the multi-cycle query task; The multiple complexity metrics include the length type, nesting type, join type, aggregation function type, condition type, and index usage type corresponding to the query statement of the multi-cycle query task.
[0008] In one embodiment, the performance of the database to be queried is determined based on the following method: Obtain the scores of multiple performance metrics of the database to be queried; Perform weighted summation on the scores of the multiple performance metrics to obtain the performance score of the database to be queried; The multiple performance metrics include the response duration type, throughput type, resource utilization type, concurrent connection number type, and lock wait duration type of the database to be queried; The resource utilization type is obtained by synthesizing the CPU usage type, memory usage type, disk I / O usage type, and network bandwidth usage type of the database to be queried.
[0009] In one embodiment, determining the number of assignable threads for the multiple single-cycle query tasks based on the complexity of the multi-cycle query task and the performance of the database to be queried includes: Obtain the complexity level corresponding to the complexity score of the multi-cycle query task; Obtain the query task quantity and query thread ratio corresponding to the complexity level; Based on the query task quantity, the query thread ratio, and the performance score of the database to be queried, determine the number of assignable threads for the multiple single-cycle query tasks.
[0010] In one embodiment, parallelly executing each single-cycle query task in the database to be queried based on the number of assignable threads to obtain the single-cycle query result of each single-cycle query task includes: Create a thread pool; Submit each single-cycle query task to the thread pool; Based on the idle threads of the number of assignable threads in the thread pool, parallelly execute each single-cycle query task in the database to be queried to obtain the single-cycle query result of each single-cycle query task.
[0011] In one embodiment, aggregating the single-cycle query results to obtain the query result of the multi-cycle query task includes: Grouping the single-cycle query results according to the query fields to obtain multiple groups of query results; Calculating the query results within each group according to the query target to obtain multiple groups of calculation results; Adjusting the format of the calculation results of each group to obtain the query result of the multi-cycle query task.
[0012] In a second aspect, an embodiment of the present application provides a query task scheduling device, including: A query task splitting module, configured to: split a multi-cycle query task into multiple single-cycle query tasks; A query thread allocation module, configured to: determine the number of allocable threads for the multiple single-cycle query tasks based on the complexity of the multi-cycle query task and the performance of the database to be queried; A query result acquisition module, configured to: based on the number of allocable threads, execute each single-cycle query task in parallel in the database to be queried to obtain the single-cycle query results of each single-cycle query task; A query result aggregation module, configured to: aggregate the single-cycle query results to obtain the query result of the multi-cycle query task.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory storing a computer program, and the processor implements the steps of the query task scheduling method described in the first aspect when executing the program.
[0014] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program, and the computer program implements the steps of the query task scheduling method described in the first aspect when executed by a processor.
[0015] In a fifth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium, including a computer program, and the computer program implements the steps of the query task scheduling method described in the first aspect when executed by a processor.
[0016] The query task scheduling method and device provided by this application split a multi-cycle query task into multiple single-cycle query tasks, determine the allocable number of threads for the multiple single-cycle query tasks based on the complexity of the multi-cycle query task and the performance of the database to be queried, and execute each single-cycle query task in parallel in the database to be queried based on the allocable number of threads to obtain the single-cycle query results of each single-cycle query task, and aggregate the single-cycle query results of each single-cycle query task to obtain the query result of the multi-cycle query task. Whether this application faces a single multi-cycle query task or multiple multi-cycle query tasks, each multi-cycle query task is split into multiple single-cycle query tasks, so that the data volume of the single-cycle query task is greatly reduced compared with the multi-cycle query task, and the operation of the single-cycle query task is also greatly simplified compared with the multi-cycle query task. Then, based on the complexity of each multi-cycle query task and the performance of the database to be queried, threads are allocated to the multiple split single-cycle query tasks for parallel query and query result aggregation, so that each thread executes the small-volume and simple single-cycle query tasks in parallel, thereby greatly improving the query efficiency and thread utilization rate, avoiding waste of computing resources. At the same time, since the improvement of the query efficiency will significantly reduce the occupation time of computing resources, it is possible to reduce the impact on other query tasks and improve the user query experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is one of the flow diagrams of the query task scheduling method provided by the embodiment of this application; Figure 2 is the second flow diagram of the query task scheduling method provided by the embodiment of this application; Figure 3 is the third flow diagram of the query task scheduling method provided by the embodiment of this application; Figure 4 is the fourth flow diagram of the query task scheduling method provided by the embodiment of this application; Figure 5 is the fifth flow diagram of the query task scheduling method provided by the embodiment of this application; Figure 6 is the sixth flow diagram of the query task scheduling method provided by the embodiment of this application; Figure 7 is the structural diagram of the query task scheduling device provided by the embodiment of this application; Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0019] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without making creative efforts shall fall within the protection scope of the present application.
[0020] Figure 1 It is one of the schematic flowcharts of the query task scheduling method provided by an embodiment of the present application. Refer to Figure 1 , an embodiment of the present application provides a query task scheduling method, which may include: 101. Split a multi-cycle query task into multiple single-cycle query tasks; 102. Determine the number of assignable threads for multiple single-cycle query tasks based on the complexity of the multi-cycle query task and the performance of the database to be queried; 103. Based on the number of assignable threads, execute each single-cycle query task in parallel in the database to be queried to obtain the single-cycle query results of each single-cycle query task; 104. Aggregate the single-cycle query results of each single-cycle query task to obtain the query result of the multi-cycle query task.
[0021] In step 101, the cycle range involved can be extracted from the query conditions of the multi-cycle query task A. For example, if the query condition is WHERE data_date BETWEEN '20240101' AND '20241031', the cycle range is determined to be from 20240101 to 20241031; According to business requirements and data characteristics, determine the time span of a single cycle. For example, if business data is recorded on a daily basis and it is more appropriate to perform single-cycle analysis on a daily basis, the time span of a single cycle is one day. If business data is recorded on a monthly basis and it is more appropriate to perform single-cycle analysis on a monthly basis, the time span of a single cycle is one month; Starting from the start point of the cycle range, gradually traverse to the end point of the cycle range according to the time span of a single cycle. For each single cycle, generate a corresponding single-cycle query task B. For each B, construct a new cycle range in the query condition to replace the cycle range in the query condition of A. For example, if the time span of a single cycle is one day and A is split into 31 Bs by day, for B with a single cycle of 20240105, the cycle range in its query condition can be set to WHERE data_date = '20240105'; Furthermore, in the query condition of B, except for the filtering conditions related to multi-cycle, other filtering conditions are the same as those of A, such as other field filtering conditions, associated table join filtering conditions, etc.; if A includes aggregate functions and grouping operations, then adaptive adjustments need to be made in B. For example, if A includes a function for calculating the number of times, it can remain unchanged in B; if A includes a function with duplicate removal logic such as calculating the number of people, then according to the specific business scenario, the function needs to be adjusted in B.
[0022] After the above steps, the multi-cycle query task A can be split into multiple single-cycle query tasks B.
[0023] In step 102, since the complexity of the multi-cycle query task and the performance of the database to be queried have a great impact on the execution of the query task, based on these two factors, thread numbers are allocated for multiple single-cycle query tasks to ensure that each single-cycle query task can be executed smoothly to the greatest extent. Among them, the database to be queried can be ClickHouse, which is mainly applied in the field of online analytical processing, supports a flexible large-scale parallel processing architecture and can achieve linear scalability. It is especially suitable for scenarios with frequent reads and few data updates, such as data warehouses and real-time analysis, and can support users to configure themselves in a visual way to achieve real-time statistical analysis of the behavior scenarios of data producers.
[0024] The query task scheduling method provided in this embodiment splits a multi-cycle query task into multiple single-cycle query tasks, determines the number of allocable threads for the multiple single-cycle query tasks based on the complexity of the multi-cycle query task and the performance of the database to be queried, and based on the number of allocable threads, parallelly executes each single-cycle query task in the database to be queried to obtain the single-cycle query results of each single-cycle query task, and aggregates the single-cycle query results of each single-cycle query task to obtain the query result of the multi-cycle query task. Whether this embodiment faces a single multi-cycle query task or multiple multi-cycle query tasks, each multi-cycle query task is split into multiple single-cycle query tasks, so that the data volume of the single-cycle query task is greatly reduced compared to the multi-cycle query task, and the operation of the single-cycle query task is also greatly simplified compared to the multi-cycle query task. Then, based on the complexity of each multi-cycle query task and the performance of the database to be queried, threads are allocated to the multiple split single-cycle query tasks for parallel query and query result aggregation, so that each thread parallelly executes small-volume and simple single-cycle query tasks, thereby greatly improving the query efficiency and thread utilization rate, avoiding waste of computing resources. At the same time, since the improvement of the query efficiency will significantly reduce the occupation time of computing resources, it is possible to reduce the impact on other query tasks and improve the user query experience.
[0025] Figure 2 It is the second flowchart of the query task scheduling method provided in the embodiment of the present application. Refer to Figure 2 , in one embodiment, the complexity of the multi-cycle query task can be determined based on the following method: 201. Obtain the scores of multiple complexity indicators of the multi-cycle query task; 202. Perform weighted summation on the scores of the multiple complexity indicators to obtain the complexity score of the multi-cycle query task.
[0026] The multiple complexity indicators include the length type, nesting type, connection type, aggregation function type, condition type, and index usage type of the query statement of the multi-cycle query task.
[0027] In step 201: 1. The length type corresponding to the query statement can be divided into short statements, medium-length statements, long statements, and very long statements, where: The short statement can be a statement with a character count less than or equal to 50. This type is usually relatively simple and does not introduce too much complexity, and a score of 0 can be assigned to it; The medium-length statement can be a statement with a character count greater than 50 and less than or equal to 200. This type begins to involve some slightly more complex logic, but is still relatively simple overall, and a score of 0.3 can be assigned to it; A long statement can be a statement with a character count greater than 200 and less than or equal to 500. This type contains more conditions, field operations, etc., and has a higher complexity. It can be assigned 0.6 points. A very long statement can be a statement with a character count greater than 500. This type is usually very complex, contains multiple clauses and complex logical relationships, and can be assigned 1 point.
[0028] 2. The nested types corresponding to query statements can be divided into non-nested, single-layer simple nesting, multi-layer simple nesting, and complex nesting. Among them: Non-nested means there is no nested query. This type is usually the most basic and can be assigned 0 points. Single-layer simple nesting can have only one layer of subquery. This type needs to consider the relationship between the inner and outer layer queries, so the complexity is increased. It can be assigned 0.4 points. Multi-layer simple nesting can include multiple layers of subqueries. This type makes the logical relationship more complex, and the difficulty of parsing and execution also increases accordingly. It can be assigned 0.7 points. Complex nesting can be nested calls in a stored procedure or include complex logics such as recursion in the nesting. This type has a high degree of complexity, and the difficulty of parsing, execution, and optimization further increases. It can be assigned 1 point.
[0029] 3. The join types corresponding to query statements can be divided into non-join, single inner join, multiple inner joins, single non-inner join, and multiple non-inner joins. Among them: Non-join means there is no join operation. This type is relatively simple and mainly operates on a single table. It can be assigned 0 points. Single inner join means joining two tables and returning the matching records of the joined fields in both tables. This type needs to consider the association relationship between the two tables, increasing a certain degree of complexity. It can be assigned 0.3 points. Multiple inner joins mean joining three or more tables and returning the matching records in these tables. This type needs to consider the association relationships between more tables, further increasing the complexity. It can be assigned 0.6 points. Single non-inner join can be at least one outer join or at least one cross join. Among them, outer join means joining two tables and returning the matching and non-matching records of the joined fields in both tables, and cross join means joining two tables and returning all combinations of the joined fields in both tables. This type includes non-matching records, further increasing the complexity. It can be assigned 0.8 points. Multiple non-inner joins can be a combination of at least one outer join and at least one cross join. This type needs to handle multiple complex join relationships and data matching situations, further increasing the complexity. It can be assigned 1 point.
[0030] 4. The aggregation function types corresponding to query statements can be divided into no aggregation function, single simple aggregation function, multiple simple aggregation functions, and complex aggregation functions, where: The no aggregation function means there is no aggregation operation. This type is relatively simple, mainly for data query and filtering, and can be assigned 0 points. The single simple aggregation function can be one of the sum function, count function, and mean function. This type requires summarizing and calculating data, increasing a certain degree of complexity, and can be assigned 0.3 points. The multiple simple aggregation functions can be a combination of multiple functions among the sum function, count function, and mean function. This type involves summarizing data in different dimensions, further increasing the complexity, and can be assigned 0.6 points. The complex aggregation function can be a window function or a nested aggregation function. This type has more advanced calculation logic, significantly increasing the complexity, and can be assigned 1 point.
[0031] 5. The condition types corresponding to query statements can be divided into single simple condition, multiple simple conditions, conditions including function calls, conditions including subqueries, and complex conditions, where: The single simple condition can be a single equal or not equal comparison condition. This type is relatively simple, and the overall complexity is not high, and can be assigned 0 points. The multiple simple conditions can be a combination of at least two simple conditions connected by AND / OR. This type needs to consider the relationship between different conditions, increasing the complexity of logical judgment, and can be assigned 0.4 points. The conditions including function calls need to consider the calculation logic and return value of the function, making the evaluation of the conditions more complex, and can be assigned 0.7 points. The conditions including subqueries will introduce additional query logic, greatly increasing the complexity, and can be assigned 0.9 points. The complex conditions can be a combination of conditions including function calls and conditions including subqueries, which involve various complex logics and query relationships, further increasing the complexity, and can be assigned 1 point.
[0032] 6. The index usage types corresponding to query statements can be divided into effective index usage, partial effective index usage, and ineffective index usage, where: Effective index usage means that all fields involved can be optimized through appropriate indexes. This type indicates that the query statement is optimal in terms of index usage and will not increase complexity due to index problems, and can be assigned 0 points. When only part of the index is effectively used, there are fields where the index is not effectively used. This type requires additional operations to retrieve data from the database, resulting in a decrease in query performance and an increase in complexity. It can be assigned a score of 0.4. When the index is not effectively used, that is, the index is completely not used or the wrong index is used. This type will lead to low query efficiency and requires more resources to execute the query task, with a higher complexity. It can be assigned a score of 0.8.
[0033] In step 202: 1. For the length type corresponding to the query statement, although it is an important influencing factor for the complexity of multi-cycle query tasks, its influence degree is smaller than that of the nested type and the join type. For example, in multi-cycle query tasks, even if the query statement is long, but without complex nesting and joining, the actual complexity of the multi-cycle query task will not be very high. Therefore, the weight of the length type can be set to 0.15.
[0034] 2. For the nested type corresponding to the query statement, since it involves multiple levels of logic and data relationships, it usually has a significant impact on the complexity of multi-cycle query tasks. Therefore, a higher weight needs to be set for it, that is, the weight of the nested type can be set to 0.2.
[0035] 3. For the join type corresponding to the query statement, since it is one of the key factors affecting the complexity of multi-cycle query tasks when it involves multi-table association and complex joins, a higher weight needs to be set for it, that is, the weight of the join type can be set to 0.2.
[0036] 4. For the aggregate function type corresponding to the query statement, although the aggregate function will increase the complexity, its influence degree on the complexity of multi-cycle query tasks is smaller than that of the nested type and the join type. Therefore, the weight of the aggregate function type can be set to 0.15.
[0037] 5. For the condition type corresponding to the query statement, since it is the key part of controlling data filtering and return in the query statement and will significantly affect the parsing and execution difficulty of the statement, a higher weight needs to be set for it, that is, the weight of the condition type can be set to 0.2.
[0038] 6. For the index usage type corresponding to the query statement, since its influence on the complexity of multi-cycle query tasks is relatively small compared to other complexity metrics that affect the logical structure of the query statement, a lower weight needs to be set for it, that is, the weight of the index usage type can be set to 0.1.
[0039] Suppose the length type of the query statement of a multi - cycle query task is medium length, the nesting type is single - layer simple nesting, the join type is single inner join, the aggregate function type is single simple aggregate function, the condition type is multiple simple conditions, and the index usage type is partial effective use of index. Then the complexity score of this multi - cycle query task is: ; The higher the score, the higher the complexity of the multi - cycle query task.
[0040] It should be noted that the weights of the above complexity metrics can be obtained based on the analysis of the query statement execution data and performance test data of historical multi - cycle query tasks.
[0041] When calculating the complexity score of a multi - cycle query task in this embodiment, multiple complexity metrics including the length type, nesting type, join type, aggregate function type, condition type, and index usage type corresponding to its query statement are comprehensively considered, so as to cover various situations of the query statement. At the same time, the weights set for multiple complexity metrics based on historical data can accurately represent the impact of the corresponding complexity metrics on the complexity of the multi - cycle query task. Therefore, the complexity score obtained by weighted summation can comprehensively reflect the true complexity of the multi - cycle query task and improve the accuracy of complexity measurement.
[0042] Figure 3 It is the third flow diagram of the query task scheduling method provided by the embodiment of the present application. Refer to Figure 3 In one embodiment, the performance of the database to be queried can be determined based on the following method: 301. Obtain the scores of multiple performance metrics of the database to be queried; 302. Perform weighted summation on the scores of multiple performance metrics to obtain the performance score of the database to be queried.
[0043] Multiple performance metrics include the response duration type, throughput type, resource utilization type, concurrent connection number type, and lock waiting duration type of the database to be queried; The resource utilization type is comprehensively obtained based on the CPU usage type, memory usage type, disk I / O usage type, and network bandwidth usage type of the database to be queried.
[0044] In step 301: 1. The response duration of the database to be queried is the duration from when the user sends a query request from its client to when the database to be queried returns the query result, which reflects the response speed of the database to be queried to the query request. Its type can be divided into fast, medium, slower, and very slow, where: The response time for quick correspondence can be less than or equal to 1 second. This type can provide a good user experience in most application scenarios, indicating that the performance of the database to be queried is good, and it can be assigned a score of 0.8. The response time for medium correspondence can be greater than 1 second and less than or equal to 3 seconds. This type will make users feel a certain time delay, but it is still within an acceptable range, and it can be assigned a score of 0.5. The response time for slower correspondence can be greater than 3 seconds and less than or equal to 5 seconds. This type will have a greater impact on the user experience, and the performance of the database to be queried needs to be optimized, and it can be assigned a score of 0.3. The response time for very slow correspondence can be greater than 5 seconds. This type will seriously affect the availability of the database to be queried and user satisfaction, and it can be assigned a score of 0.1.
[0045] 2. The throughput of the database to be queried is the number of transactions or data volume that the database to be queried can process per unit time, reflecting the processing capacity of the database to be queried within a certain period of time. Its types can be divided into high throughput, medium throughput, low throughput, and extremely low throughput, where: High throughput can be a throughput greater than or equal to 90% of the maximum designed throughput. This type indicates that the database to be queried can efficiently process a large number of transactions or data per unit time, with excellent performance, and it can be assigned a score of 0.8. Medium throughput can be a throughput greater than 60% of the maximum designed throughput and less than 90% of the maximum designed throughput. This type indicates that the database to be queried can process transactions or data normally, but there is still room for improvement in processing capacity, and it can be assigned a score of 0.5. Low throughput can be a throughput greater than 30% of the maximum designed throughput and less than or equal to 60% of the maximum designed throughput. This type indicates that the processing capacity of the database to be queried has decreased, which will affect the performance of the database to be queried, and it can be assigned a score of 0.3. Extremely low throughput can be a throughput less than or equal to 30% of the maximum designed throughput. This type indicates that the processing capacity of the database to be queried is seriously insufficient, with a performance bottleneck, and it can be assigned a score of 0.1.
[0046] 3. The resource utilization rate of the database to be queried includes CPU usage rate, memory usage rate, disk I / O usage rate, and network bandwidth usage rate, reflecting the consumption of server resources by the database to be queried. Its types are jointly determined by the CPU usage rate type, memory usage rate type, disk I / O usage rate type, and network bandwidth usage rate type, where: 3.1. The CPU usage rate type can be divided into low usage rate, medium usage rate, high usage rate, and excessive usage rate, where: The CPU usage rate corresponding to low usage can be less than or equal to 40%. This type indicates that there is still a large margin of CPU resources, the database to be queried has good performance, and it can be assigned 0.8 points; The CPU usage rate corresponding to medium usage can be greater than 40% and less than or equal to 70%. This type indicates that the CPU has a certain load, but it is still within a reasonable range, and it can be assigned 0.5 points; The CPU usage rate for high usage can be greater than 70% and less than or equal to 90%. This type indicates that the CPU has a high load, which will affect the performance of the database to be queried, and it can be assigned 0.3 points; The CPU usage rate for excessive usage can be greater than 90%. This type indicates that the CPU load has caused a performance bottleneck, and it can be assigned 0.1 points; 3.2. The memory usage rate types can be divided into sufficient, relatively tight, tight, and severely tight, where: The memory usage rate corresponding to sufficient can be less than or equal to 60%. This type indicates that the memory resources are relatively abundant, and it can be assigned 0.8 points; The memory usage rate corresponding to relatively tight can be greater than 60% and less than or equal to 80%. This type indicates that the memory is starting to become tight, and the memory usage situation needs to be concerned, and it can be assigned 0.5 points; The memory usage rate corresponding to tight can be greater than 80% and less than or equal to 90%. This type indicates that the memory pressure is relatively large, resulting in a decline in the performance of the database to be queried, and it can be assigned 0.3 points; The memory usage rate corresponding to severely tight can be greater than 90%. This type indicates that the memory is insufficient, seriously affecting the performance of the database to be queried, and it can be assigned 0.1 points; 3.3. The disk I / O usage rate types can be divided into low usage rate, medium usage rate, high usage rate, and excessive usage rate, where: The disk I / O usage rate corresponding to low usage can be less than or equal to 30%. This type of disk I / O has a small load, and the database to be queried has good performance, and it can be assigned 0.8 points; The disk I / O usage rate corresponding to medium usage can be greater than 30% and less than or equal to 60%. This type of disk I / O has a certain load, but it is still within the normal range, and it can be assigned 0.5 points; The disk I / O usage rate corresponding to high usage can be greater than 60% and less than or equal to 80%. This type of disk I / O has a high load, which will become a performance bottleneck for the database to be queried, and it can be assigned 0.3 points; The disk I / O utilization rate corresponding to an excessively high utilization rate can be greater than 80%. This type of disk I / O load will seriously affect the performance of the database to be queried, and it can be assigned 0.1 points; 3.4. The network bandwidth utilization rate type can be divided into low utilization rate, medium utilization rate, high utilization rate, and excessively high utilization rate, where: The network bandwidth utilization rate corresponding to a low utilization rate can be less than or equal to 40%. This type of network bandwidth resource is sufficient and will not affect the performance of the database to be queried, and it can be assigned 0.8 points; The network bandwidth utilization rate corresponding to a medium utilization rate can be greater than 40% and less than or equal to 60%. This type of network bandwidth has a certain load but is still within a reasonable range, and it can be assigned 0.5 points; The network bandwidth utilization rate corresponding to a high utilization rate can be greater than 60% and less than or equal to 80%. This type of network bandwidth has a relatively high pressure and will affect the data transmission speed, and it can be assigned 0.3 points; The network bandwidth utilization rate corresponding to an excessively high utilization rate can be greater than 80%. This type of network bandwidth load will cause a performance bottleneck in the database to be queried, and it can be assigned 0.1 points.
[0047] 4. The concurrent connection number of the database to be queried is the number of clients simultaneously connected to the database to be queried, which reflects the number of concurrent users that the database to be queried can support. Its type can be divided into sufficient, normal, busy, and overloaded, where: The concurrent connection number corresponding to sufficient can be less than or equal to 30% of the maximum designed concurrent connection number. This type indicates that the database to be queried still has a large margin for concurrent processing ability, and it can be assigned 0.8 points; The concurrent connection number corresponding to normal can be greater than 30% of the maximum designed concurrent connection number and less than or equal to 70% of the maximum designed concurrent connection number. This type indicates that the database to be queried can normally process concurrent connections but is close to the busy state, and it can be assigned 0.5 points; The concurrent connection number corresponding to busy can be greater than 70% of the maximum designed concurrent connection number and less than or equal to 90% of the maximum designed concurrent connection number. This type indicates that the database to be queried is in a high-concurrency state and its performance is affected to a certain extent, and it can be assigned 0.3 points; The concurrent connection number corresponding to overloaded can be greater than 90% of the maximum designed concurrent connection number. This type indicates that the database to be queried cannot respond to all connection requests in a timely manner, and its performance drops significantly, and it can be assigned 0.1 points.
[0048] 5. The lock waiting duration of the database to be queried is the duration experienced when a transaction waits to acquire a lock held by another transaction, which reflects the concurrent performance of the database to be queried. Its type can be divided into short, medium, long, and very long, where: The short corresponding lock waiting duration can be less than or equal to 1 second. This type has less impact on the performance of the database to be queried and can be assigned a score of 0.8; The medium corresponding lock waiting duration can be greater than 1 second and less than or equal to 3 seconds. This type begins to affect the execution efficiency of transactions and can be assigned a score of 0.5; The long corresponding lock waiting duration can be greater than 3 seconds and less than or equal to 5 seconds. This type significantly affects the concurrency performance of the database to be queried and can be assigned a score of 0.3; The very long corresponding lock waiting duration can be greater than 5 seconds. This type can cause serious problems such as deadlocks and severely affect the performance of the database to be queried and can be assigned a score of 0.1.
[0049] In step 302: 1. For the response duration type of the database to be queried, since it is directly related to the user experience and is a key indicator for measuring the database to be queried, a relatively high weight needs to be set for it. That is, the weight of the response duration type can be set to 0.3.
[0050] 2. For the throughput type of the database to be queried, since it reflects the processing capacity of the database to be queried and is very important for high-load application scenarios, a relatively high weight needs to be set for it. That is, the weight of the throughput type can be set to 0.25.
[0051] 3. For the resource utilization rate type of the database to be queried, its weight can be set to 0.2. Further, it is also necessary to set the weights of the CPU usage rate type, memory usage rate type, disk I / O usage rate type, and network bandwidth usage rate type in the resource utilization rate type respectively: 3.1. For the CPU usage rate type, since it is a key resource for the operation of the database to be queried and has an important impact on the performance of the database to be queried, its weight in the resource utilization rate type can be set to 0.3; 3.2. For the memory usage rate type, since it is also very important for the running stability and performance of the database to be queried, its weight in the resource utilization rate type can be set to 0.3; 3.3. For the disk I / O usage rate type, since it has an important impact on the read and write operations of the database to be queried, its weight in the resource utilization rate type can be set to 0.2; 3.4. For the network bandwidth usage rate type, since it has an important impact on the remote access and data transmission performance of the database to be queried, its weight in the resource utilization rate type can be set to 0.2.
[0052] 4. For the concurrent connection number type of the database to be queried, since it reflects the support ability of the database to be queried for multi-user access and is very important for user-oriented applications, the weight of the concurrent connection number type can be set to 0.15.
[0053] 5. For the lock wait duration type of the database to be queried, since it is an important indicator for measuring the concurrent performance of the database to be queried, although its impact is not as direct as the response duration type and the throughput type, it is also very critical in high-concurrency scenarios. Therefore, the weight of the lock wait duration type can be set to 0.1.
[0054] Suppose the response duration of a certain database to be queried is 2 seconds (i.e., its response duration type is medium), the throughput is 70% of the maximum designed throughput (i.e., its throughput type is medium throughput), the CPU usage rate is 60% (i.e., its CPU usage rate type is medium usage), the memory usage rate is 70% (i.e., its memory usage rate type is relatively tight), the disk I / O usage rate is 40% (i.e., its disk I / O usage rate type is medium usage), the network bandwidth usage rate is 50% (i.e., its network bandwidth usage rate type is medium usage), the concurrent connection number is 60% of the maximum designed concurrent connection number (i.e., its concurrent connection number type is normal), and the lock wait duration is 2 seconds (i.e., its lock wait duration type is medium). Then, first calculate the score of the resource utilization type as follows: ; Then calculate the performance score of the database to be queried as follows: ; The higher the score, the better the performance of the database to be queried.
[0055] It should be noted that the weights of the above performance indicators can be allocated based on different application scenarios and business requirements, or can be determined by methods such as analytic hierarchy process, which is not limited here.
[0056] In this embodiment, when calculating the performance score of the database to be queried, multiple performance indicators including its response duration type, throughput type, resource utilization type, concurrent connection number type, and lock wait duration type are comprehensively considered. And for the resource utilization type, the CPU usage rate type, memory usage rate type, disk I / O usage rate type, and network bandwidth usage rate type it covers are fully considered, so as to cover various situations of the database to be queried. At the same time, by allocating weights to multiple performance indicators according to different application scenarios and business requirements or determining weights by methods such as analytic hierarchy process, it can accurately characterize the impact of the corresponding performance indicators on the performance of the database to be queried, so that the performance score obtained by weighted summation can comprehensively reflect the true performance of the database to be queried and improve the accuracy of performance measurement.
[0057] Figure 4 This is the fourth flowchart of the query task scheduling method provided by the embodiments of the present application. Refer to Figure 4 , in one embodiment, based on the complexity of the multi-cycle query task and the performance of the database to be queried, determining the number of allocable threads for multiple single-cycle query tasks may include: 401. Obtain the complexity level corresponding to the complexity score of the multi-cycle query task; 402. Obtain the query task quantity and query thread ratio corresponding to the complexity level; 403. Based on the query task quantity, query thread ratio, and the performance score of the database to be queried, determine the number of allocable threads for multiple single-cycle query tasks.
[0058] In step 401, the multi-cycle query task may be divided into three levels of high complexity, medium complexity, and low complexity based on the complexity score of the multi-cycle query task. For example, a multi-cycle query task with a complexity score greater than or equal to 0.7 is divided into the high complexity level, a multi-cycle query task with a complexity score greater than 0.2 and less than 0.7 is divided into the medium complexity level, and a multi-cycle query task with a complexity score less than or equal to 0.2 is divided into the low complexity level.
[0059] In step 402, to avoid resource contention among multiple query tasks, the resources may be divided based on the complexity of the query tasks. For example, the query thread ratio corresponding to the high complexity level may be set to 50%, the query thread ratio corresponding to the medium complexity level may be set to 30%, and the query thread ratio corresponding to the low complexity level may be set to 20%. Different ratios may also be set according to their own requirements and actual thread conditions, which are not limited here.
[0060] In step 403, assuming that the complexity score of the multi-cycle query task is 0.35 and the performance score of the database to be queried is 0.5, then the complexity level of the multi-cycle query task is medium, and its corresponding query thread ratio is 30%. When the number of query tasks corresponding to this level is 3 and the maximum designed number of threads is 100, the number of allocable threads for multiple single-cycle query tasks split from this multi-cycle query task , where is the weight of the performance score, and the value range may be , assuming , then , which can be approximately taken as 7, that is, multiple single-cycle query tasks will be executed in parallel on 7 threads.
[0061] When allocating the number of threads for multiple single - cycle query tasks in this embodiment, not only the complexity of the original multi - cycle query tasks is considered, but also the performance of the database to be queried and the number of query tasks with different complexity levels are fully considered. Combining the non - linear formula and the query thread ratio rule, different allocation strategies are adopted according to different situations, and the allocable number of threads for the multiple split single - cycle query tasks is calculated. This method isolates the threads for query tasks with different complexity levels, ensures that a certain number of threads can be allocated for each complexity level, avoids the situation where threads are occupied by high - complexity query tasks, thus affecting the execution of low - complexity query tasks, realizes the fine and reasonable allocation of threads, and improves resource utilization and query efficiency.
[0062] Furthermore, this embodiment also has a dynamic adjustment mechanism. As the database to be queried runs and the business changes, the complexity of the multi - cycle query tasks may change, the performance of the database to be queried may also fluctuate, and the number of query tasks corresponding to different complexity levels will also change. This embodiment can recalculate the allocable number of threads in real - time or regularly according to the changes of these factors to ensure the efficient execution of query tasks.
[0063] Figure 5 It is the fifth flow diagram of the query task scheduling method provided by the embodiments of the present application. Referring to Figure 5 , in one embodiment, based on the allocable number of threads, each single - cycle query task is executed in parallel in the database to be queried, and the single - cycle query results of each single - cycle query task can be obtained, which may include: 501. Create a thread pool; 502. Submit each single - cycle query task to the thread pool; 503. Based on the idle threads in the thread pool with the allocable number of threads, each single - cycle query task is executed in parallel in the database to be queried, and the single - cycle query results of each single - cycle query task are obtained.
[0064] In step 501, after creating the thread pool, the size of the thread pool needs to be reasonably configured according to the calculated allocable number of threads for multiple single - cycle query tasks, and parameters such as the maximum number of threads that can be accommodated, the minimum number of threads, and the thread idle duration in the thread pool are determined. This thread pool can be used to manage the internal threads.
[0065] In steps 502 to 503, each single - cycle query task can be submitted to the thread pool one by one. Assuming the allocable number of threads is 7, when the thread pool receives a query task request, it will allocate threads to execute the query task according to the current state of the thread pool, that is, whether there are idle threads: If there are currently idle threads in the thread pool and the number of idle threads is greater than or equal to 7, then directly allocate 7 idle threads to execute each single - cycle query task in parallel; If there are currently idle threads in the thread pool, but the number of idle threads is less than 7, for example, there are only 5 idle threads, and the current number of threads in the thread pool has not reached the maximum number of threads, then directly allocate these 5 idle threads and create 2 new threads to form 7 idle threads for parallel execution of each single-cycle query task; If there are currently no idle threads in the thread pool and the current number of threads in the thread pool has not reached the maximum number of threads, then create 7 new threads for parallel execution of each single-cycle query task; If the current number of threads in the thread pool has reached the maximum number of threads, the unexecuted single-cycle query tasks will wait in the queue until idle threads are available.
[0066] It should be noted that any thread in the thread pool is responsible for a single-cycle query task. When any thread is assigned a single-cycle query task, it will initiate a query request to the database to be queried, execute the single-cycle query task in the database to be queried, and the database to be queried retrieves and processes data according to the query statement of the single-cycle query task, obtains the query result, and returns the query result to the thread.
[0067] Before creating the thread pool, according to the running environment and requirements of the application, an appropriate local cache technology can be selected to create a local cache. For example, use ConcurrentHashMap in Java to create a local cache named taskACache in the memory space of the application, set the initial capacity of this local cache to 1000, and set the load factor to the default value of 0.75. Then initialize this local cache, and some default values or flags can be set so that the subsequent operations can accurately identify the status of the cache. For example, set a key-value pair with the key "status" and the value "initializing" to indicate that the cache is in the initialization phase.
[0068] After each thread obtains the query result of its single-cycle query task, the query result can be recorded in the local cache taskACache after initialization. When recording data, it is necessary to ensure that the data format matches the cache storage structure, and add necessary identification information to each record, such as the corresponding single-cycle identifier (such as date), so that different cycle data can be accurately identified and distinguished during subsequent aggregation.
[0069] Furthermore, the execution status of threads can also be monitored through the relevant mechanisms of the thread pool to determine whether all single-cycle query tasks have been executed completely. For example, the waiting methods provided by the thread pool can be used, such as the awaitTermination method in Java, to set a waiting duration, such as 300 seconds, and continuously check whether all threads in the thread pool have been executed completely within 30 seconds. If all threads have been executed completely within 30 seconds, the next aggregation operation can be carried out; if there are still threads that have not been executed completely after 30 seconds, corresponding measures need to be taken, such as checking the task status of the threads that have not been executed completely and whether there are database connection problems, to ensure that all query tasks can be executed completely eventually.
[0070] In this embodiment, the relevant mechanisms of the thread pool are used to allocate idle threads to each single-cycle query task to ensure the parallel execution of each single-cycle query task, and the execution status of the threads is monitored to ensure that all query tasks can be executed completely, so that the query results of each single-cycle query task can be obtained efficiently. At the same time, by creating a local cache to record each query result, it is ensured that each query result can be obtained accurately and without error in the subsequent aggregation operation.
[0071] Figure 6 is the sixth flow diagram of the query task scheduling method provided by the embodiment of the present application. Refer to Figure 6 , in one embodiment, aggregating each single-cycle query result to obtain the query result of the multi-cycle query task may include: 601. Group the query results of each single cycle according to the query fields to obtain multiple groups of query results; 602. Calculate the query results within each group according to the query target to obtain multiple groups of calculation results; 603. Adjust the format of each group of calculation results to obtain the query result of the multi-cycle query task.
[0072] In step 601, that is, obtain all the recorded query results from the local cache taskACache. Assume that the multi-cycle query task requires calculating the total number of page views in different provinces and different pages from January 1, 2024 to October 31, 2024. Then the query fields are province and page. According to the single-cycle identifier added before, assume it is day, then group the query results of each day according to province and page.
[0073] In step 602, the query target of this multi-cycle query task is the total number of page views, so sum the number of page views within each group to obtain the total number of page views corresponding to this multi-cycle query task.
[0074] In step 603, the obtained total number of views can be formatted, including data type conversion, field order adjustment, etc., to make it conform to the format expected by the multi-cycle query task. Finally, through methods such as function call return values and message queue transmission, the formatted total number of views is accurately transmitted to the multi-cycle query task based on the convention between the multi-cycle query task and the execution environment, completing the entire query process.
[0075] In this embodiment, the single-cycle query results are first grouped according to the query fields, then the query results within the group are calculated according to the query target, and finally the calculation results are formatted, so as to complete the aggregation of the single-cycle query results and efficiently obtain the multi-cycle query results based on the single-cycle query results.
[0076] The query task scheduling device provided by the embodiments of the present application will be described below. The query task scheduling device described below can be correspondingly referred to the query task scheduling method described above.
[0077] Figure 7 is a schematic structural diagram of the query task scheduling device provided by the embodiments of the present application. Refer to Figure 7 , the embodiments of the present application provide a query task scheduling device, which may include: A query task splitting module 701, configured to: split a multi-cycle query task into multiple single-cycle query tasks; A query thread allocation module 702, configured to: determine the number of allocable threads for the multiple single-cycle query tasks based on the complexity of the multi-cycle query task and the performance of the database to be queried; A query result obtaining module 703, configured to: based on the number of allocable threads, execute each single-cycle query task in parallel in the database to be queried, and obtain the single-cycle query results of each single-cycle query task; A query result aggregation module 704, configured to: aggregate the single-cycle query results to obtain the query results of the multi-cycle query task.
[0078] The query task scheduling device provided in this embodiment splits a multi-cycle query task into multiple single-cycle query tasks, determines the allocable number of threads for the multiple single-cycle query tasks based on the complexity of the multi-cycle query task and the performance of the database to be queried, and based on the allocable number of threads, parallelly executes each single-cycle query task in the database to be queried to obtain the single-cycle query results of each single-cycle query task, and aggregates the single-cycle query results of each single-cycle query task to obtain the query result of the multi-cycle query task. Whether this embodiment faces a single multi-cycle query task or multiple multi-cycle query tasks, each multi-cycle query task is split into multiple single-cycle query tasks, so that the data volume of the single-cycle query task is greatly reduced compared with the multi-cycle query task, and the operation of the single-cycle query task is also greatly simplified compared with the multi-cycle query task. Then, based on the complexity of each multi-cycle query task and the performance of the database to be queried, threads are allocated to the multiple split single-cycle query tasks for parallel query and query result aggregation, so that each thread parallelly executes small-volume and simple single-cycle query tasks, thereby greatly improving the query efficiency and thread utilization rate, avoiding waste of computing resources. At the same time, since the improvement of the query efficiency will significantly reduce the occupation duration of computing resources, it is possible to reduce the impact on other query tasks and improve the user query experience.
[0079] In one embodiment, it further includes a complexity determination module (not shown in the figure) for: Obtain the scores of multiple complexity metrics of the multi-cycle query task; Perform weighted summation on the scores of the multiple complexity metrics to obtain the complexity score of the multi-cycle query task; The multiple complexity metrics include the length type, nesting type, connection type, aggregation function type, condition type, and index usage type corresponding to the query statement of the multi-cycle query task.
[0080] In one embodiment, it further includes a performance determination module (not shown in the figure) for: Obtain the scores of multiple performance metrics of the database to be queried; Perform weighted summation on the scores of the multiple performance metrics to obtain the performance score of the database to be queried; The multiple performance metrics include the response duration type, throughput type, resource utilization type, concurrent connection number type, and lock waiting duration type of the database to be queried; The resource utilization type is comprehensively obtained based on the CPU usage type, memory usage type, disk I / O usage type, and network bandwidth usage type of the database to be queried.
[0081] In one embodiment, the query thread allocation module 702 is specifically used for: Obtain the complexity level corresponding to the complexity score of the multi-cycle query task; Obtain the number of query tasks and the query thread ratio corresponding to the complexity level; Based on the number of query tasks, the query thread ratio, and the performance score of the database to be queried, determine the number of allocable threads for the multiple single-cycle query tasks.
[0082] In one embodiment, the query result acquisition module 703 is specifically configured to: Create a thread pool; Submit each single-cycle query task to the thread pool; Based on the idle threads with the allocable number of threads in the thread pool, execute each single-cycle query task in parallel in the database to be queried, and obtain the single-cycle query results of each single-cycle query task.
[0083] In one embodiment, the query result aggregation module 704 is specifically configured to: Group each single-cycle query result according to the query field to obtain multiple groups of query results; Calculate the query results within each group according to the query target to obtain multiple groups of calculation results; Adjust the format of each group of calculation results to obtain the query result of the multi-cycle query task.
[0084] Figure 8 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 8 shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 may call a computer program in the memory 830 to execute the steps of the query task scheduling method, for example, including: Split a multi-cycle query task into multiple single-cycle query tasks; Based on the complexity of the multi-cycle query task and the performance of the database to be queried, determine the number of allocable threads for the multiple single-cycle query tasks; Based on the number of allocable threads, execute each single-cycle query task in parallel in the database to be queried, and obtain the single-cycle query results of each single-cycle query task; Aggregate each single-cycle query result to obtain the query result of the multi-cycle query task.
[0085] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0086] On the other hand, an embodiment of this application also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the steps of the query task scheduling method provided in the above-mentioned various embodiments, for example, including: Split a multi-cycle query task into multiple single-cycle query tasks; Based on the complexity of the multi-cycle query task and the performance of the database to be queried, determine the number of assignable threads for the multiple single-cycle query tasks; Based on the number of assignable threads, parallelly execute each single-cycle query task in the database to be queried to obtain the single-cycle query results of each single-cycle query task; Aggregate the single-cycle query results of each single-cycle query task to obtain the query result of the multi-cycle query task.
[0087] On the other hand, an embodiment of this application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. The computer program is used to cause a processor to execute the steps of the query task scheduling method provided in the above-mentioned various embodiments, for example, including: Split a multi-cycle query task into multiple single-cycle query tasks; Based on the complexity of the multi-cycle query task and the performance of the database to be queried, determine the number of assignable threads for the multiple single-cycle query tasks; Based on the number of assignable threads, parallelly execute each single-cycle query task in the database to be queried to obtain the single-cycle query results of each single-cycle query task; Aggregate the single-cycle query results of each single-cycle query task to obtain the query result of the multi-cycle query task.
[0088] The non-transitory computer-readable storage medium may be any available medium or data storage device accessible by the processor, including but not limited to magnetic memories (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical memories (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor memories (such as ROMs, EPROMs, EEPROMs, non-volatile memories (NAND FLASH), solid state drives (SSD)), etc.
[0089] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0090] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disks, optical disks, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A query task scheduling method, characterized in that: include: Split a multi-cycle query task into multiple single-cycle query tasks; Determining the number of allocatable threads for the multiple single-cycle query tasks based on the complexity of the multi-cycle query tasks and the performance of the database to be queried; Based on the number of allocatable threads, executing the single-cycle query tasks in parallel in the database to be queried, and obtaining the single-cycle query results of the single-cycle query tasks; Aggregate the single-cycle query results to obtain the query result of the multi-cycle query task.
2. The query task scheduling method according to claim 1, characterized in that: The complexity of the multi-cycle query task is determined based on the following method: Obtaining scores of multiple complexity indicators of the multi-period query task; Performing a weighted summation on the scores of the multiple complexity indicators to obtain a complexity score of the multi-period query task; The multiple complexity indicators include length type, nesting type, connection type, aggregate function type, condition type and index usage type corresponding to the query statement of the multi-cycle query task.
3. The query task scheduling method according to claim 1, characterized in that: The performance of the database to be queried is determined based on the following method: Obtaining scores of multiple performance indicators of the database to be queried; Performing a weighted summation on the scores of the multiple performance indicators to obtain a performance score of the database to be queried; The multiple performance indicators include the response time type, throughput type, resource utilization type, concurrent connection number type and lock waiting time type of the database to be queried; The resource utilization type is obtained based on the CPU utilization type, memory utilization type, disk I / O utilization type and network bandwidth utilization type of the database to be queried.
4. The query task scheduling method according to claim 1, characterized in that: The determining, based on the complexity of the multi-cycle query task and the performance of the database to be queried, the number of threads that can be allocated to the multiple single-cycle query tasks includes: Obtaining a complexity level corresponding to the complexity score of the multi-period query task; Obtain the number of query tasks and query thread ratio corresponding to the complexity level; Based on the number of query tasks, the query thread ratio, and the performance score of the database to be queried, the number of threads that can be allocated to the multiple single-cycle query tasks is determined.
5. The query task scheduling method according to claim 1, characterized in that: The executing of each single-cycle query task in parallel in the to-be-queried database based on the number of allocatable threads to obtain a single-cycle query result of each single-cycle query task includes: Create a thread pool; Submitting each single-cycle query task to the thread pool; Based on the idle threads of the allocatable number of threads in the thread pool, each single-cycle query task is executed in parallel in the database to be queried to obtain a single-cycle query result of each single-cycle query task.
6. The query task scheduling method according to claim 1, characterized in that: The step of aggregating the single-cycle query results to obtain the query result of the multi-cycle query task includes: Grouping each single-cycle query result according to the query field to obtain multiple groups of query results; Calculate the query results in each group according to the query target to obtain multiple groups of calculation results; The format of each group of calculation results is adjusted to obtain the query result of the multi-period query task.
7. A query task scheduling device, characterized in that: include: A query task splitting module is used to split a multi-cycle query task into multiple single-cycle query tasks; A query thread allocation module, used to: determine the number of allocable threads for the multiple single-cycle query tasks based on the complexity of the multi-cycle query tasks and the performance of the database to be queried; A query result acquisition module, used to: execute each single-cycle query task in parallel in the to-be-queried database based on the number of allocatable threads, and obtain a single-cycle query result of each single-cycle query task; The query result aggregation module is used to aggregate the single-cycle query results to obtain the query result of the multi-cycle query task.
8. An electronic device comprising a processor and a memory storing a computer program, characterized in that: When the processor executes the computer program, the steps of the query task scheduling method according to any one of claims 1 to 6 are implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the query task scheduling method according to any one of claims 1 to 6 are implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the query task scheduling method according to any one of claims 1 to 6 are implemented.