Data query request processing method and device, equipment and storage medium

By matching the priority and querying thresholds based on the metadata information requested by the data query, counting the current query rate per second and parallelism, the problem of concurrent traffic control in the database system is solved, and efficient traffic management and user experience improvement is achieved.

CN120128636APending Publication Date: 2025-06-10BIGO TECH PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510146378.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing technology cannot effectively control the concurrent traffic of the database system, resulting in a decline in service quality, and may even cause abnormal system crashes and unavailability of services, and have poor user experience.

Method used

By obtaining the pending data query request, matching the priority according to its metadata information, and querying the query rate threshold per second and the parallelism threshold. Statistics the current query rate per second and parallelism, and executes data query requests when the threshold allows to adapt to flow control and parallelism management of different priorities.

Benefits of technology

Effectively prevent overloading system resources caused by traffic peaks, dynamically adjust the number of parallel tasks, improve service quality and user experience, and avoid system crashes and unavailability of services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128636A_ABST
    Figure CN120128636A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data query request processing method and device, equipment and a storage medium, and the method comprises the steps: obtaining a to-be-processed data query request, matching a corresponding priority according to metadata information of the data query request, and querying a per-second query rate threshold and a parallelism degree threshold corresponding to the priority; counting a first query rate per second currently corresponding to the user identifier and a second query rate per second currently corresponding to the cluster identifier, and when the first query rate per second is smaller than an individual query rate per second threshold value and the second query rate per second is smaller than a cluster query rate per second threshold value, sending the cluster identifier to the user identifier; counting a cluster parallelism degree currently corresponding to the cluster identifier and a personal parallelism degree currently corresponding to the user identifier; under the condition that the cluster parallelism degree is smaller than the first cluster parallelism degree threshold value and the individual parallelism degree is smaller than the individual parallelism degree threshold value, the data query request is executed. According to the scheme, flow control is carried out by adapting to different priorities of the data query requests, and the concurrent flow of a database system is effectively controlled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technologies, and in particular, to a method, apparatus, device, and storage medium for processing data query requests. Background Art

[0002] In today's big data era, due to data characteristics such as explosive growth of data volume and variety of data types, the collection, storage, analysis, and application of data have gradually become the core competitiveness of all industries. Among them, high-performance databases have become an indispensable component in enterprises. While continuously developing and popularizing, they are also facing more and more challenges. To ensure system stability, it is necessary to perform traffic control and management on data query requests accessing the database, so as to avoid system crashes or delays caused by excessive requests, thereby protecting server resources, optimizing user experience, and improving system security.

[0003] In related technologies, access traffic control of a database system is performed by restricting the number of accesses per user per unit time, which cannot effectively control the concurrent traffic of the database system, resulting in a decline in service quality, and even abnormal situations such as system crashes and service unavailability may occur, and the user experience is poor. Summary of the Invention

[0004] The embodiments of the present application provide a method, apparatus, device, and storage medium for processing data query requests, which solve the problem in related technologies that the concurrent traffic of the database system cannot be effectively controlled, resulting in a decline in service quality, and even abnormal situations such as system crashes and service unavailability may occur, and the user experience is poor. The method adapts to different priorities of data query requests for traffic control, flexibly adjusts the allowed amount of requests according to the actual request situations of the cluster and users, effectively prevents overload of system resources caused by traffic peaks, dynamically adjusts the number of parallel tasks, effectively controls the concurrent traffic of the database system, and is beneficial to improving service quality and user experience.

[0005] In a first aspect, the embodiments of the present application provide a method for processing a data query request, and the method includes:

[0006] Obtain a data query request to be processed, match a corresponding priority according to the metadata information of the data query request, and query the query rate threshold per second and the parallelism threshold corresponding to the priority. The metadata information includes a user identifier and a cluster identifier. The query rate threshold per second includes a personal query rate threshold per second and a cluster query rate threshold per second. The parallelism threshold includes a personal parallelism threshold and a first cluster parallelism threshold;

[0007] Statistically calculate the current first queries per second rate corresponding to the user identifier and the current second queries per second rate corresponding to the cluster identifier. When the first queries per second rate is less than the personal queries per second rate threshold and the second queries per second rate is less than the cluster queries per second rate threshold, statistically calculate the current cluster parallelism corresponding to the cluster identifier and the current personal parallelism corresponding to the user identifier;

[0008] When the cluster parallelism is less than the first cluster parallelism threshold and the personal parallelism is less than the personal parallelism threshold, execute the data query request.

[0009] In a second aspect, an embodiment of the present application further provides a data query request processing device, which includes:

[0010] A first threshold determination module, configured to obtain a data query request to be processed, match a corresponding priority according to the metadata information of the data query request, and query the queries per second rate threshold and the parallelism threshold corresponding to the priority. The metadata information includes a user identifier and a cluster identifier. The queries per second rate threshold includes a personal queries per second rate threshold and a cluster queries per second rate threshold. The parallelism threshold includes a personal parallelism threshold and a first cluster parallelism threshold;

[0011] A queries per second rate verification module, configured to statistically calculate the current first queries per second rate corresponding to the user identifier and the current second queries per second rate corresponding to the cluster identifier. When the first queries per second rate is less than the personal queries per second rate threshold and the second queries per second rate is less than the cluster queries per second rate threshold, statistically calculate the current cluster parallelism corresponding to the cluster identifier and the current personal parallelism corresponding to the user identifier;

[0012] A first parallelism verification module, configured to execute the data query request when the cluster parallelism is less than the first cluster parallelism threshold and the personal parallelism is less than the personal parallelism threshold.

[0013] In a third aspect, an embodiment of the present application further provides a data query request processing device, which includes:

[0014] One or more processors;

[0015] A storage device, configured to store one or more programs,

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the data query request processing method described in the embodiments of the present application.

[0017] Fourthly, an embodiment of the present application further provides a non-volatile storage medium storing computer-executable instructions, and the computer-executable instructions are configured to execute the data query request processing method described in the embodiment of the present application when executed by a computer processor.

[0018] Fifthly, an embodiment of the present application further provides a computer program product, which includes a computer program. The computer program is stored in a computer-readable storage medium, and at least one processor of the device reads and executes the computer program, so that the device executes the data query request processing method described in the embodiment of the present application.

[0019] In the embodiment of the present application, by obtaining a data query request to be processed, matching a corresponding priority according to the metadata information of the data query request, and querying the query rate threshold per second and the parallelism threshold corresponding to the priority, where the metadata information includes a user identifier and a cluster identifier, the query rate threshold per second includes a personal query rate threshold per second and a cluster query rate threshold per second, and the parallelism threshold includes a personal parallelism threshold and a first cluster parallelism threshold; counting the first query rate per second corresponding to the user identifier currently and the second query rate per second corresponding to the cluster identifier currently, and when the first query rate per second is less than the personal query rate threshold per second and the second query rate per second is less than the cluster query rate threshold per second, counting the cluster parallelism corresponding to the cluster identifier currently and the personal parallelism corresponding to the user identifier currently; and when the cluster parallelism is less than the first cluster parallelism threshold and the personal parallelism is less than the personal parallelism threshold, executing the data query request. In the above solution, by matching a corresponding priority according to the metadata information of the data query request, the data query requests to be processed can be pre-divided according to different priorities, and the query rate threshold per second and the parallelism threshold corresponding to the priority are queried to adapt to the traffic control of the data query requests with different priorities, ensuring the priority processing of the data query requests with high priorities. By respectively verifying the first query rate per second corresponding to the user identifier and the second query rate per second corresponding to the cluster identifier, the allowable amount of the requests can be flexibly adjusted according to the actual request conditions of the cluster and the user, effectively preventing the overload of system resources caused by traffic peaks. By respectively verifying the personal parallelism corresponding to the user identifier and the cluster parallelism corresponding to the cluster identifier, parallelism control in different dimensions is performed, and the number of parallel tasks is dynamically adjusted, effectively controlling the concurrent traffic of the database system, which is beneficial to improving the service quality and user experience. Description of the Drawings

[0020] Figure 1 It is a flowchart of a data query request processing method provided by an embodiment of the present application;

[0021] Figure 2Flowchart of a data query request processing method provided by an embodiment of the present application, which includes a process of combining the parallelism threshold of the second cluster and a preset ratio threshold to judge request processing;

[0022] Figure 3 Flowchart of a data query request processing method provided by an embodiment of the present application, which includes a process of combining the request error rate and the number of request errors to judge request processing;

[0023] Figure 4 Flowchart of a data query request processing method provided by an embodiment of the present application, which includes a process of judging whether a data query request is a duplicate invalid request;

[0024] Figure 5 Flowchart of a data query request processing method provided by an embodiment of the present application, which includes a process of delaying the processing of data query requests;

[0025] Figure 6 Flowchart of a data query request processing method provided by an embodiment of the present application, which includes a process of rejecting data query requests;

[0026] Figure 7 Block diagram of a data query request processing device provided by an embodiment of the present application;

[0027] Figure 8 Structural schematic diagram of a data query request processing device provided by an embodiment of the present application. Detailed implementation manners

[0028] The following further elaborates on the embodiments of the present application in conjunction with the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the embodiments of the present application, rather than limiting the embodiments of the present application. Additionally, it should be noted that for the sake of description, only parts related to the embodiments of the present application are shown in the drawings, rather than all structures.

[0029] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.

[0030] The data query request processing method provided by the embodiments of the present application can be applied to the access scenario of a database, and can adapt to different priorities of data query requests for database access traffic control. First, by respectively verifying the first queries per second rate corresponding to the user identifier and the second queries per second rate corresponding to the cluster identifier, the allowable amount of requests can be flexibly adjusted based on the actual request situations of the cluster and the user. Then, by respectively verifying the personal parallelism corresponding to the user identifier and the cluster parallelism corresponding to the cluster identifier, the number of parallel tasks can be dynamically adjusted, effectively alleviating the access pressure on the database.

[0031] In the data query request processing method provided by the embodiments of the present application, the execution subject of each step can be a computer device, which refers to any electronic device with data calculation, processing, and storage capabilities, such as a server and other devices. The embodiments of the present application do not make any limitations in this regard.

[0032] Figure 1 It is a flowchart of a data query request processing method provided by the embodiments of the present application. As Figure 1 shown, it includes the following steps:

[0033] Step S101: Obtain the data query request to be processed, match the corresponding priority according to the metadata information of the data query request, and query the queries per second rate threshold and parallelism threshold corresponding to the priority. Among them, the metadata information includes a user identifier and a cluster identifier, the queries per second rate threshold includes a personal queries per second rate threshold and a cluster queries per second rate threshold, and the parallelism threshold includes a personal parallelism threshold and a first cluster parallelism threshold.

[0034] Among them, the data query request can be a request sent by the client to retrieve or extract target data that meets specific conditions from the database. The metadata information of the data query request may include a user identifier, a cluster identifier, a source identifier, and a database table identifier, etc., wherein the user identifier is used to identify the user who initiated the data query request, the cluster identifier is used to identify the target server cluster in the database system that needs to process the data query request, the source identifier is used to identify the source data end of the data query request, for example, different business platforms, etc., and the database table identifier is used to identify the target database table corresponding to the target data that needs to be queried by the data query request. In one embodiment, one or more of the user identifier, cluster identifier, source identifier, and database table identifier can be used for priority matching, which is not limited in this application. Taking the use of user ID, cluster ID, source ID and database table ID for priority matching as an example, developers can distinguish different users, different clusters, different source ends and different database tables with different levels of importance in advance according to the business characteristics of the actual application scenario, and build corresponding priority comparison tables respectively. Thus, according to the user ID, cluster ID, source end ID and database table ID of the data query request to be processed, the corresponding priority comparison table can be queried respectively to determine multiple candidate priorities, and the largest one among the multiple candidate priorities is selected as the final priority corresponding to the data query request. In addition, each priority is pre-set with a corresponding query rate threshold per second and parallelism threshold. The query rate threshold per second includes a personal query rate threshold per second and a cluster query rate threshold per second. The parallelism threshold includes a personal parallelism threshold and a first cluster parallelism threshold. The personal query rate threshold per second is used to determine whether the number of query requests currently triggered by the user per second exceeds the limit. The cluster query rate threshold per second is used to determine whether the number of query requests currently processed per second by the target server cluster that needs to process the data query request exceeds the limit. The personal parallelism threshold is used to determine whether the number of query requests currently triggered by the user in the execution state exceeds the limit. The first cluster parallelism threshold is used to determine whether the number of query requests currently processed by the target server cluster in the execution state exceeds the limit. The personal query rate threshold per second, cluster query rate threshold per second, personal parallelism threshold and first cluster parallelism threshold set for different priorities are different. The higher the priority, the larger the corresponding personal query rate threshold per second, cluster query rate threshold per second, personal parallelism threshold and first cluster parallelism threshold are. The lower the priority, the smaller the corresponding personal query rate threshold per second, cluster query rate threshold per second, personal parallelism threshold and first cluster parallelism threshold are. The values ​​of the personal query rate threshold per second, cluster query rate threshold per second, personal parallelism threshold and first cluster parallelism threshold for specific priority settings can be set by the developer according to the business characteristics of the actual application scenario and the characteristics of the database system, and are not limited in this application.

[0035] Step S102: Statistically calculate the first queries per second corresponding to the current user identifier and the second queries per second corresponding to the current cluster identifier. When the first queries per second is less than the personal queries per second threshold and the second queries per second is less than the cluster queries per second threshold, statistically calculate the cluster parallelism corresponding to the current cluster identifier and the personal parallelism corresponding to the current user identifier.

[0036] Among them, the first queries per second can be used to represent the number of query requests triggered by the user per second currently, and the second queries per second can be used to represent the number of query requests processed by the target server cluster per second currently. If the first queries per second is less than the personal queries per second threshold, it can be considered that the user's query traffic is still within the acceptable range. If the first queries per second is greater than or equal to the personal queries per second threshold, it can be considered that the user's query traffic exceeds the expected limit and may cause instantaneous traffic overload. If the second queries per second is less than the cluster queries per second threshold, it can be considered that the query traffic of the target server cluster is still within the acceptable range. If the second queries per second is greater than or equal to the cluster queries per second threshold, it can be considered that the query traffic of the target server cluster exceeds the expected limit and may also cause instantaneous traffic overload. Thus, the personal and global query traffic can be comprehensively considered to more comprehensively determine whether the data query request can be continued to be processed. When the first queries per second is less than the personal queries per second threshold and the second queries per second is less than the cluster queries per second threshold, it can be considered that both the personal and global query traffic are within the acceptable range, and the subsequent parallelism verification process can be continued to determine whether to execute the data query request. When the first queries per second is greater than or equal to the personal queries per second threshold, or the second queries per second is greater than or equal to the cluster queries per second threshold, it can be considered that the personal or global query traffic exceeds the limit and may cause system overload, and the data query request can be rejected.

[0037] Step S103: When the cluster parallelism is less than the first cluster parallelism threshold and the personal parallelism is less than the personal parallelism threshold, execute the data query request.

[0038] Among them, the individual parallelism can be used to characterize the number of query requests currently triggered by the user that are in the execution state, and the cluster parallelism can be used to characterize the number of query requests currently processed by the target server cluster that are in the execution state. If the cluster parallelism is less than the first cluster parallelism threshold, it can be considered that the concurrent traffic of the target server cluster is still within the acceptable range. If the cluster parallelism is greater than or equal to the first cluster parallelism threshold, it can be considered that the concurrent traffic of the target server cluster exceeds the expected limit and may cause overload. If the individual parallelism is less than the individual parallelism threshold, it can be considered that the concurrent traffic triggered by the user is still within the acceptable range. If the individual parallelism is greater than or equal to the individual parallelism threshold, it can be considered that the concurrent traffic triggered by the user exceeds the expected limit and may also cause overload. Thus, the concurrent traffic of both the individual and the global can be comprehensively considered to effectively determine whether the data query request can be continued to be processed. When the cluster parallelism is less than the first cluster parallelism threshold and the individual parallelism is less than the individual parallelism threshold, it can be considered that the concurrent traffic of both the individual and the global is within the acceptable range, and the system is not in an overloaded state, and the data query request can be continued to be executed. When the cluster parallelism is greater than or equal to the first cluster parallelism threshold, or the individual parallelism is greater than or equal to the individual parallelism threshold, it can be considered that the concurrent traffic of the individual or the global exceeds the limit, and the system is likely to be in an overloaded state, and the data query request is rejected.

[0039] As described above, by obtaining a data query request to be processed, matching the corresponding priority according to the metadata information of the data query request, and querying the query rate threshold per second and the parallelism threshold corresponding to the priority, where the metadata information includes a user identifier and a cluster identifier, the query rate threshold per second includes a personal query rate threshold per second and a cluster query rate threshold per second, and the parallelism threshold includes a personal parallelism threshold and a first cluster parallelism threshold; counting the first query rate per second corresponding to the current user identifier and the second query rate per second corresponding to the current cluster identifier, and when the first query rate per second is less than the personal query rate threshold per second and the second query rate per second is less than the cluster query rate threshold per second, counting the cluster parallelism corresponding to the current cluster identifier and the personal parallelism corresponding to the current user identifier; when the cluster parallelism is less than the first cluster parallelism threshold and the personal parallelism is less than the personal parallelism threshold, executing the data query request. In the above solution, by matching the corresponding priority according to the metadata information of the data query request, the data query requests to be processed can be pre-divided according to different priorities, and the query rate threshold per second and the parallelism threshold corresponding to the priority can be queried to adapt to the traffic control of the data query requests with different priorities, ensuring the priority processing of the high-priority data query requests. By separately verifying the first query rate per second corresponding to the user identifier and the second query rate per second corresponding to the cluster identifier, the allowable amount of the request can be flexibly adjusted according to the actual request conditions of the cluster and the user, effectively preventing the overload of the system resources caused by the traffic peak. By separately verifying the personal parallelism corresponding to the user identifier and the cluster parallelism corresponding to the cluster identifier, performing parallelism control in different dimensions, dynamically adjusting the number of parallel tasks, and effectively controlling the concurrent traffic of the database system, which is beneficial to improving the service quality and user experience.

[0040] Figure 2 The figure is a flowchart of a data query request processing method provided by an embodiment of the present application, which includes a process of judging a request process by combining a second cluster parallelism threshold and a preset ratio threshold. As Figure 2 shown, it includes the following steps:

[0041] Step S201: Obtain a data query request to be processed, match the corresponding priority according to the metadata information of the data query request, and query the query rate threshold per second and the parallelism threshold corresponding to the priority, where the metadata information includes a user identifier and a cluster identifier, the query rate threshold per second includes a personal query rate threshold per second and a cluster query rate threshold per second, and the parallelism threshold includes a personal parallelism threshold and a first cluster parallelism threshold.

[0042] Step S202: Statistically calculate the first queries per second corresponding to the current user identifier and the second queries per second corresponding to the current cluster identifier. When the first queries per second is less than the personal queries per second threshold and the second queries per second is less than the cluster queries per second threshold, statistically calculate the cluster parallelism corresponding to the current cluster identifier, and query the second cluster parallelism threshold and the preset ratio threshold corresponding to the query priority.

[0043] Among them, the second cluster parallelism threshold can be used to preliminarily determine whether the number of query requests currently being processed by the target server cluster in the execution state is in a high state and whether the system load is too high. The preset ratio threshold can be used to determine whether the average execution speed of the system for processing requests has slowed down. The second cluster parallelism threshold and the preset ratio threshold set for different priorities are different. The higher the priority, the larger the set second cluster parallelism threshold and the smaller the preset ratio threshold. The lower the priority, the smaller the set second cluster parallelism threshold and the larger the preset ratio threshold.

[0044] Step S203: When the cluster parallelism is less than the second cluster parallelism threshold, calculate the first proportion of the number of data query requests whose execution time is less than the reference execution time within a preset time range relative to the total number of requests.

[0045] Among them, if the cluster parallelism is greater than or equal to the second cluster parallelism threshold, it can be considered that the concurrent traffic of the system is in a high state and the system load is relatively high. If the cluster parallelism is less than the second cluster parallelism threshold, it can be considered that the concurrent traffic of the system is in a reasonable state and the system load is relatively stable. Thus, when the cluster parallelism is greater than or equal to the second cluster parallelism threshold, the data query request can be rejected. When the cluster parallelism is less than the second cluster parallelism threshold, the average speed change of the system processing the request can be continuously determined in combination with the request execution time. The execution time is the duration from receiving a certain data query request to processing and completing the data query request. The reference execution time can be determined based on the execution time data of the system history. For example, it can be the 90th percentile of the execution times corresponding to all data query requests in the past year. This application does not make any limitations here. The preset time range can be within the most recent 5 minutes, or within the most recent 10 minutes, etc. By counting the first proportion of the number of data query requests with execution times less than the reference execution time within the preset time range relative to the total number of requests, the execution time distribution of the data query requests within the preset time range can be determined. If the first proportion is less than the preset proportion threshold, it can be considered that the average execution time of the data query requests within the preset time range is relatively long, the average speed of the system processing requests is slow, and the load may be relatively high. If the first proportion is greater than or equal to the preset proportion threshold, it can be considered that the average execution time of the data query requests within the preset time range is relatively short, the average speed of the system processing requests is fast, and the load is relatively stable. Specifically, the values of the second cluster parallelism threshold, the preset proportion threshold, and the reference execution time for different priority settings can be set by developers according to the business characteristics of the actual application scenario and the characteristics of the database system. This application does not make any limitations here.

[0046] Step S204, when the first proportion is greater than or equal to the preset proportion threshold, count the personal parallelism currently corresponding to the user identifier.

[0047] Among them, when the first proportion is less than the preset proportion threshold, the data query request can be rejected. When the first proportion is greater than or equal to the preset proportion threshold, different-dimensional parallelism verification can be continued to determine whether to execute the data query request.

[0048] Step S205, when the cluster parallelism is less than the first cluster parallelism threshold and the personal parallelism is less than the personal parallelism threshold, execute the data query request.

[0049] As described above, by comparing the cluster parallelism with the second cluster parallelism threshold, it is possible to preliminarily determine whether the system is in a state of high concurrent traffic. Additionally, by calculating the first proportion of the number of data query requests with execution times less than the reference execution time within a preset time range relative to the total number of requests, it is possible to determine the execution time distribution of the requests currently being processed by the system. Furthermore, it is possible to determine whether there are problems such as a slowdown in processing speed and high load in the system, allocate system resources reasonably, and ensure the priority processing of important tasks while effectively controlling traffic.

[0050] Figure 3 This is a flowchart of a data query request processing method provided by an embodiment of the present application, which includes a process of judging request processing by combining the request error rate and the number of request errors. As Figure 3 shown, it includes the following steps:

[0051] Step S301: Obtain the data query request to be processed, match the corresponding priority according to the metadata information of the data query request, and query the corresponding queries per second threshold and parallelism threshold for the priority. Among them, the metadata information includes the user identifier and the cluster identifier. The queries per second threshold includes the personal queries per second threshold and the cluster queries per second threshold, and the parallelism threshold includes the personal parallelism threshold and the first cluster parallelism threshold.

[0052] Step S302: Statistically calculate the first queries per second corresponding to the current user identifier and the second queries per second corresponding to the current cluster identifier. When the first queries per second is less than the personal queries per second threshold and the second queries per second is less than the cluster queries per second threshold, statistically calculate the request error rate and the number of request errors corresponding to the user identifier within a preset time range according to the recorded request status information.

[0053] Among them, the request status information can record the processing status of each data query request. For example, the request processing is successful, or the request processing fails due to errors such as abnormal query information or abnormal query statement syntax. The request error rate can be used to represent the proportion of data query requests with request errors among all data query requests triggered by the user within the preset time range, and the number of request errors can be used to represent the number of data query requests with request errors among all data query requests triggered by the user within the preset time range.

[0054] Step S303: When the request error rate is less than the preset error rate threshold or the number of request errors is less than the preset number threshold, statistically calculate the cluster parallelism corresponding to the current cluster identifier and the personal parallelism corresponding to the current user identifier.

[0055] Among them, if the request error rate is less than the preset error rate threshold, or the number of request errors is less than the preset number threshold, it can be considered that the possibility of this data query request being an invalid request is relatively small, and the parallelism verification can be continued to determine whether to process this data query request. If the request error rate is greater than or equal to the preset error rate threshold, and the number of request errors is greater than or equal to the preset number threshold, it can be considered that the possibility of this data query request being an invalid request is relatively large, and it can be selected to continue to determine whether this data query request is a duplicate, so as to determine whether to continue to process this data query request.

[0056] Step S304, execute the data query request when the cluster parallelism is less than the first cluster parallelism threshold and the personal parallelism is less than the personal parallelism threshold.

[0057] As described above, by adding the statistics of the request error rate and the number of request errors and threshold judgment, invalid requests can be effectively identified, which is beneficial to excluding third-party malicious attacks or incorrect operations, reducing the waste of system resources caused by invalid requests, and thus improving the security and defense capabilities of the system.

[0058] Figure 4 It is a flowchart of a data query request processing method provided by an embodiment of the present application, which includes a process of determining whether a data query request is a duplicate invalid request. As Figure 4 shown, it includes the following steps:

[0059] Step S401, obtain the data query request to be processed, match the corresponding priority according to the metadata information of the data query request, and query the per-second query rate threshold and parallelism threshold corresponding to the priority. Among them, the metadata information includes the user identifier and the cluster identifier, the per-second query rate threshold includes the personal per-second query rate threshold and the cluster per-second query rate threshold, and the parallelism threshold includes the personal parallelism threshold and the first cluster parallelism threshold.

[0060] Step S402, count the first per-second query rate corresponding to the user identifier currently and the second per-second query rate corresponding to the cluster identifier currently. When the first per-second query rate is less than the personal per-second query rate threshold and the second per-second query rate is less than the cluster per-second query rate threshold, count the request error rate and the number of request errors corresponding to the user identifier within the preset time range according to the recorded request status information.

[0061] Step S403, when the request error rate is less than the preset error rate threshold, or the number of request errors is less than the preset number threshold, count the cluster parallelism corresponding to the cluster identifier currently and the personal parallelism corresponding to the user identifier currently.

[0062] Step S404: When the request error rate is greater than or equal to the preset error rate threshold and the request error count is greater than or equal to the preset count threshold, determine whether the data query request is a duplicate request.

[0063] Among them, if the request error rate is greater than or equal to the preset error rate threshold and the request error count is greater than or equal to the preset count threshold, it can be considered that the possibility of this data query request being an invalid request is relatively high. By determining whether the data query request is a duplicate request, it can be determined whether the invalid request with the same request content is received repeatedly. If the data query request is a duplicate request, the data query request can be rejected. If the data query request is not a duplicate request, the parallelism verification can be continued to determine whether to continue processing the data query request.

[0064] Step S405: When the data query request is not a duplicate request, count the cluster parallelism corresponding to the current cluster identifier and the personal parallelism corresponding to the current user identifier.

[0065] Step S406: When the cluster parallelism is less than the first cluster parallelism threshold and the personal parallelism is less than the personal parallelism threshold, execute the data query request.

[0066] As described above, by combining the determination of whether the data query request is a duplicate request, it can be further determined whether the data query request is a duplicate invalid request, which is beneficial to identifying repeated malicious attacks or incorrect operations and ensuring the information processing security of the system.

[0067] Figure 5 The figure is a flowchart of a data query request processing method provided by an embodiment of the present application, which includes a process of delaying the processing of data query requests. As Figure 5 shown, it includes the following steps:

[0068] Step S501: Obtain the data query request to be processed, match the corresponding priority according to the metadata information of the data query request, and query the query rate threshold per second and the parallelism threshold corresponding to the priority. Among them, the metadata information includes the user identifier and the cluster identifier. The query rate threshold per second includes the personal query rate threshold per second and the cluster query rate threshold per second. The parallelism threshold includes the personal parallelism threshold and the first cluster parallelism threshold.

[0069] Step S502: Count the first query rate per second corresponding to the current user identifier and the second query rate per second corresponding to the current cluster identifier.

[0070] Step S503: When the first query rate per second is less than the personal query rate threshold per second and the second query rate per second is less than the cluster query rate threshold per second, count the cluster parallelism corresponding to the current cluster identifier and the personal parallelism corresponding to the current user identifier.

[0071] Step S504: Execute the data query request when the cluster parallelism is less than the first cluster parallelism threshold and the personal parallelism is less than the personal parallelism threshold.

[0072] Step S505: When the first queries per second rate is greater than or equal to the personal queries per second rate threshold, or the second queries per second rate is greater than or equal to the cluster queries per second rate threshold, pause for a preset time interval and then re - perform the queries per second rate verification for the data query request a preset number of times.

[0073] Among them, if the first queries per second rate is greater than or equal to the personal queries per second rate threshold, or the second queries per second rate is greater than or equal to the cluster queries per second rate threshold, it can be considered that the query traffic of the current user or the target server cluster exceeds the expected limit, and the system may be in a traffic overload state. The processing of this data query request can be postponed. The preset time interval can be 5s, 10s, etc. By pausing for the preset time interval, the processing of this data query request can be delayed, and the queries per second rate verification is re - performed a preset number of times. The preset number of times can be 1 time, 2 times, etc., to relieve the instantaneous traffic overload of the system.

[0074] As described above, by delaying the processing of data query requests that exceed the queries per second rate limit, it is possible to effectively avoid instantaneous traffic overload, smooth traffic peaks, and further reduce the situation where user requests cannot be executed due to system instantaneous overload, improving the user experience.

[0075] Figure 6 This is a flowchart of a data query request processing method provided by an embodiment of the present application, which includes a process of rejecting a data query request. As Figure 6 shown, it includes the following steps:

[0076] Step S601: Obtain the data query request to be processed, match the corresponding priority according to the metadata information of the data query request, and query the queries per second rate threshold and parallelism threshold corresponding to the priority. Among them, the metadata information includes the user identifier and the cluster identifier, the queries per second rate threshold includes the personal queries per second rate threshold and the cluster queries per second rate threshold, and the parallelism threshold includes the personal parallelism threshold and the first cluster parallelism threshold.

[0077] Step S602: Statistically calculate the first queries per second rate corresponding to the current user identifier and the second queries per second rate corresponding to the current cluster identifier. When the first queries per second rate is less than the personal queries per second rate threshold and the second queries per second rate is less than the cluster queries per second rate threshold, statistically calculate the cluster parallelism corresponding to the current cluster identifier and the personal parallelism corresponding to the current user identifier.

[0078] Step S603: Execute the data query request when the cluster parallelism is less than the first cluster parallelism threshold and the individual parallelism is less than the individual parallelism threshold.

[0079] Step S604: Reject the data query request when the cluster parallelism is greater than or equal to the first cluster parallelism threshold or the individual parallelism is greater than or equal to the individual parallelism threshold.

[0080] As described above, when the cluster parallelism is greater than or equal to the first cluster parallelism threshold or the individual parallelism is greater than or equal to the individual parallelism threshold, by rejecting the data query request, the parallelism in each dimension can be effectively controlled, preventing the system from experiencing excessive concurrency and reasonably coordinating system resources.

[0081] Figure 7 The block diagram of a data query request processing device provided by an embodiment of the present application. The device is configured to execute the data query request processing method provided in the above embodiment, and has functional modules and beneficial effects corresponding to the execution of the method. As Figure 7 shown, the device includes:

[0082] The first threshold determination module 101 is configured to obtain the data query request to be processed, match the corresponding priority according to the metadata information of the data query request, and query the per-second query rate threshold and parallelism threshold corresponding to the priority. Among them, the metadata information includes the user identifier and the cluster identifier. The per-second query rate threshold includes the individual per-second query rate threshold and the cluster per-second query rate threshold. The parallelism threshold includes the individual parallelism threshold and the first cluster parallelism threshold;

[0083] The per-second query rate verification module 102 is configured to count the first per-second query rate corresponding to the current user identifier and the second per-second query rate corresponding to the current cluster identifier. When the first per-second query rate is less than the individual per-second query rate threshold and the second per-second query rate is less than the cluster per-second query rate threshold, count the cluster parallelism corresponding to the current cluster identifier and the individual parallelism corresponding to the current user identifier;

[0084] The first parallelism verification module 103 is configured to execute the data query request when the cluster parallelism is less than the first cluster parallelism threshold and the individual parallelism is less than the individual parallelism threshold.

[0085] As described above, by obtaining a data query request to be processed, matching a corresponding priority according to the metadata information of the data query request, and querying the query rate threshold per second and the parallelism threshold corresponding to the priority, where the metadata information includes a user identifier and a cluster identifier, the query rate threshold per second includes a personal query rate threshold per second and a cluster query rate threshold per second, and the parallelism threshold includes a personal parallelism threshold and a first cluster parallelism threshold; statistically calculating the first query rate per second corresponding to the user identifier currently and the second query rate per second corresponding to the cluster identifier currently, and when the first query rate per second is less than the personal query rate threshold per second and the second query rate per second is less than the cluster query rate threshold per second, statistically calculating the cluster parallelism corresponding to the cluster identifier currently and the personal parallelism corresponding to the user identifier currently; when the cluster parallelism is less than the first cluster parallelism threshold and the personal parallelism is less than the personal parallelism threshold, executing the data query request. In the above solution, by matching a corresponding priority according to the metadata information of the data query request, the data query requests to be processed can be pre-divided according to different priorities, and the query rate threshold per second and the parallelism threshold corresponding to the priority are queried to adapt to the traffic control of data query requests with different priorities, ensuring the priority processing of high-priority data query requests. By respectively verifying the first query rate per second corresponding to the user identifier and the second query rate per second corresponding to the cluster identifier, the allowable amount of requests can be flexibly adjusted according to the actual request situations of the cluster and the user, effectively preventing the overload of system resources caused by traffic peaks. By respectively verifying the personal parallelism corresponding to the user identifier and the cluster parallelism corresponding to the cluster identifier, parallelism control in different dimensions is performed, dynamically adjusting the number of parallel tasks, effectively controlling the concurrent traffic of the database system, which is beneficial to improving the service quality and user experience.

[0086] In a possible embodiment, it further includes a second threshold determination module, configured to

[0087] query the second cluster parallelism threshold and the preset ratio threshold corresponding to the priority;

[0088] a proportion calculation module, configured to:

[0089] when the cluster parallelism is less than the second cluster parallelism threshold, calculate the first proportion of the number of data query requests whose execution time is less than the reference execution time within a preset time range relative to the total number of requests;

[0090] Correspondingly, the query rate verification module 102 is further configured to:

[0091] when the first proportion is greater than or equal to the preset ratio threshold, statistically calculate the personal parallelism corresponding to the user identifier currently.

[0092] In a possible embodiment, it further includes a request error statistics module, configured to:

[0093] Statistically calculate the request error rate and the number of request errors corresponding to the user identifier within a preset time range according to the recorded request status information;

[0094] Correspondingly, the per-second query rate verification module 102 is further configured to:

[0095] When the request error rate is less than the preset error rate threshold or the number of request errors is less than the preset number threshold, statistically calculate the current cluster parallelism corresponding to the cluster identifier and the current personal parallelism corresponding to the user identifier.

[0096] In a possible embodiment, it further includes a request duplication determination module, configured to:

[0097] When the request error rate is greater than or equal to the preset error rate threshold and the number of request errors is greater than or equal to the preset number threshold, determine whether the data query request is a duplicate request;

[0098] When the data query request is not a duplicate request, statistically calculate the current cluster parallelism corresponding to the cluster identifier and the current personal parallelism corresponding to the user identifier.

[0099] In a possible embodiment, it further includes a request latency processing module, configured to:

[0100] When the first per-second query rate is greater than or equal to the personal per-second query rate threshold or the second per-second query rate is greater than or equal to the cluster per-second query rate threshold, pause for a preset time interval and then re-perform the per-second query rate verification on the data query request a preset number of times.

[0101] In a possible embodiment, it further includes a second parallelism verification module, configured to:

[0102] When the cluster parallelism is greater than or equal to the first cluster parallelism threshold or the personal parallelism is greater than or equal to the personal parallelism threshold, reject the data query request.

[0103] Figure 8 The structural schematic diagram of a data query request processing device provided by an embodiment of the present application is as Figure 8 shown. The device includes a processor 201, a memory 202, an input device 203, and an output device 204; the number of processors 201 in the device can be one or more, Figure 8 taking one processor 201 as an example; the processor 201, the memory 202, the input device 203, and the output device 204 in the device can be connected through a bus or other means, Figure 8Take the bus connection as an example. The memory 202, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the data query request processing method in the embodiments of the present application. The processor 201 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 202, that is, implements the above-mentioned data query request processing method. The input device 203 can be configured to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the device. The output device 204 may include display devices such as a display screen.

[0104] The embodiments of the present application also provide a non-volatile storage medium containing computer-executable instructions. The computer-executable instructions are configured to execute a data query request processing method described in the above embodiments when executed by a computer processor. The method includes: obtaining a data query request to be processed, matching a corresponding priority according to the metadata information of the data query request, and querying the query rate threshold per second and parallelism threshold corresponding to the priority. The metadata information includes a user identifier and a cluster identifier. The query rate threshold per second includes a personal query rate threshold per second and a cluster query rate threshold per second. The parallelism threshold includes a personal parallelism threshold and a first cluster parallelism threshold; counting the first query rate per second corresponding to the user identifier currently and the second query rate per second corresponding to the cluster identifier currently. When the first query rate per second is less than the personal query rate threshold per second and the second query rate per second is less than the cluster query rate threshold per second, counting the cluster parallelism corresponding to the cluster identifier currently and the personal parallelism corresponding to the user identifier currently; when the cluster parallelism is less than the first cluster parallelism threshold and the personal parallelism is less than the personal parallelism threshold, execute the data query request.

[0105] It should be noted that in the embodiments of the above data query request processing device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and are not configured to limit the protection scope of the embodiments of the present application.

[0106] In some possible implementation manners, various aspects of the method provided by the present application can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is configured to enable the computer device to execute the steps in the methods according to various exemplary embodiments of the present application described above in this specification. For example, the computer device can execute the data query request processing method recorded in the embodiments of the present application. The program product can be implemented by any combination of one or more readable media.

Claims

1. A method for processing a data query request, characterized in that: include: Obtain a data query request to be processed, match a corresponding priority according to metadata information of the data query request, and query a query rate threshold per second and a parallelism threshold corresponding to the priority, wherein the metadata information includes a user identifier and a cluster identifier, the query rate threshold per second includes a personal query rate threshold per second and a cluster query rate threshold per second, and the parallelism threshold includes a personal parallelism threshold and a first cluster parallelism threshold; Counting a first query rate per second currently corresponding to the user identifier and a second query rate per second currently corresponding to the cluster identifier, and when the first query rate per second is less than the personal query rate per second threshold, and the second query rate per second is less than the cluster query rate per second threshold, counting the cluster parallelism currently corresponding to the cluster identifier and the personal parallelism currently corresponding to the user identifier; When the cluster parallelism is less than the first cluster parallelism threshold and the individual parallelism is less than the individual parallelism threshold, the data query request is executed.

2. The data query request processing method according to claim 1, characterized in that: After counting the cluster parallelism currently corresponding to the cluster identifier, the method further includes: Querying a second cluster parallelism threshold and a preset ratio threshold corresponding to the priority; When the cluster parallelism is less than the second cluster parallelism threshold, calculating a first ratio of the number of data query requests whose execution time is less than the reference execution time within a preset time range to the total number of requests; Correspondingly, the counting of the personal parallelism currently corresponding to the user identifier includes: When the first proportion is greater than or equal to the preset proportion threshold, the personal parallelism currently corresponding to the user identification is counted.

3. The data query request processing method according to claim 1, characterized in that: Before counting the cluster parallelism currently corresponding to the cluster identifier and the personal parallelism currently corresponding to the user identifier, the method further includes: Counting the request error rate and the number of request errors corresponding to the user identifier within a preset time range according to the recorded request status information; Accordingly, counting the cluster parallelism currently corresponding to the cluster identifier and the personal parallelism currently corresponding to the user identifier includes: When the request error rate is less than a preset error rate threshold, or the number of request errors is less than a preset number threshold, the cluster parallelism currently corresponding to the cluster identifier and the personal parallelism currently corresponding to the user identifier are counted.

4. The data query request processing method according to claim 3, characterized in that: Also includes: When the request error rate is greater than or equal to a preset error rate threshold, and the number of request errors is greater than or equal to a preset number threshold, determining whether the data query request is a repeated request; When the data query request is not a repeated request, the cluster parallelism currently corresponding to the cluster identifier and the personal parallelism currently corresponding to the user identifier are counted.

5. The data query request processing method according to claim 1, characterized in that: Also includes: When the first query rate per second is greater than or equal to the personal query rate per second threshold, or the second query rate per second is greater than or equal to the cluster query rate per second threshold, the query rate per second of the data query request is rechecked a preset number of times after pausing for a preset time interval.

6. The data query request processing method according to claim 1, characterized in that: Also includes: When the cluster parallelism is greater than or equal to the first cluster parallelism threshold, or the individual parallelism is greater than or equal to the individual parallelism threshold, the data query request is rejected.

7. A data query request processing device, characterized in that: include: A first threshold determination module is configured to obtain a data query request to be processed, match a corresponding priority according to metadata information of the data query request, and query a query rate threshold per second and a parallelism threshold corresponding to the priority, wherein the metadata information includes a user identifier and a cluster identifier, the query rate threshold per second includes a personal query rate threshold per second and a cluster query rate threshold per second, and the parallelism threshold includes a personal parallelism threshold and a first cluster parallelism threshold; a query rate per second verification module, configured to count a first query rate per second currently corresponding to the user identifier and a second query rate per second currently corresponding to the cluster identifier, and when the first query rate per second is less than the personal query rate per second threshold, and the second query rate per second is less than the cluster query rate per second threshold, count the cluster parallelism currently corresponding to the cluster identifier and the personal parallelism currently corresponding to the user identifier; The first parallelism checking module is configured to execute the data query request when the cluster parallelism is less than the first cluster parallelism threshold and the individual parallelism is less than the individual parallelism threshold.

8. A data query request processing device, the device comprising: one or more processors; A storage device is configured to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the data query request processing method described in any one of claims 1 to 6.

9. A non-volatile storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the data query request processing method according to any one of claims 1 to 6 when executed by a computer processor.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the data query request processing method described in any one of claims 1 to 6 is implemented.