Request processing method and device, electronic equipment and storage medium
By deploying data collection agents and machine learning models on API gateway nodes, real-time collection and prediction of load data, and dynamic adjustment of request priorities, the problem that traditional API gateways cannot adapt to traffic changes in real time is solved, priority processing and resource optimization of key requests are achieved, and system performance and service quality are improved.
Patent Information
- Application Number
- CN202510895996.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-12
AI Technical Summary
Traditional API gateways lack dynamic adjustment mechanisms and are unable to perceive changes in network and service status in real time, resulting in delayed processing of important requests, affecting user experience and service quality. They also lack automated support, increasing the complexity and cost of system maintenance.
By deploying data collection agents on API gateway nodes, business request data and load indicator data are collected in real time, and the request priority is dynamically adjusted in combination with machine learning models to predict future system loads and optimize resource allocation.
Ensure that critical requests are prioritized, optimize resource allocation, avoid resource waste and overload, and improve overall system performance and service quality.
Smart Images

Figure CN120639864A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of request processing, and in particular to a request processing method, device, electronic device and storage medium. Background Art
[0002] In modern internet applications, APIs (Application Programming Interfaces) serve as a crucial means of communication and service integration between applications, and are widely used in various microservice architectures and cross-platform integration. As the intermediary between service consumption and provision, the performance and reliability of API gateways play a crucial role in the overall system's service quality.
[0003] Commonly used API gateways often use fixed priority allocation strategies to handle different requests. Fixed priority strategies cannot dynamically adapt to the needs of different application scenarios. For example, during peak hours, certain critical business requests need to be processed first, while during low traffic periods, all requests need to be processed evenly. Due to the lack of a dynamic adjustment mechanism, when the service load or environment changes, fixed priorities can easily cause important requests to be delayed, affecting user experience and service quality. Traditional API gateways lack the ability to monitor real-time traffic and service quality, and are unable to instantly adjust priorities based on current traffic and request characteristics. Since they cannot perceive changes in network and service status in real time, the system struggles to achieve efficient resource scheduling and traffic management, and is insufficiently prepared for sudden traffic and abnormal situations. Although these problems can be alleviated to a certain extent by manually adjusting configurations or adopting complex strategies, these methods often require a lot of human intervention and lack automated support, increasing the complexity and cost of system maintenance. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a request processing method, apparatus, electronic device, and storage medium to prioritize critical requests, optimize resource allocation, avoid resource waste and overload, and improve overall system performance and service quality. The specific technical solutions are as follows:
[0005] In a first aspect of the implementation of the present application, a request processing method is first provided, comprising:
[0006] Use the data collection agent deployed on the API gateway node to collect business request data and load indicator data;
[0007] Classifying and sorting the service request data according to the initial priority rule to obtain classified and sorted request data;
[0008] Determining a scheduling strategy for the classification and sorting request data at a current time based on the load indicator data, so as to schedule the classification and sorting request data according to the scheduling strategy;
[0009] Inputting API request data and system load data of the API gateway node within a first preset time period from the current time into a system load prediction model to obtain a predicted system load of the API gateway node at a future target time;
[0010] Based on the predicted system load, the scheduling strategy of the classified sorting request data at the target time is adjusted to obtain the target scheduling strategy of the target time, and when the target time is reached, the classified sorting request data at the target time is requested to be scheduled based on the target scheduling strategy.
[0011] In a second aspect of the present application, a request processing device is provided, comprising:
[0012] The data acquisition module is used to collect business request data and load indicator data using the data collection agent deployed on the API gateway node;
[0013] A request data acquisition module is used to classify and sort the service request data according to the initial priority rule to obtain classified and sorted request data;
[0014] a scheduling strategy determining module, configured to determine a scheduling strategy for the classified and sorted request data according to the load indicator data, so as to perform request scheduling on the classified and sorted request data according to the scheduling strategy;
[0015] A predicted load acquisition module, configured to input API request data and system load data of the API gateway node within a first preset time period from the current time into a system load prediction model to obtain a predicted system load of the API gateway node at a future target time;
[0016] A scheduling strategy adjustment module is used to adjust the scheduling strategy of the classified sorting request data at the target time based on the predicted system load, obtain the target scheduling strategy of the target time, and, when the target time is reached, request scheduling of the classified sorting request data at the target time based on the target scheduling strategy.
[0017] In another aspect of the present application, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;
[0018] Memory for storing computer programs;
[0019] The processor is configured to implement any of the above-mentioned request processing methods when executing a program stored in the memory.
[0020] In another aspect of the present application, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, and when the computer-readable storage medium is run on the computer, the computer executes any of the request processing methods described above.
[0021] In another aspect of the implementation of the present application, a computer program product comprising instructions is provided, on which a computer program is stored. When the computer program is run on a computer, the computer is enabled to execute any of the request processing methods described above.
[0022] The solution provided in the embodiment of the present application collects request data of API gateway nodes in real time through a data collection agent program, and dynamically adjusts the processing priority of API requests in combination with a machine learning model, which can ensure that key requests are processed first, optimize resource allocation, avoid resource waste and overload, and improve overall system performance and service quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art.
[0024] Figure 1 A flowchart of a request processing method provided in an embodiment of the present application;
[0025] Figure 2 A schematic diagram of a data collection and storage process provided in an embodiment of the present application;
[0026] Figure 3 A flowchart of a request classification and sorting method provided in an embodiment of the present application;
[0027] Figure 4 A schematic diagram of a request classification architecture provided in an embodiment of the present application;
[0028] Figure 5 A flowchart of the steps of a scheduling strategy adjustment method provided in an embodiment of the present application;
[0029] Figure 6 A flow chart of the steps of a load prediction model training method provided in an embodiment of the present application;
[0030] Figure 7 A flowchart of the steps of a policy resource adjustment method provided in an embodiment of the present application;
[0031] Figure 8A schematic diagram of a model training and strategy adjustment process provided in an embodiment of the present application;
[0032] Figure 9 A flowchart of the steps of a scheduling strategy testing method provided in an embodiment of the present application;
[0033] Figure 10 A schematic diagram of a test process provided in an embodiment of the present application;
[0034] Figure 11 A schematic diagram of a strategy optimization process provided in an embodiment of the present application;
[0035] Figure 12 A flowchart of another method for adjusting a scheduling strategy provided in an embodiment of the present application;
[0036] Figure 13 A schematic diagram of the structure of a request processing device provided in an embodiment of the present application;
[0037] Figure 14 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0038] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0039] Figure 1 A flowchart of a request processing method provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the request processing method may include: step 101, step 102, step 103, step 104 and step 105.
[0040] Step 101: Use the data collection agent deployed on the API gateway node to collect business request data and load indicator data.
[0041] In this embodiment, the API gateway node is a server that acts as an API front end and is located between the backend service and the client. It is responsible for receiving API requests, executing throttling and security policies, passing requests to the backend service, and returning appropriate results to the client.
[0042] A data collection agent is deployed on the API gateway node. This agent is software or script deployed on the API gateway node to collect business request data and load metric data passing through the gateway. This agent can monitor and capture network traffic flowing through the API gateway and extract valuable information from it, such as request content, request frequency, and response time.
[0043] When dynamically adjusting the request priority of an API gateway node, a data collection agent deployed on the API gateway node can be used to collect business request data and load metric data in real time. Business request data can include: request content (such as HTTP (HyperText Transfer Protocol) request header, request method, request parameters, and request body), request source (such as the IP address, MAC address, and user agent of the client or device initiating the request), request time (such as the timestamp when the request arrives at the API gateway, which can be used to analyze the real-time nature of the request), and request result (such as the HTTP response status code, response content, and response time). Load metric data can include: throughput (the number of requests the API gateway can process per second, reflecting the gateway's processing capacity), response time (the time from when a request arrives at the API gateway to when a response is returned to the client, used to evaluate the service's responsiveness), concurrency (the number of requests processed simultaneously, reflecting the gateway's performance under high concurrency), error rate (the proportion of request failures, used to evaluate the stability and reliability of the service), and resource utilization (such as the usage of resources such as CPU, memory, and disk, used to evaluate whether the gateway's hardware resources are fully utilized).
[0044] After collecting service request data and load metric data, they can be stored in a time series database (such as Prometheus) to support rapid query and analysis. Furthermore, the alarm system can pre-set alarm rules. When monitoring data exceeds a threshold, an alarm is triggered to notify operations and maintenance personnel. For example, in the case of the bullet message service, the error rate of these requests determines whether the service is normal. If the error rate is particularly high, the priority of the bullet message service is lowered, and high-priority services are promoted.
[0045] The data collection and storage process can be as follows Figure 2 As shown in the figure, the data collection agent can collect business request data and load index data flowing through the API gateway node every N seconds (N is a positive integer) and send it to the data processor. The data processor can then process the collected business request data and load index data, such as cleaning, conversion, and aggregation. The processed business request data and load index data are then sent to the time series database via the HTTP interface.
[0046] Of course, when storing data, the data format can be converted into JSON format and stored in the form of key-value pairs (such as using timestamp as key and data content and load indicator content as value, etc.) to facilitate subsequent query and analysis.
[0047] The embodiment of the present application deploys a lightweight data collection agent program on the API gateway node, which can collect API request data and load indicator data in real time, centrally collect and analyze this data, form a data analysis system with a global perspective, and support rapid response and decision-making.
[0048] Step 102: Classify and sort the service request data according to the initial priority rule to obtain classified and sorted request data.
[0049] The initial priority rule may be a preset fixed priority rule, which defines rules such as the priority level and scheduling order of the service request.
[0050] After obtaining the business request data collected by the data collection agent pre-deployed on the API gateway node, the business request data can be classified and sorted according to the pre-set initial priority rules to obtain classified and sorted request data.
[0051] After obtaining the sorted request data, the sorted request data can be stored in the corresponding priority queue. For example, preliminary priority rules can be set to initialize the request priority based on the request type and user role. For example, payment requests and VIP user operations are set to high priority and can be stored in the high priority queue. Query requests and ordinary user operations are set to medium priority and can be stored in the medium priority queue. Data synchronization requests are set to low priority and can be stored in the low priority queue, etc.
[0052] It can be understood that the above examples are merely examples listed for a better understanding of the technical solutions of the embodiments of the present application, and are not intended to be the sole limitation on the embodiments.
[0053] The processing flow for the allocation and sorting of service request data will be combined in the following embodiments. Figure 3 The present embodiment will be described in detail and will not be described in detail here.
[0054] Step 103: Determine a scheduling strategy for the classified and sorted request data according to the load indicator data, and perform request scheduling on the classified and sorted request data according to the scheduling strategy.
[0055] A scheduling policy is a set of decision-making rules and execution strategies developed to achieve efficient request scheduling. Its core purpose is to optimize resource allocation and request processing by dynamically adapting request characteristics and system status. In this embodiment, the scheduling policy can include scheduling priorities for categorized and sorted request data (i.e., higher-priority data is scheduled first) and scheduling resource information (such as GPU computing resources, CPU resources, etc.).
[0056] Request scheduling is the process of assigning and routing classified and sorted request data to appropriate processing resources during business request processing. Its core goal is to ensure that requests are processed efficiently and in an orderly manner, while optimizing system resource utilization, ensuring service quality, and meeting business needs.
[0057] After the business request data is classified and sorted according to the initial priority rule to obtain the classified and sorted request data, the scheduling strategy of the obtained classified and sorted request data can be determined according to the load index data, and the classified and sorted request data can be scheduled according to the determined scheduling strategy. Specifically, the current system load of the API gateway node can be determined according to the load index data, and the scheduling strategy of the classified and sorted request data at the current time can be determined according to the current system load. The implementation process will be combined with the following embodiments. Figure 5 Of course, the system load of the API gateway node can also be determined based on the load index data, and the optimal scheduling strategy for the test can be determined as the scheduling strategy for the classified sorting request data at the current time. Figure 12 The present embodiment will be described in detail and will not be described in detail here.
[0058] Step 104: Input the API request data and system load data of the API gateway node within a first preset time period from the current time into a system load prediction model to obtain the predicted system load of the API gateway node at a future target time.
[0059] The system load prediction model refers to a pre-trained model used to predict the system load of the API gateway node at a certain moment in the future. The training process of the system load prediction model will be combined with the following embodiments. Figure 6 The present embodiment will be described in detail and will not be described in detail here.
[0060] The first preset duration is a preset duration used to obtain API request data and load data for future system load prediction. In this example, the first preset duration can be 10 minutes, 20 minutes, or the like. The specific value of the first preset duration can be determined based on service requirements and is not limited in this embodiment.
[0061] When it is necessary to predict the system load of the API gateway node at a future target time, the API request data and system load data of the API gateway node within a first preset time period from the current time can be obtained. For example, if the first preset time period is 10 minutes and the current time is 12:00 on December 30, 2024, the API request data and system load data of the API gateway node from 11:50 on December 30, 2024 to 12:00 on December 30, 2024 can be obtained to serve as input to the system load prediction model to predict the future system load.
[0062] After obtaining API request data and system load data from the API gateway node within a first preset time period from the current time, the API request data and system load data within the first preset time period can be input into a system load prediction model. The system load prediction model processes the API request data and system load data to predict the predicted system load of the API gateway node at a future target time. Specifically, the prediction process can include: 1. Data preprocessing: removing duplicate, invalid, or abnormal data and converting data of different magnitudes to the same magnitude to facilitate model processing. Then, the features most relevant to the predicted system load can be selected and the data converted into a time series format. 2. The data converted into the time series format is input into the system load prediction model, and a future target time is specified, so that the model outputs the predicted system load at that time.
[0063] Step 105: Based on the predicted system load, adjust the scheduling strategy of the classification and sorting request data at the target time to obtain the target scheduling strategy at the target time, and when the target time is reached, perform request scheduling on the classification and sorting request data at the target time based on the target scheduling strategy.
[0064] After the predicted system load of the API gateway node at the target time in the future is obtained, the scheduling policy of the classified and sorted request data at the target time can be adjusted based on the predicted system load to obtain the target scheduling policy at the target time. Thus, when the target time is reached, the classified and sorted request data at the target time can be scheduled based on the target scheduling policy. Specifically, the scheduling priority can be adjusted based on the predicted system load and the pre-set load threshold, and the scheduling resource information can be pre-allocated. This implementation process will be described in the following embodiments in conjunction with Figure 7 The present embodiment will be described in detail and will not be described in detail here.
[0065] The embodiment of the present application uses a data collection agent to collect request data from the API gateway node in real time, and dynamically adjusts the processing priority of the API request in combination with a machine learning model, which can ensure that key requests are processed first, optimize resource allocation, avoid resource waste and overload, and improve overall system performance and service quality.
[0066] Next, combine Figure 3 The allocation, sorting and storage process of business request data is described in detail.
[0067] Reference Figure 3 , shows a flowchart of the steps of a request classification and sorting method provided by an embodiment of the present application. Figure 3 As shown, the request classification and sorting method may include: step 301, step 302 and step 303.
[0068] Step 301: Determine the request priority corresponding to each of the service request data according to the initial priority rule, the service type of each of the service request data, and the client identifier corresponding to each of the service request data.
[0069] In this embodiment, the service type refers to the type corresponding to the service request data. In this example, the service type may be, but is not limited to, a payment type, a query type, a synchronization type, and the like.
[0070] The client ID refers to the unique ID of the client that sends the service request data.
[0071] After obtaining the business request data of the API gateway node, the request priority corresponding to each business request data can be determined based on the initial priority rules, the business type of each business request data and the client identifier corresponding to each business request data. Among them, the initial priority rules are a set of criteria for determining the order of processing tasks, events or projects, which may include various factors such as time sensitivity, resource dependence, urgency, etc. In a specific implementation, preliminary priority rules can be set in advance, and the priority of the request can be initialized according to the request type and user role (i.e., client identifier). For example, payment requests and VIP user operations are set to high priority, query requests and ordinary user operations are set to medium priority, and data synchronization requests are set to low priority. Or, taking a multimedia player as an example, the priority of the video playback service is higher than the priority of the barrage playback service, and the priority of the barrage playback service is higher than the priority of the comment service, etc.
[0072] It can be understood that the above examples are merely examples listed for a better understanding of the technical solutions of the embodiments of the present application, and are not intended to be the sole limitation on the embodiments.
[0073] Step 302: Sort the service request data according to the request priorities to obtain sorted request data.
[0074] After determining the request priority corresponding to each service request data according to the initial priority rule, the service type of each service request data, and the client identifier corresponding to each service request data, the service request data can be sorted according to the request priority to obtain sorted request data. Specifically, the service request data can be sorted in descending order of request priority.
[0075] Step 303: Classify the sorting request data according to the priority threshold corresponding to at least one preset priority queue to obtain the classified sorting request data.
[0076] The priority queue may be a max-heap data structure for quickly determining the highest priority request.
[0077] After sorting each business request data according to each request priority and obtaining the sorted request data, the sorted request data can be allocated according to the priority threshold corresponding to at least one pre-set priority queue to obtain classified sorted request data. At the same time, the classified sorted request data can be stored in the corresponding priority queue. Specifically, multiple priority queues can be created first, such as a high priority queue, a medium priority queue, a low priority queue, etc., and a priority threshold corresponding to each priority queue can be set. The request data list is traversed, and each request is compared with the threshold according to its priority, and is allocated to the corresponding priority queue to implement the classification, sorting and storage process of the business request data.
[0078] Of course, in actual applications, a priority queue can be pre-set, and the service request data can be stored in the priority queue in descending order. Figure 4 As shown, first, request classification can be performed, that is, preliminary priority rules are set, and the priority of the request is initialized according to the request type and user role, so as to divide the business request data into three types of request data: high priority, medium priority, and low priority, and store the business request data in the priority queue in order from high to low priority. Through this priority queue, the highest priority business request data and the second priority business request data can be clearly identified, which is convenient for request scheduling. Specifically, high-priority request data is scheduled first, medium-priority request data is scheduled after the high-priority request data is scheduled, and low-priority request data is scheduled last.
[0079] This embodiment of the application combines initial priority rules, the service type of the service request data, the client identifier corresponding to the service request data, and pre-set priority queues and priority thresholds for request storage. This processing flow can efficiently determine, sort, and assign the priority of service request data. This not only improves the efficiency and response speed of system request processing, but also ensures that critical service requests are prioritized, thereby improving the overall service processing quality and user experience.
[0080] Next, combine Figure 5 The process of adjusting the scheduling strategy based on load indicator data is described in detail.
[0081] Reference Figure 5 , shows a flowchart of the steps of a scheduling strategy adjustment method provided by an embodiment of the present application. Figure 5 As shown, the scheduling strategy adjustment method may include: step 501 and step 502.
[0082] Step 501: Determine the current system load of the API gateway node according to the load indicator data.
[0083] In this embodiment, after obtaining load indicator data for an API gateway node, the current system load of the API gateway node can be determined based on the load indicator data. The load indicator data may include indicators such as CPU utilization, memory utilization, disk I / O utilization, and network I / O utilization. When determining the system load, each load indicator can be individually evaluated to determine whether it is in a normal, warning, or critical state, such as by setting a threshold. The evaluation results of all indicators are summarized to comprehensively consider the overall system load. Specifically, comprehensive evaluation rules can be set: for example, if multiple key indicators (such as CPU utilization and memory utilization) are simultaneously in a warning or critical state, the overall system load is considered to be relatively severe. Furthermore, based on the comprehensive evaluation results, the system load can be categorized into three levels: low, medium, and high. Low load: All indicators are within the normal range; medium load: Some indicators are in a warning state, but the system still operates normally; and high load: Multiple key indicators are in a critical state, impacting system performance.
[0084] Step 502: Determine a scheduling strategy for the classified and sorted request data according to the current system load.
[0085] After determining the current system load of the API gateway node based on the load indicator data, a scheduling policy for the classified sorting request data can be determined based on the current system load. In this embodiment, the priority queues may include: a first priority queue, a second priority queue, and the third priority queue. The priority of the classified sorting request data in the first priority queue is higher than the priority of the classified sorting request data in the second priority queue, and the priority of the classified sorting request data in the second priority queue is higher than the priority of the classified sorting request data in the third priority queue. That is, the first priority queue is a high priority queue, the second priority queue is a medium priority queue, and the third priority queue is a low priority queue. It is understandable that within the same priority queue, the priorities of multiple service request data may be different. Specifically, this can be determined based on actual conditions, and this embodiment does not impose any restrictions on this.
[0086] The method for determining the scheduling strategy can be described in combination with the following two methods.
[0087] 1. When the current system load is greater than or equal to the first load threshold, the scheduling strategy for the classified request data in the priority queue is determined to be the first scheduling strategy. The first scheduling strategy is: in the first priority queue, the priority of a part of the classified sorting request data is increased; in the second priority queue, the priority of a part of the classified sorting request data is lowered; and in the third priority queue, a part of the classified sorting request data is intercepted.
[0088] Specifically, when the system load is high, some medium-priority requests will be downgraded to low priority, while the priority of some high-priority requests will be further increased. Furthermore, some low-priority requests will be blocked to ensure that high-priority requests are processed first. Understandably, requests in the high-priority queue are not of the same type. When the system load is excessively high, priority can be further refined. For example, in financial trading systems, "real-time stop-loss orders" require a higher priority than "normal transaction queries." This secondary increase ensures that the former is executed first during CPU preemption. In e-commerce systems, "payment link requests" (which impact cash flow) have a higher priority than "product browsing requests" (which impact user experience). Under high load, the former can be further increased. Similarly, in the medium-priority queue, "user review storage" in e-commerce systems can be downgraded under high load. Since it does not affect the main transaction process, its processing can be delayed. In the low-priority queue, background tasks such as data backup and log compression are directly blocked during high load to avoid consuming CPU and I / O resources.
[0089] In the above implementation, the "partial quantity" is determined based on the current system load. Proportional control is based on the load threshold. When the system load reaches the first threshold, the "partial quantity" of each queue is adjusted according to the preset ratio. For example, if the high-priority queue has 100 requests, the load factor is 1.2 (the load exceeds the threshold by 20%), and the promotion weight is 0.2, then the promotion quantity = 100 × 1.2 × 0.2 = 24. If the medium-priority queue has 200 requests, the demotion weight is 0.4, and the demotion quantity = 200 × 1.2 × 0.4 = 96, etc.
[0090] Of course, it is also possible to dynamically adjust the "partial quantity" based on dynamic resource feedback, that is, by monitoring the remaining amount of system resources in real time (such as CPU idle rate, available memory, etc.), dynamically calculate the number of requests that can be processed, and perform adjustments on the excess. High priority queue: If the available resources can only handle 70% of the requests, increase the priority of the most critical 30% of requests to ensure that they obtain resources first. Medium priority queue: If the available resources are insufficient, calculate the number that needs to be downgraded = total number of queues - (available resources × average resource consumption per request). Low priority queue: Directly intercept the number of requests that exceed the available resources, etc.
[0091] It can be understood that the above examples are merely examples listed for a better understanding of the technical solutions of the embodiments of the present application, and are not intended to be the sole limitation on the embodiments.
[0092] 2. When the current system load is less than the first load threshold, the scheduling strategy for the classified sorting request data in the priority queue is determined to be the second scheduling strategy. The second scheduling strategy is: the scheduling priority of the classified sorting request data in the priority queue remains unchanged.
[0093] That is, when the system load is not high, the scheduling priority of the classified sorted request data in the priority queue is kept unchanged, and the business request data is processed in the order of high, medium and low priority.
[0094] The embodiment of the present application determines the scheduling strategy of business requests through the load index data of the API gateway node collected in real time, which can ensure the priority processing of key business requests and improve the overall system performance and service quality.
[0095] Next, combine Figure 6 The training process of the system load prediction model is described in detail.
[0096] Reference Figure 6 , shows a flow chart of the steps of a load prediction model training method provided by an embodiment of the present application. Figure 6 As shown, the load prediction model training method may include: step 601, step 602, step 603 and step 604.
[0097] Step 601: Acquire historical request data of the API gateway node within a second preset time period from the current time, and system load data corresponding to the historical request data, where the second preset time period is greater than the first preset time period.
[0098] In this embodiment, the second preset time period is longer than the first preset time period. The second preset time period can be 1 month, 1 year, etc., which is not limited in this embodiment.
[0099] Historical request data may include: request time, processing time, user role, request type, and other data.
[0100] System load data can be a system performance index of the API gateway node reflected by data such as CPU usage, memory usage, and request queue length, such as high load, medium load, and low load.
[0101] When training the system load prediction model, historical request data of the API gateway node within a second preset time period from the current time and the system load data corresponding to the historical request data can be obtained. For example, if the current time is 12:00 on December 31, 2024, and the second preset time period is 1 year, then historical request data of the API gateway node from 12:00 on December 31, 2023 to 12:00 on December 31, 2024, and the corresponding system load data can be obtained from the database.
[0102] Step 602: Extract the request data features of the historical request data and the load data features of the system load data.
[0103] After obtaining the historical request data of the API gateway node within a second preset time period from the current time, and the system load data corresponding to the historical request data, the request data features of the historical request data and the load data features of the system load data can be extracted. The extracted features may include: time features, request features, and system load features, wherein the time features may include: the timestamp of the request (i.e., the specific time of each request) and periodic features (such as the day, week, month, etc. of the request, which help capture seasonal or periodic changes), and the request features may include: request type (such as payment type, query type, etc., which may reflect the impact of different types of requests on the system load) and user role (such as VIP user, ordinary user, etc., which may affect the frequency and priority of the request). The system load feature can be a specific indicator value of the current system load of the API gateway node, etc.
[0104] In the above solution, a rule-based feature engineering approach can be used for feature extraction. That is, by defining feature extraction logic through preset business rules or experience, valuable features can be directly parsed from the raw data. Among them, time feature extraction: parses periodic attributes such as day, week, and month from the request timestamp (such as extracting "day of the week" and "whether it is a weekday" through time functions), or identifies peak request periods (such as peak request volume at 9:00 and 18:00 every day). Request feature extraction: directly classifies based on request type fields (such as "payment" and "query") and user role tags ("VIP" and "ordinary user"), or determines the request type based on keywords in the request URL path (such as " / api / pay"). System load feature extraction: directly extracts specific values of indicators such as CPU utilization, memory usage, and QPS (requests per second) from load data, or divides load levels (such as "high load" and "medium load") based on thresholds.
[0105] Statistical analysis methods can also be used for feature extraction, that is, to calculate the distribution, trend or correlation of data through mathematical statistical means to generate descriptive features. Among them, time feature extraction: calculate the statistics of historical requests in the time dimension (such as the mean and variance of daily request volume, and determine the periodic fluctuation pattern), or calculate the recent request frequency through a sliding window (such as the number of requests in the past 10 minutes). Request feature extraction: count the proportion of different request types (such as payment requests account for 30% of total requests), the request frequency distribution of user roles (the average daily request volume of VIP users is twice that of ordinary users), or analyze the correlation between request type and system load (such as payment requests have a greater impact on CPU load). System load feature extraction: calculate the historical statistical values of load indicators (such as the maximum, minimum, and average values of CPU utilization in the past hour), or divide the load interval by quantile (such as defining the top 10% of QPS as "high load period"), etc.
[0106] It can be understood that the above two feature extraction methods are merely examples listed to better understand the technical solutions of the embodiments of the present application, and are not intended to be the sole limitations of this embodiment.
[0107] Step 603: Standardize and normalize the request data features and the load data features to obtain model training samples.
[0108] After extracting the request data features of the historical request data and the load data features of the system load data, the request data features and the load data features can be standardized and normalized to obtain model training samples. Specifically, the feature data can be standardized and normalized to convert them to the same magnitude to obtain model training samples.
[0109] Step 604: The system load prediction model of the API gateway node is obtained based on the model training sample training.
[0110] After standardizing and normalizing the request data features and the load data features to obtain model training samples, a system load prediction model for the API gateway node can be trained based on the model training samples.
[0111] When training a model, you first select an appropriate model, such as a random forest regression model or a gradient boosting regression model. Then, you train the model using the training set data and adjust the model parameters to minimize prediction error. During training, you can use cross-validation (such as K-fold cross-validation) to evaluate the model's generalization ability and avoid overfitting.
[0112] During model training, you can choose an appropriate loss function to measure the model's prediction error. Common loss functions include mean squared error (MSE) and mean absolute error (MAE). At the end of each training cycle, the model's loss on the validation set is calculated to evaluate the model's training effectiveness based on the change in loss. When optimizing model parameters, you can use optimization algorithms such as gradient descent to adjust the model's parameters to minimize the loss.
[0113] After training the system load prediction model, the trained model can be deployed to the API gateway node for real-time prediction in the production environment. When the API gateway node is running, feature data can be collected in real time and input into the model for prediction, so as to dynamically adjust the priority of API requests or perform other optimization operations based on the prediction results. Figure 7 As shown in the figure, first, a deep learning model can be trained based on historical request data. The deep learning model can be used to predict the system load of the API gateway node at a certain time in the future. The priority scheduling mechanism can be used to dynamically adjust the scheduling policy of the priority queue based on the prediction results.
[0114] By training the system load prediction model, the embodiment of the present application can predict the system load in the future in real time when the system is running, and use the prediction results to dynamically adjust the priority of API requests, thereby preparing for resource scheduling and priority adjustment in advance.
[0115] Next, combine Figure 8 The process of adjusting strategies and preparing scheduling resources based on the predicted system load at the target time is described in detail.
[0116] Reference Figure 8 , shows a flow chart of the steps of a policy resource adjustment method provided by an embodiment of the present application. Figure 8 As shown, the policy resource adjustment method may include: step 801, step 802 and step 803.
[0117] Step 801: When the predicted system load is greater than or equal to a second load threshold, the scheduling priority of the classified sorting request data in the priority queue at the target time is adjusted to the target scheduling priority.
[0118] In this embodiment, the priority queue may include: a first priority queue, a second priority queue, and the third priority queue. The priority of the classified sorting request data in the first priority queue is higher than the priority of the classified sorting request data in the second priority queue, and the priority of the classified sorting request data in the second priority queue is higher than the priority of the classified sorting request data in the third priority queue. That is, the first priority queue is a high priority queue, the second priority queue is a medium priority queue, and the third priority queue is a low priority queue. It can be understood that in the same priority queue, the priorities of multiple service request data may be different. Specifically, it can be determined according to actual conditions, and this embodiment does not limit this.
[0119] The second load threshold refers to a preset system load threshold for determining whether the system load is too high. The specific value of the second load threshold can be determined according to business requirements, and this embodiment does not impose any restrictions on this.
[0120] After the predicted system load of the API gateway node at the future target time is predicted, it can be determined whether the predicted system load is greater than or equal to a second load threshold.
[0121] If the predicted system load threshold is less than the second load threshold, it means that the API gateway node is in a normal load state. At this time, the business data request at the target time can be scheduled according to the pre-defined fixed priority rules and pre-allocated scheduling resources.
[0122] If the predicted system load threshold is greater than or equal to the second load threshold, it means that the API gateway node is in a high load state. At this time, the scheduling priority of the classified sorting request data in the priority queue at the target time can be adjusted to the target scheduling priority.
[0123] The target scheduling priority may be: raising the priority of a portion of the classified sorting request data in the first priority queue, lowering the priority of a portion of the classified sorting request data in the second priority queue, and intercepting a portion of the classified sorting request data in the third priority queue. This implementation process is similar to the implementation process of determining the scheduling policy in step 502 above, and will not be further described in this embodiment.
[0124] Step 802: Determine target scheduling resource information corresponding to the classified sorting request data in the priority queue at the target time according to the target scheduling priority.
[0125] After obtaining the target scheduling priority, the target scheduling resource information corresponding to the classified sorting request data in the priority queue at the target moment can be determined according to the target scheduling priority. Among them, the target scheduling resource information is: increase the scheduling resources corresponding to the first priority queue, and lower the scheduling resources of the second priority queue and the third priority queue. In a specific implementation, according to the adjusted target priority, the system can allocate more resources to the first priority queue, such as increasing the number of CPU cores, memory size or network bandwidth. Correspondingly, the system can reduce the resource allocation of the second priority queue, which may include reducing the CPU time slice, memory usage limit or lowering the network priority. For the third priority queue, since the request is fully or partially intercepted, the system can completely stop allocating resources to the queue or only allocate very few resources, etc.
[0126] Step 803: Determine the target scheduling priority and the target scheduling resource information as the target scheduling policy at the target time.
[0127] After obtaining the target scheduling priority and target scheduling resource information, the target scheduling priority and target scheduling resource information can be used as a target scheduling policy at a target time. The target scheduling policy can schedule the classified and sorted request data when the target time is reached.
[0128] The embodiment of the present application can prioritize critical requests under high load conditions while ensuring the stability and response speed of the overall system through the above-mentioned scheduling priority adjustment and scheduling resource allocation methods.
[0129] In this embodiment, all data related to priority adjustment and request processing can also be collected and stored, and an analysis system can be established to generate reports regularly to evaluate the effectiveness of the scheduling strategy. By continuously monitoring and analyzing all priority adjustment and request processing results, the adjustment strategy can be continuously optimized to ensure continuous improvement and optimization of the system.
[0130] Next, combine Figure 9 The process of scheduling strategy testing is described in detail.
[0131] Reference Figure 9 , shows a flowchart of the steps of a scheduling strategy testing method provided by an embodiment of the present application. Figure 9 As shown, the scheduling strategy testing method may include: step 901, step 902 and step 903.
[0132] Step 901: Obtain API request data of the API gateway node under different system loads.
[0133] In this embodiment, when testing the optimal request scheduling strategy, API request data of the API gateway node under different system loads may be obtained.
[0134] Step 902: Use multiple request scheduling strategies to test the API request data under each system load respectively, and obtain performance indicators of the API gateway node under the multiple request scheduling strategies under each system load.
[0135] Furthermore, multiple request scheduling strategies can be used to test API request data under each system load, obtaining performance metrics for the API gateway node under each of these multiple request scheduling strategies. These multiple request scheduling strategies can include strategies such as adjusting the number of high-priority, medium-priority, and low-priority business data requests. In a simulation environment, different request scheduling strategies can be applied to API request data within each system load interval, and the performance metrics of each strategy under different loads, such as response time, throughput, and error rate, can be recorded.
[0136] Step 903: Determine the optimal request scheduling strategy corresponding to each system load based on the performance indicator.
[0137] Finally, the optimal request scheduling strategy for each system load can be determined based on the performance metrics. Specifically, the performance metrics of different strategies within each system load range can be compared and analyzed, focusing on the changing trends of key performance indicators, such as whether response time is shortened and throughput is improved. Then, based on the analysis of performance metrics, the optimal request scheduling strategy for each system load range can be determined. The optimal request scheduling strategy should be the request scheduling method that maximizes system performance (e.g., shortest response time, highest throughput, etc.) under a given load.
[0138] The test process can be as follows Figure 10 As shown, design and implement A / B testing of scheduling strategies to compare the effects of different strategies. Divide user traffic into multiple experimental groups and control groups, apply different priority scheduling strategies, and apply strategy A and strategy B to each experimental group to test through the A / B testing system. Compare the system performance indicators after the test. When performing statistics, statistical analysis methods (such as t-test, ANOVA, etc.) can be used to evaluate the significant differences between different strategies. Based on feedback data and A / B test results, automatically generate and adjust new priority scheduling strategies. Run the automated optimization engine to adjust the system's scheduling strategy in real time to cope with traffic changes.
[0139] The specific implementation process of A / B testing can be as follows Figure 11 As shown, the A / B testing system and its feedback data priority scheduling mechanism can generate new policies based on system load. This means the testing process can be similar to the one described above. The automated optimization engine automatically generates and adjusts new priority scheduling policies based on feedback data and A / B testing results, adjusting the system's scheduling policy in real time to accommodate traffic changes. Furthermore, the real-time priority scheduling mechanism can apply new policies in real time, enhancing the system's intelligence and adaptability.
[0140] The embodiment of the present application can determine the optimal request scheduling strategy for the API gateway node under different system loads by testing the optimal request scheduling strategy, thereby improving the overall performance of the system and user experience.
[0141] Next, combine Figure 12 The process of adjusting the current system scheduling strategy according to the current system load and the optimal request scheduling strategy is described in detail.
[0142] Reference Figure 12 , shows a flowchart of another method for adjusting the scheduling strategy provided by an embodiment of the present application. Figure 12 As shown, the scheduling strategy adjustment method may include: step 1201 and step 1202.
[0143] Step 1201: Determine the node load of the API gateway node according to the load indicator data.
[0144] In this embodiment, after obtaining the load index data of the API gateway node at the current time, the node load of the API gateway node can be determined according to the load index data.
[0145] Step 1202: Determine the scheduling strategy for the classified and sorted request data as the optimal request scheduling strategy corresponding to the node load.
[0146] After obtaining the node load of the API gateway node, the scheduling policy for the classified and sorted request data can be determined as the optimal request scheduling policy corresponding to the node load. Specifically, based on the correspondence between the load and the optimal request scheduling policy previously determined through testing or accumulated experience, the corresponding optimal request scheduling policy can be selected for the currently determined node load level. Furthermore, the selected optimal request scheduling policy can be applied to the request processing system of the API gateway node, and the processing priority of the classified and sorted request data in the priority queue can be adjusted according to the new scheduling policy. For example, under high load conditions, the priority of critical service requests can be increased, while the processing of non-critical requests can be reduced or suspended.
[0147] The embodiment of the present application determines the request scheduling strategy of the current node load according to the optimal request scheduling strategy under different loads tested in advance, thereby ensuring that the API gateway node can process requests in the optimal manner under different load conditions, thereby providing stable and efficient services.
[0148] In another specific implementation of the present application, after the predicted system load of the API gateway node at a future target time is obtained using a system load prediction model, the scheduling policy for the classified and sorted request data at the target time can be adjusted to the optimal request scheduling policy corresponding to the predicted system load. This ensures that the API gateway node can optimally process requests at the future target time, thereby providing stable and efficient service. This also provides a foundation for continuous system optimization and performance improvement.
[0149] The embodiment of the present application uses a data collection agent to collect request data from the API gateway node in real time, and dynamically adjusts the processing priority of the API request in combination with a machine learning model, which can ensure that key requests are processed first, optimize resource allocation, avoid resource waste and overload, and improve overall system performance and service quality.
[0150] Reference Figure 13 , shows a schematic diagram of the structure of a request processing device provided by an embodiment of the present application, such as Figure 13 As shown, the request processing device 1300 may include the following modules:
[0151] The data acquisition module 1310 is used to collect business request data and load indicator data using a data collection agent deployed on the API gateway node;
[0152] The request data acquisition module 1320 is used to classify and sort the service request data according to the initial priority rule to obtain classified and sorted request data;
[0153] a scheduling strategy determining module 1330, configured to determine a scheduling strategy for the classified and sorted request data based on the load indicator data, so as to schedule the classified and sorted request data according to the scheduling strategy;
[0154] The predicted load acquisition module 1340 is configured to input the API request data and system load data of the API gateway node within a first preset time period from the current time into a system load prediction model to obtain the predicted system load of the API gateway node at a future target time;
[0155] The scheduling strategy adjustment module 1350 is used to adjust the scheduling strategy of the classified sorting request data at the target time based on the predicted system load, obtain the target scheduling strategy of the target time, and, when the target time is reached, perform request scheduling on the classified sorting request data at the target time based on the target scheduling strategy.
[0156] Optionally, the request data acquisition module includes:
[0157] a request priority determination unit, configured to determine the request priority corresponding to each of the service request data according to the initial priority rule, the service type of each of the service request data, and the client identifier corresponding to each of the service request data;
[0158] a sorting request data acquiring unit, configured to sort each of the service request data according to the request priority to obtain sorting request data;
[0159] The request data storage unit is used to classify the sorting request data according to the priority threshold corresponding to at least one preset priority queue to obtain the classified sorting request data.
[0160] Optionally, the scheduling strategy determination module includes:
[0161] A system load determination unit, configured to determine a current system load of the API gateway node based on the load indicator data;
[0162] The scheduling strategy determining unit is used to determine the scheduling strategy of the classified sorting request data at the current time according to the current system load.
[0163] Optionally, the classification and sorting request data is cached in a priority queue, and the priority queue includes: the first priority queue, the second priority queue, and the third priority queue. The priority of the classification and sorting request data in the first priority queue is higher than the priority of the classification and sorting request data in the second priority queue, and the priority of the classification and sorting request data in the second priority queue is higher than the priority of the classification and sorting request data in the third priority queue.
[0164] The scheduling strategy determination unit includes:
[0165] a first policy determination subunit, configured to, when the current system load is greater than or equal to a first load threshold, determine that a scheduling policy for the classification and sorting request data in the priority queue is a first scheduling policy, wherein the first scheduling policy is: raising the priority of a portion of the classification and sorting request data in the first priority queue, lowering the priority of a portion of the classification and sorting request data in the second priority queue, and intercepting a portion of the classification and sorting request data in the third priority queue;
[0166] The second strategy determination subunit is used to determine that the scheduling strategy of the classified sorting request data in the priority queue is a second scheduling strategy when the current system load is less than the first load threshold. The second scheduling strategy is: the scheduling priority of the classified sorting request data in the priority queue remains unchanged.
[0167] Optionally, the device further comprises:
[0168] a historical data acquisition module, configured to acquire historical request data of the API gateway node within a second preset time period from the current time, and system load data corresponding to the historical request data, where the second preset time period is greater than the first preset time period;
[0169] a data feature extraction module, configured to extract request data features of the historical request data and load data features of the system load data;
[0170] A model sample acquisition module is used to standardize and normalize the request data features and the load data features to obtain model training samples;
[0171] The load prediction model training module is used to train the system load prediction model of the API gateway node based on the model training samples.
[0172] Optionally, the scheduling strategy adjustment module includes:
[0173] A priority adjustment unit, configured to adjust the scheduling priority of the classified sorting request data in the priority queue at the target time to a target scheduling priority when the predicted system load is greater than or equal to a second load threshold;
[0174] a scheduling resource determination unit, configured to determine target scheduling resource information corresponding to the classified sorting request data in the priority queue at the target time according to the target scheduling priority;
[0175] a target policy determining unit, configured to determine the target scheduling priority and the target scheduling resource information as the target scheduling policy at the target moment;
[0176] The target scheduling priority is: raising the priority of some classified sorting request data in the first priority queue, lowering the priority of some classified sorting request data in the second priority queue, and intercepting some classified sorting request data in the third priority queue;
[0177] The target scheduling resource information is: increasing the scheduling resources corresponding to the first priority queue, and decreasing the scheduling resources corresponding to the second priority queue and the third priority queue.
[0178] Optionally, the device further comprises:
[0179] An API request acquisition module is used to obtain API request data of the API gateway node under different system loads;
[0180] A performance indicator acquisition module is used to test the API request data under each system load using multiple request scheduling strategies to obtain performance indicators of the API gateway node under the multiple request scheduling strategies for each system load;
[0181] The optimal strategy determination module is used to determine the optimal request scheduling strategy corresponding to each system load based on the performance indicator.
[0182] Optionally, the scheduling strategy determination module includes:
[0183] A node load determination unit, configured to determine the node load of the API gateway node based on the load indicator data;
[0184] A policy determination unit is used to determine the scheduling policy of the classified and sorted request data as the optimal request scheduling policy corresponding to the node load.
[0185] Optionally, the scheduling strategy adjustment module includes:
[0186] The fourth policy adjustment unit is configured to adjust the scheduling policy of the classified and sorted request data at the target time to an optimal request scheduling policy corresponding to the predicted system load.
[0187] The embodiment of the present application uses a data collection agent to collect request data from the API gateway node in real time, and dynamically adjusts the processing priority of the API request in combination with a machine learning model, which can ensure that key requests are processed first, optimize resource allocation, avoid resource waste and overload, and improve overall system performance and service quality.
[0188] The present application also provides an electronic device, such as Figure 14As shown, it includes a processor 1401, a communication interface 1402, a memory 1403 and a communication bus 1404, wherein the processor 1401, the communication interface 1402, and the memory 1403 communicate with each other through the communication bus 1404.
[0189] Memory 1403, used for storing computer programs;
[0190] The processor 1401 is configured to execute the program stored in the memory 1403 by performing the following steps:
[0191] Use the data collection agent deployed on the API gateway node to collect business request data and load indicator data;
[0192] Classifying and sorting the service request data according to the initial priority rule to obtain classified and sorted request data;
[0193] Determining a scheduling strategy for the classification and sorting request data at a current time based on the load indicator data, so as to schedule the classification and sorting request data according to the scheduling strategy;
[0194] Inputting API request data and system load data of the API gateway node within a first preset time period from the current time into a system load prediction model to obtain a predicted system load of the API gateway node at a future target time;
[0195] Based on the predicted system load, the scheduling strategy of the classified sorting request data at the target time is adjusted to obtain the target scheduling strategy of the target time, and when the target time is reached, the classified sorting request data at the target time is requested to be scheduled based on the target scheduling strategy.
[0196] The communication bus mentioned in the terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.
[0197] The communication interface is used for communication between the above terminal and other devices.
[0198] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.
[0199] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0200] In another embodiment provided by the present application, a computer-readable storage medium is further provided. The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the request processing method described in any one of the above embodiments.
[0201] In another embodiment provided by the present application, a computer program product including instructions is further provided, on which a computer program is stored. When the computer program is run on a computer, the computer is enabled to execute any of the request processing methods described above.
[0202] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0203] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0204] Each embodiment in this specification is described in a related manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For related parts, refer to the description of the method embodiment.
[0205] The above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the scope of protection of the present application.
Claims
1. A request processing method, characterized in that: include: Use the data collection agent deployed on the API gateway node to collect business request data and load indicator data; Classifying and sorting the service request data according to the initial priority rule to obtain classified and sorted request data; Determining a scheduling strategy for the classification and sorting request data at a current time based on the load indicator data, so as to schedule the classification and sorting request data according to the scheduling strategy; Inputting API request data and system load data of the API gateway node within a first preset time period from the current time into a system load prediction model to obtain a predicted system load of the API gateway node at a future target time; Based on the predicted system load, the scheduling strategy of the classified sorting request data at the target time is adjusted to obtain the target scheduling strategy of the target time, and when the target time is reached, the classified sorting request data at the target time is requested to be scheduled based on the target scheduling strategy.
2. The method according to claim 1, characterized in that The process of classifying and sorting the service request data according to the initial priority rule to obtain classified and sorted request data includes: determining a request priority corresponding to each of the service request data according to the initial priority rule, a service type of each of the service request data, and a client identifier corresponding to each of the service request data; Sorting the service request data according to the request priorities to obtain sorted request data; The sorting request data is classified according to a priority threshold corresponding to at least one preset priority queue to obtain the classified sorting request data.
3. The method according to claim 1, characterized in that The determining, based on the load indicator data, a scheduling strategy for the classification and sorting request data at the current time includes: Determine the current system load of the API gateway node based on the load indicator data; According to the current system load, a scheduling strategy for the classified and sorted request data at the current time is determined.
4. The method according to claim 3, characterized in that The classification and sorting request data is cached in a priority queue, the priority queue including: the first priority queue, the second priority queue, and the third priority queue, the priority of the classification and sorting request data in the first priority queue is higher than the priority of the classification and sorting request data in the second priority queue, and the priority of the classification and sorting request data in the second priority queue is higher than the priority of the classification and sorting request data in the third priority queue; The determining, based on the current system load, a scheduling strategy for the classified and sorted request data at the current time includes: When the current system load is greater than or equal to a first load threshold, determining that a scheduling strategy for the classification and sorting request data in the priority queue is a first scheduling strategy, wherein the first scheduling strategy is: in the first priority queue, the priority of a portion of the classification and sorting request data is increased, in the second priority queue, the priority of a portion of the classification and sorting request data is decreased, and in the third priority queue, a portion of the classification and sorting request data is intercepted; When the current system load is less than the first load threshold, the scheduling policy for the classified sorting request data in the priority queue is determined to be a second scheduling policy, and the second scheduling policy is: the scheduling priority of the classified sorting request data in the priority queue remains unchanged.
5. The method according to claim 1, wherein Before inputting the API request data and system load data of the API gateway node within a first preset time period from the current time into the system load prediction model to obtain the predicted system load of the API gateway node at a future target time, the method further includes: Obtain historical request data of the API gateway node within a second preset time period from the current time, and system load data corresponding to the historical request data, where the second preset time period is greater than the first preset time period; extracting request data features of the historical request data and load data features of the system load data; Standardizing and normalizing the request data features and the load data features to obtain model training samples; The system load prediction model of the API gateway node is obtained based on the model training sample training.
6. The method according to claim 4, characterized in that The step of adjusting the scheduling strategy of the classified sorted request data at the target time based on the predicted system load to obtain the target scheduling strategy at the target time includes: When the predicted system load is greater than or equal to a second load threshold, adjusting the scheduling priority of the classified sorting request data in the priority queue at the target time to a target scheduling priority; Determining target scheduling resource information corresponding to the classified sorting request data in the priority queue at the target time according to the target scheduling priority; Determining the target scheduling priority and the target scheduling resource information as the target scheduling policy at the target time; The target scheduling priority is: raising the priority of some classified sorting request data in the first priority queue, lowering the priority of some classified sorting request data in the second priority queue, and intercepting some classified sorting request data in the third priority queue; The target scheduling resource information is: increasing the scheduling resources corresponding to the first priority queue, and decreasing the scheduling resources corresponding to the second priority queue and the third priority queue.
7. The method according to claim 1, characterized in that The method further comprises: Obtaining API request data of the API gateway node under different system loads; Using multiple request scheduling strategies to test the API request data under each system load, respectively, to obtain performance indicators of the API gateway node under the multiple request scheduling strategies for each system load; An optimal request scheduling strategy corresponding to each system load is determined based on the performance indicators.
8. The method according to claim 7, characterized in that The determining, based on the load indicator data, a scheduling strategy for the classification and sorting request data at the current time includes: Determine the node load of the API gateway node based on the load indicator data; The scheduling strategy for the classified and sorted request data is determined as the optimal request scheduling strategy corresponding to the node load.
9. The method according to claim 7, characterized in that The step of adjusting the scheduling strategy of the classified sorted request data at the target time based on the predicted system load to obtain the target scheduling strategy at the target time includes: The scheduling strategy of the classified and sorted request data at the target time is adjusted to the optimal request scheduling strategy corresponding to the predicted system load.
10. A request processing device, characterized in that: include: The data acquisition module is used to collect business request data and load indicator data using the data collection agent deployed on the API gateway node; A request data acquisition module is used to classify and sort the service request data according to the initial priority rule to obtain classified and sorted request data; a scheduling strategy determining module, configured to determine a scheduling strategy for the classified and sorted request data according to the load indicator data, so as to perform request scheduling on the classified and sorted request data according to the scheduling strategy; A predicted load acquisition module, configured to input API request data and system load data of the API gateway node within a first preset time period from the current time into a system load prediction model to obtain a predicted system load of the API gateway node at a future target time; A scheduling strategy adjustment module is used to adjust the scheduling strategy of the classified sorting request data at the target time based on the predicted system load, obtain the target scheduling strategy of the target time, and, when the target time is reached, request scheduling of the classified sorting request data at the target time based on the target scheduling strategy.
11. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; A processor, configured to implement the method steps described in any one of claims 1 to 9 when executing a program stored in a memory.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
13. A computer program product comprising instructions, on which a computer program is stored, characterized in that When the computer program is run on a computer, the computer is caused to execute the method according to any one of claims 1 to 9.
Citation Information
Cited By
API request flow control method, system and device and medium
CN121077980A