Adaptive scheduling method based on real-time traffic characteristic sensing and distribution forwarding server
Patent Information
- Application Number
- CN202610935786.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-06-26
AI Technical Summary
[0007]1、调度策略静态固化,无法感知实时流量变化:
[0065] 1. By performing online clustering of multi-dimensional features of front-end traffic, real-time identification of traffic patterns and business semantic awareness are achieved, breaking through the limitation of traditional load balancing that can only perceive basic traffic indicators. It can accurately distinguish different traffic types such as normal business, sudden flash sales, and malicious crawlers, providing core decision-making basis for subsequent adaptive intelligent scheduling.
Smart Images

Figure CN122476066B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of computer networks and distributed systems, and in particular to an adaptive scheduling method and distribution forwarding server based on real-time traffic feature awareness. Background Technology
[0002] In current mainstream distributed service architectures, reverse proxy servers or cloud vendor load balancers are typically used as the request distribution entry point. Their typical workflow is as follows:
[0003] 1. The client initiates a request to the distribution and forwarding server;
[0004] 2. The server selects a backend service node based on a preset strategy (such as round-robin, weighted round-robin, minimum number of connections, IP hash).
[0005] 3. Forward the request to that node and return a response.
[0006] However, the aforementioned prior art has the following significant drawbacks:
[0007] 1. The scheduling strategy is static and fixed, making it impossible to detect real-time traffic changes:
[0008] Preset strategies (such as polling) do not take into account the real-time status of backend nodes, such as current CPU, memory, network I / O, and request processing latency, which can lead to some nodes being overloaded while others are idle, resulting in uneven resource utilization.
[0009] 2. Lacks adaptive capability to sudden traffic spikes or abnormal requests:
[0010] In scenarios such as flash sales, live streaming, and DDoS attacks, the surge in requests or the emergence of a large number of slow requests can cause traditional schedulers to be unable to dynamically isolate abnormal traffic or switch to backup links, resulting in an overall service collapse.
[0011] 3. Failure to differentiate between request type and service quality requirements:
[0012] All requests (such as login authentication, file download, and heartbeat keep-alive) are treated equally, while high-priority or low-latency sensitive requests (such as game commands) may time out due to queuing.
[0013] 4. Weak support for cross-regional / multi-link scenarios:
[0014] In hybrid cloud or multi-carrier network environments, it is impossible to intelligently select the optimal access path based on the client's geographical location and network quality (RTT, packet loss rate). Summary of the Invention
[0015] To achieve efficient and intelligent scheduling of massive client requests, this application provides an adaptive scheduling method and distribution forwarding server based on real-time traffic feature awareness.
[0016] Firstly, this application provides an adaptive scheduling method based on real-time traffic feature awareness, employing the following technical solution:
[0017] An adaptive scheduling method based on real-time traffic feature awareness includes the following steps:
[0018] Collect multi-dimensional real-time metrics to obtain front-end traffic characteristics and back-end node status;
[0019] Sliding window clustering is performed based on the aforementioned front-end traffic characteristics to identify traffic patterns;
[0020] Based on the traffic pattern, a corresponding scheduling template is selected and a node health score is calculated in combination with the backend node status. Based on the node health score, a dynamic weight for the allocation of traffic to the node is generated and adjusted.
[0021] Obtain the protocol and / or content characteristics corresponding to the request to select a forwarding strategy, and perform differentiated forwarding of the request in combination with the dynamic weight.
[0022] In some embodiments, the front-end traffic characteristics include requests per second, request type distribution, average request size, client IP geographic distribution, and TLS handshake time. Sliding window clustering based on these front-end traffic characteristics to identify traffic patterns includes the following steps:
[0023] Based on the front-end traffic characteristics within the preset time window, a multi-dimensional feature vector is constructed;
[0024] Clustering algorithms are used to cluster the multidimensional feature vectors, and the corresponding traffic patterns are identified based on the output results.
[0025] In some embodiments, the backend node status includes CPU utilization, memory usage, number of active connections, average response latency, and error rate. The process of selecting a corresponding scheduling template based on the traffic pattern and calculating a node health score in conjunction with the backend node status includes the following steps:
[0026] The node health score is calculated based on the following formula;
[0027]
[0028] Wherein, α, β, γ are configurable weight coefficients, and their sum is 1. The magnitude of each weight coefficient is determined by the scheduling template corresponding to the traffic pattern. This represents the average response latency of the i-th backend node within the most recent sliding window; This represents the minimum average response latency among all currently available nodes; Indicates the maximum average response latency among all currently available nodes; ϵ is a preset constant; This represents the percentage of CPU usage. This represents the percentage of recent erroneous requests.
[0029] In some embodiments, a corresponding scheduling template is selected based on the traffic pattern, including offline evaluation during the training phase, specifically including:
[0030] Historical traffic data is acquired to replay tests on several different scheduling templates.
[0031] Based on the historical data, the backend node status is used to calculate key indicators, including the average response latency reduction rate, CPU utilization standard deviation, and error rate fluctuation range.
[0032] A comprehensive score is calculated based on the key indicators, and the scheduling template with the highest comprehensive score is selected as the default configuration.
[0033] In some embodiments, a corresponding scheduling template is selected based on the traffic pattern, including online feedback during the runtime phase, specifically including:
[0034] The traffic is split into a small proportion of traffic and a large proportion of traffic, and the small proportion of traffic is allocated to the new scheduling template, while the large proportion of traffic is allocated to the old scheduling template.
[0035] The measured indicators corresponding to the new scheduling template are continuously collected and compared with the benchmark indicators of the old scheduling template to calculate the template score.
[0036] The decision to allocate full traffic to the new scheduling template is based on the numerical relationship between the template score and the score threshold.
[0037] In some embodiments, obtaining the protocol and / or content characteristics corresponding to the request to select a forwarding strategy, and combining the dynamic weight to perform differentiated forwarding of the request, includes the following steps:
[0038] By default, the request is forwarded to the node with the highest health score, and the configurable weight coefficient in the scheduling template is forcibly configured based on the protocol objective.
[0039] And / or,
[0040] The request is routed to a dedicated high-bandwidth egress link node pool, and the upper limit of bandwidth utilization of the high-bandwidth egress link node pool is adjusted.
[0041] And / or,
[0042] The source IP of the request is subject to token bucket rate limiting control, and subsequent new requests are directed to a dedicated anti-scraping processing cluster for deep verification.
[0043] In some embodiments, the following steps are also included:
[0044] Continuously monitor the status of each backend node under the corresponding traffic mode;
[0045] When any preset abnormal condition is met within multiple consecutive monitoring windows, the circuit breaker and degradation mechanism is triggered.
[0046] Executing the circuit breaker and degradation mechanism includes: pausing the forwarding of the request to the original target backend node or node pool, and enabling a backup processing scheme. The backup processing scheme includes switching the paused request to the backup node pool and starting a fast retry strategy or returning to the prompt interface and recording the event log.
[0047] In each subsequent sliding window, the recovery status of the original target backend node is monitored, and the traffic weight is gradually restored when the recovery benchmark is met.
[0048] In some embodiments, for the traffic pattern, the following steps are also included:
[0049] Configure a time series library, which stores full time series data of each traffic pattern change in historical time.
[0050] Collect the change paths corresponding to several current consecutive traffic patterns, and match them with historical paths in the time series database based on time and content features to output spatiotemporal prediction information, which includes predicted traffic patterns, predicted peak values, and predicted arrival times.
[0051] A pre-scheduling strategy is generated and executed based on the spatiotemporal prediction information, while deviation calculation is performed on the spatiotemporal prediction information based on the real-time collected front-end traffic characteristics.
[0052] Based on the result of the deviation calculation, it is determined whether to continue executing the pre-scheduling strategy.
[0053] In some embodiments, the pre-scheduling strategy includes:
[0054] Resource pre-positioning: Pre-lock a preset proportion of computing power resources in the backup node pool for the predicted peak value and gradually increase the dynamic weight of the corresponding backend node;
[0055] Link pre-planning: Pre-establish connections to the backend node with the highest dynamic weight;
[0056] Template preloading: The scheduling template corresponding to the predicted traffic pattern is preloaded and added to the local cache.
[0057] Secondly, this application provides a distribution and forwarding server, which adopts the following technical solution:
[0058] A distribution and forwarding server for implementing the above method includes:
[0059] The traffic acquisition module is used to collect multi-dimensional real-time metrics of requests to obtain front-end traffic characteristics and back-end node status.
[0060] The feature analysis engine is used to perform sliding window clustering based on the front-end traffic features to identify traffic patterns. It is also used to select the corresponding scheduling template based on the traffic patterns and calculate the node health score in combination with the back-end node status. Based on the node health score, it generates and adjusts the dynamic weights for node traffic allocation.
[0061] Strategy Decision Center: Used to store the scheduling templates and generate the final scheduling strategy;
[0062] The intelligent forwarding engine is used to obtain the protocol and / or content characteristics corresponding to the request to select a forwarding strategy, and to perform differentiated forwarding of the request in combination with the dynamic weight;
[0063] Configure the management interface and support dynamic updates of scheduling parameters via API.
[0064] The technical solutions provided by the embodiments of this application have the following technical effects:
[0065] 1. By performing online clustering of multi-dimensional features of front-end traffic, real-time identification of traffic patterns and business semantic awareness are achieved, breaking through the limitation of traditional load balancing that can only perceive basic traffic indicators. It can accurately distinguish different traffic types such as normal business, sudden flash sales, and malicious crawlers, providing core decision-making basis for subsequent adaptive intelligent scheduling.
[0066] 2. An automatic matching mechanism of "traffic pattern - scheduling template" was constructed, and forwarding weight was dynamically calculated in combination with the real-time health status of backend nodes, realizing real-time linkage between scheduling strategy and business scenario and node load, and solving the problem of uneven node load caused by traditional static scheduling strategy.
[0067] 3. Differentiated traffic splitting and forwarding are performed based on request protocol type and content characteristics. Dedicated node pools and forwarding rules are matched for long connections, large files and ordinary requests respectively, which realizes accurate matching between request characteristics and backend node resource capabilities and greatly reduces request tail latency.
[0068] 4. Based on the sliding window-based multidimensional anomaly detection and automatic circuit breaking and degradation, as well as the smooth back-switch closed-loop mechanism, the system achieves automatic isolation of abnormal nodes and malicious traffic, thus avoiding fault propagation and service avalanche. Attached Figure Description
[0069] Figure 1 This is a schematic diagram illustrating the steps of an adaptive scheduling method based on real-time traffic feature awareness provided in this embodiment.
[0070] Figure 2 This is a schematic diagram of the module of the adaptive scheduling system based on real-time traffic feature perception provided in the embodiments of this application. Detailed Implementation
[0071] To better understand the purpose, technical solutions, and advantages of this application, it has been described and illustrated below with reference to the accompanying drawings and embodiments. However, those skilled in the art should understand that this application can be implemented without these details. In some cases, to avoid obscuring various aspects of this application due to unnecessary description, well-known methods, processes, systems, components, and / or circuits already described at a higher level will not be elaborated upon. It will be apparent to those skilled in the art that various modifications can be made to the embodiments disclosed in this application, and the general principles defined in this application can be applied to other embodiments and application scenarios without departing from the principles and scope of this application. Therefore, this application is not limited to the illustrated embodiments, but conforms to the broadest scope consistent with the scope of protection claimed in this application.
[0072] It should be noted that the descriptions of these embodiments are for the purpose of aiding understanding the present invention, but do not constitute a limitation thereof. Furthermore, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0073] In the description of this application, "several" means one or more, "more than" means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. The use of "first" and "second" in the description is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features.
[0074] In the description of this application, the terms "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples.
[0075] like Figure 1 As shown in the figure, this application discloses an adaptive scheduling method based on real-time traffic feature awareness, including the following steps:
[0076] The S100 performs multi-dimensional real-time indicator collection to obtain front-end traffic characteristics and back-end node status.
[0077] The traffic acquisition module captures network and application layer metrics in real time, and the eBPF program is mounted to the kernel network stack to collect raw TCP / HTTP layer metrics, supporting Netlink to transmit application logs back.
[0078] This is used to sample the traffic characteristics of front-end requests and the status parameters of back-end nodes.
[0079] S200 uses sliding window clustering based on front-end traffic characteristics to identify traffic patterns.
[0080] Specifically, front-end traffic characteristics include the following data:
[0081] Requests per second (QPS), request type distribution (HTTP Method / path), average request size, client IP geographic distribution, and TLS handshake time.
[0082] Specifically, requests per second are measured by packet capture using the eBPF program to assess traffic scale; request type distribution is obtained through HTTP header parsing for business semantic identification; average request size is measured by TCP packet length to identify traffic type; client IP geographic distribution is obtained through IP database matching and entropy calculation for geographic concentration analysis; and TLS handshake time is obtained through eBPF hook recording for network quality assessment.
[0083] Specifically, the following steps are included:
[0084] S210 constructs a multi-dimensional feature vector based on the front-end traffic characteristics within a preset time window.
[0085] S220 uses a clustering algorithm to cluster multidimensional feature vectors and identifies the corresponding traffic patterns based on the output results.
[0086] A lightweight sliding window K-Means++ algorithm is used to cluster the request stream within the last 5 seconds according to the feature vector [QPS, avg_size, login_ratio, geo_entropy].
[0087] QPS is the average number of requests per second within the sliding window, avg_size is the average request body size within the sliding window, and login_ratio represents the proportion of login / authentication requests within the window.
[0088] geo_entropy is the entropy value of the client IP address distribution, calculated using the following formula: n represents the total number of geographical regions to which the current window's access requests belong. Represented as the proportion of requests from the i-th region to the total number of requests, this metric is used to quantify the geographical dispersion of traffic and determine whether there are targeted single-point clustering attacks.
[0089] Based on the clustering results, corresponding traffic patterns are identified, and a sliding window W of length L is maintained. An improved lightweight K-Means++ algorithm is used for real-time clustering.
[0090] First, initialize the cluster centers: randomly select a feature vector from the sliding window W as the first cluster center; for each subsequent center, select it from W with a probability proportional to the square of the distance to the previously selected center. Then, iteratively perform the following steps until the center points are stable or the maximum number of iterations is reached: (a) assign the nearest cluster center to each feature point within the window; (b) recalculate the mean of all points within each cluster as the new cluster center.
[0091] Define the traffic pattern affiliation degree. When the weight ratio of a certain traffic pattern exceeds a preset threshold (such as 0.7), the system will automatically switch to the corresponding scheduling template.
[0092] Traffic patterns include "normal browsing", "flash sale", and "web crawler scanning".
[0093] Traditional load balancers only perceive the number of connections or QPS, failing to identify different traffic patterns. This application combines the semantic features of frontend requests (such as path distribution entropy, authentication ratio, geographical concentration, and average request body size) with a sliding time window, using an online clustering algorithm to output traffic pattern labels in real time. This mechanism is a prerequisite for subsequent intelligent scheduling.
[0094] S300 selects the corresponding scheduling template based on the traffic pattern and calculates the node health score in combination with the backend node status. Based on the node health score, it generates and adjusts the dynamic weight for the allocation of traffic to the node.
[0095] Backend node status includes CPU utilization, memory usage, number of active connections, average response latency (P95), and error rate (5xx percentage).
[0096] The node health score is calculated based on the following formula;
[0097]
[0098] Wherein, α, β, and γ are configurable weight coefficients, and their sum is 1. The magnitude of each weight coefficient is input through the configuration management interface, which controls the degree of influence of each indicator and is determined by the scheduling template corresponding to the traffic mode.
[0099] The scheduling templates include, but are not limited to, the aforementioned "normal browsing", "flash sale", and "web crawler scanning". Each scheduling template corresponds to a set of weight coefficients and traffic distribution rules, which can be dynamically updated through the configuration management interface. The system automatically matches the optimal template based on the traffic pattern to achieve fine-grained scheduling.
[0100] Template example:
[0101] Normal browsing mode: α=0.4, β=0.3, γ=0.3 (balancing latency and load);
[0102] Flash sale mode: α=0.6, β=0.2, γ=0.2 (low latency priority);
[0103] Crawler scanning mode: α=0.1, β=0.8, γ=0.1 (primarily rate limiting);
[0104] API call pattern: α=0.3, β=0.5, γ=0.2 (stability first).
[0105] This represents the P95 response latency (in milliseconds) of the i-th backend node within the most recent sliding window. It is calculated using eBPF packet capture and timestamps to reflect the node's processing speed, with lower values taking precedence.
[0106] This represents the minimum P95 response latency among all currently available nodes. Within the same sliding window, it iterates through all available backend nodes and selects the minimum P95 response latency. The minimum value is used as the lower bound of the normalization benchmark to ensure that the delay index is mapped to the [0,1] interval.
[0107] This represents the maximum average response latency among all currently available nodes. Within the same sliding window, it iterates through all available backend nodes and selects the maximum response latency. The maximum value is used as the upper limit of the normalization benchmark to eliminate the impact of absolute latency differences under different business scenarios;
[0108] The preset constant is 1ms by default. It can be adjusted through the configuration management interface, but usually no modification is needed to prevent the denominator from becoming zero when all node delays are equal, thus ensuring the robustness of the formula.
[0109] This represents the percentage of CPU utilization, obtained via system commands. It is used to prevent overload, and high values reduce the weight of CPU usage.
[0110] The percentage of recent erroneous requests is obtained through HTTP status codes (5xx), which is used to control the impact of various indicators.
[0111] The scheduling weight is represented as the probability of selecting a certain backend node. Therefore, through the above scheme, the scheduling weight no longer depends on the backend CPU or latency, but integrates both aspects of information:
[0112] Backend node health status (including kernel-level metrics such as I / O wait and context switching);
[0113] The system automatically matches the optimal weight allocation logic based on the currently identified traffic patterns (such as "flash sales" requiring low latency and "large file uploads" requiring high bandwidth).
[0114] At the same time, the selected scheduling template is further scored. Note that the scoring here is not directly on the scheduling template itself, but rather indirectly evaluated through system performance indicators after its execution. The scoring mechanism includes two types:
[0115] In other embodiments, a corresponding scheduling template is selected based on the traffic pattern, including offline evaluation during the training phase, specifically including:
[0116] S320 acquires historical traffic data to perform replay tests on several different scheduling templates.
[0117] S321 calculates key metrics based on the status of backend nodes using historical data. These key metrics include the average response (P99) latency decrease rate, CPU utilization standard deviation, and error rate fluctuation.
[0118] S322 calculates a comprehensive score based on key indicators and selects the scheduling template with the highest comprehensive score as the default configuration.
[0119] Each key metric is assigned a sub-score based on its numerical value. Generally, lower latency corresponds to a higher score, more balanced CPU performance corresponds to a higher score, and smaller error rate fluctuations correspond to a higher score.
[0120] Then, the overall score is obtained by accumulating or weighting the scores of each sub-score, and the scheduling template with the highest overall score is selected as the corresponding default configuration.
[0121] In other embodiments, a corresponding scheduling template is selected based on the traffic pattern, including online feedback during the runtime phase, specifically including:
[0122] S330 splits traffic into small and large proportions, allocating the small proportions to the new scheduling template and the large proportions to the old scheduling template.
[0123] S331 continuously collects the measured indicators corresponding to the new scheduling template and compares them with the benchmark indicators of the old scheduling template to calculate the template score.
[0124] S332, based on the numerical relationship between template score and score threshold, selects whether to allocate the full amount of traffic to the new scheduling template.
[0125] An A / B testing mechanism is introduced, where 10% of the traffic is initially allocated to the newly generated scheduling template, while the remaining 90% continues to use the current old online template.
[0126] During the test period (default 5 minutes), key performance indicators under the new scheduling template are collected, mainly including P99 latency and CPU utilization. At the same time, the template score is calculated based on the above indicators. When the score is higher than the score threshold, the new template is determined to be better, and the traffic on the new scheduling template is automatically switched to full. If the score is not higher than the score threshold, a rollback operation is performed to restore the original template.
[0127] Specifically, the formula for template scoring is:
[0128] .
[0129] in, and This serves as the baseline performance metric for the current online templates. and These are the actual measured metrics of the new template within the A / B test window.
[0130] When the calculated template score is greater than 0.15, the new template is considered to be better.
[0131] in, and The weights represent the preference for latency and stability.
[0132] This reflects the importance the business places on response speed. The higher the value, the more the system tends to choose the scheduling template with lower P99 latency.
[0133] This reflects the requirement for balanced backend resources. The larger the value, the more the system tends to select the template with smaller fluctuations in CPU utilization.
[0134] and These are not arbitrarily set constants, but configurable parameters reflecting business service quality preferences. Their values are determined by the business type and can be dynamically adjusted via API. This mechanism aligns A / B testing results with real business objectives, avoiding situations where technical metrics are optimized but business performance suffers. Two weights. and The sum is 1.
[0135] S400: Obtain the protocol and / or content characteristics corresponding to the request to select a forwarding strategy, and perform differentiated forwarding of the request by combining dynamic weights.
[0136] Different request protocol types and different request content characteristics (such as size, priority, etc.) correspond to different forwarding and distribution strategies.
[0137] in,
[0138] In certain situations, when the request is made using the HTTP / 2 or WebSocket protocol, the distribution and forwarding server will prioritize allocating the request to the backend node with the highest health score. At the same time, the health score is calculated using a scheduling template for low-latency scenarios, which means that the parameter configuration satisfies γ=0 (i.e., ignoring the error rate item) and sets α to be no less than 0.7 to enhance the sensitivity to response latency.
[0139] Therefore, through this scoring mechanism, the system can prioritize nodes with low latency and stable load to process high-priority requests.
[0140] And / or,
[0141] In certain situations, when the size of the requested content exceeds a preset threshold, such as when the content-length (HTTP message length) of the request is greater than 10MB and the request path matches the preset large file download rules, the distribution and forwarding server will route the request to a dedicated high-bandwidth egress link node pool. The bandwidth utilization limit of the node pool is set to 70% to ensure transmission quality.
[0142] And / or,
[0143] When a client from the same / 24 network segment (i.e., a Class C IP segment) initiates more than 500 requests within 1 second, it is considered a potential malicious act. The system automatically implements token bucket rate limiting control on the source IP (initial token count = 100, replenished by 50 per second) and redirects subsequent requests to a dedicated anti-scraping processing cluster for in-depth verification.
[0144] In cases where other protocols or content do not match any requirements, they are normally assigned to the default node pool.
[0145] The aforementioned protocol identification and content feature extraction are both completed by the intelligent forwarding engine based on HTTP header fields and TCP traffic patterns, without the need to parse complete application layer data packets.
[0146] After determining the target node pool, fine-grained routing is further implemented based on request protocol type (such as HTTP / 2, WebSocket) or content characteristics (such as Content-Length > 10MB, URL containing / pay). For example, long-lived connection requests are directed to the low-latency pool, and large file requests are directed to the high-throughput pool. This strategy significantly improves the efficiency of backend resource matching.
[0147] In other embodiments, the following steps are also included:
[0148] S500 continuously monitors the status of each backend node under the corresponding traffic mode.
[0149] S510 triggers the circuit breaker and degradation mechanism when any preset abnormal condition is met within multiple consecutive monitoring windows.
[0150] If a traffic cluster meets any of the preset abnormal conditions within three consecutive sliding windows, the circuit breaker degradation strategy will be triggered. The abnormal conditions include:
[0151] Error rate (5xx status code percentage) consistently exceeds 20%;
[0152] The P99 response latency increased by more than 300% compared to the previous window;
[0153] The number of active connections increased by more than 200%, and CPU utilization also increased by more than 50%.
[0154] It should be noted that the threshold values corresponding to the above conditions can be adjusted based on the actual situation.
[0155] S520 executes the circuit breaker and degradation mechanism, including: suspending the forwarding of requests to the original target backend node or node pool, and enabling the backup handling scheme. The backup handling scheme includes switching the suspended requests to the backup node pool and starting the fast retry strategy or returning the prompt interface and recording the event log.
[0156] When the circuit breaker degradation mechanism is triggered, the system will execute the following actions:
[0157] First, all new requests for this traffic cluster are paused from being forwarded to the original node pool. Then, the system automatically switches to the backup node pool and initiates a fast retry strategy (retrying immediately after the first failure). If the backup pool is overloaded, a preset friendly prompt page (such as "Service busy, please try again later") is returned, and event logs are recorded for operation and maintenance analysis.
[0158] The S530 monitors the recovery status of the original target backend node in each subsequent sliding window and gradually restores the traffic weight when the recovery benchmark is met.
[0159] In each subsequent sliding window, the recovery status of the original node is monitored. When its error rate is less than the preset value (e.g., 5%) and the P99 delay is less than the baseline multiplied by the preset coefficient (e.g., 1.2), the traffic weight is gradually restored to achieve smooth regression.
[0160] In current traditional load balancing solutions, traffic arrives first, then features are collected, patterns are identified, scheduling policies are generated, and forwarding is executed. The scheduling policy always lags behind traffic changes. In scenarios with sudden traffic surges, such as flash sales or live streams, the problem inevitably arises where traffic overwhelms nodes before scheduling follows. To address this issue, this application constructs a pre-emptive scheduling closed loop that includes pattern-level spatiotemporal prediction, pre-scheduling, pre-positioning, and real-time calibration.
[0161] In other embodiments, for traffic patterns, the following steps are also included:
[0162] The S600 is configured with a time series library, which stores full time series data of traffic pattern changes over historical periods.
[0163] First, the time-series sequence of traffic pattern changes is persistently stored. The full time-series data of [traffic pattern - duration - peak characteristics - change path] is recorded at the granularity of hour, minute or second. For example, the complete time-series of [normal browsing - flash sale preheating - peak purchase - decline and stabilization] is formed to create a spatiotemporal feature library of traffic patterns.
[0164] S610 collects the change paths corresponding to several current continuous traffic patterns, and matches them with historical paths in the time series database based on time and content features to output spatiotemporal prediction information, including predicted traffic patterns, predicted peak values, and predicted arrival times.
[0165] Based on the Transformer time series model (lightweight inference version, running entirely within the original feature analysis engine), the model matches the traffic pattern transition paths of the current three consecutive sliding windows with historical time series data in the feature library, outputting the predicted traffic pattern, peak size, and arrival time for the next three time windows: 10s, 30s, and 1 minute. For example, if the current window is identified as "flash sale pre-heating mode," the model predicts that it will enter "peak purchase mode" in 30 seconds, with a peak QPS 12 times the current value.
[0166] The S620 generates and executes a pre-scheduling strategy based on spatiotemporal prediction information, and simultaneously calculates the deviation of the spatiotemporal prediction information based on real-time collected front-end traffic characteristics.
[0167] Based on the prediction results, a pre-scheduling strategy is generated in advance, and the original template matching and health scoring logic is reused.
[0168] Specifically, the pre-scheduling strategies include:
[0169] Resource pre-positioning: To anticipate peak traffic, 70% of the computing power resources of the backup node pool are locked in advance, and the traffic weight of the corresponding nodes is increased in advance to reserve capacity and avoid temporary weight adjustments when the peak arrives.
[0170] Link pre-planning: For anticipated high-concurrency patterns, pre-establish connections to the lowest-latency nodes with the highest weight for HTTP / 2 and WebSocket long connection requests, reducing handshake overhead during peak periods.
[0171] Template preloading: The scheduling template corresponding to the prediction mode is preloaded into the local cache of the intelligent forwarding engine, eliminating the need to request the policy center in real time and reducing decision latency.
[0172] S630 determines whether to continue executing the pre-scheduling strategy based on the result of the deviation calculation.
[0173] Meanwhile, the intelligent forwarding engine executes a pre-scheduling strategy and performs real-time calibration of the prediction results based on real-time collected traffic data: if the actual traffic deviates from the prediction by more than 20%, pre-scheduling is immediately terminated, and the original real-time scheduling logic is switched; if the deviation is within 20%, the pre-scheduling weights are continuously optimized. The configuration management interface supports dynamic configuration of the prediction window, pre-occupancy ratio, and calibration threshold.
[0174] like Figure 2 As shown, this application also discloses a distribution and forwarding server for implementing the above method, specifically including:
[0175] The traffic acquisition module is used to collect multi-dimensional real-time metrics of requests to obtain front-end traffic characteristics and back-end node status.
[0176] It is mounted to the kernel network stack via the eBPF program, collects raw metrics of the TCP / HTTP layer, and supports Netlink to transmit application logs back.
[0177] The feature analysis engine is used to perform sliding window clustering based on front-end traffic features to identify traffic patterns. It is also used to select the corresponding scheduling template based on the traffic pattern and calculate the node health score in combination with the back-end node status. Based on the node health score, it generates and adjusts the dynamic weights for node traffic allocation.
[0178] Strategy Decision Center: Used to store scheduling templates and generate final scheduling policies.
[0179] It stores multiple scheduling templates, dynamically matches templates and generates the final scheduling policy, and supports hot policy updates at runtime.
[0180] The intelligent forwarding engine is used to obtain the protocol and / or content characteristics of the request to select a forwarding strategy and perform differentiated forwarding of the request by combining dynamic weights.
[0181] Implemented using DPDK or Linux TC, it features a lock-free circular buffer queue and a work-stealing thread pool, with a forwarding latency of <1ms and a throughput of ≥2M PPS.
[0182] Configure the management interface and support dynamic updates of scheduling parameters via API.
[0183] Provides a RESTful API that supports dynamic adjustment of clustering window length, health score coefficients (α, β, γ), and circuit breaker threshold.
[0184] It should be understood that although the steps in the flowcharts in the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise expressly stated herein, there is no strict order in which these steps are performed, and they may be performed in other orders.
[0185] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. An adaptive scheduling method based on real-time traffic feature awareness, characterized in that, Includes the following steps: Collect multi-dimensional real-time metrics to obtain front-end traffic characteristics and back-end node status; Sliding window clustering is performed based on the aforementioned front-end traffic characteristics to identify traffic patterns, including normal browsing, flash sales, and web crawler scanning. Based on the traffic pattern, a corresponding scheduling template is selected. Each scheduling template corresponds to a set of weight coefficients and traffic distribution rules, which can be dynamically updated through the configuration management interface to automatically match the optimal template according to the traffic pattern. A node health score is calculated based on the backend node status, and dynamic weights for node traffic allocation are generated and adjusted according to the node health score. The calculation of the node health score includes the following steps: The backend node status includes CPU utilization, memory usage, number of active connections, average response latency, and error rate, and the node health score is calculated based on the above backend node status combined with preset weight coefficients. The node health score is calculated based on the following formula; ; Wherein, α, β, γ are configurable weight coefficients, and their sum is 1. The magnitude of each weight coefficient is determined by the scheduling template corresponding to the traffic pattern. This represents the average response latency of the i-th backend node within the most recent sliding window; This represents the minimum average response latency among all currently available nodes; Indicates the maximum average response latency among all currently available nodes; ϵ is a preset constant; This represents the percentage of CPU usage. This represents the percentage of recent erroneous requests. Obtain the protocol and / or content characteristics corresponding to the request to select a forwarding strategy, and perform differentiated forwarding of the request in combination with the dynamic weight.
2. The adaptive scheduling method based on real-time traffic feature awareness according to claim 1, characterized in that, The front-end traffic characteristics include requests per second, request type distribution, average request size, client IP geographic distribution, and TLS handshake time. Sliding window clustering is performed based on these front-end traffic characteristics to identify traffic patterns, including the following steps: Based on the front-end traffic characteristics within a preset time window, a multi-dimensional feature vector is constructed. Clustering algorithms are used to cluster the multidimensional feature vectors, and the corresponding traffic patterns are identified based on the output results.
3. The adaptive scheduling method based on real-time traffic feature awareness according to claim 2, characterized in that, Based on the traffic pattern, a corresponding scheduling template is selected, including offline evaluation during the training phase, specifically including: Historical traffic data is acquired to replay tests on different scheduling templates. Based on the historical traffic data, the backend node status is used to calculate key indicators, including average response latency reduction rate, CPU utilization standard deviation, and error rate fluctuation range. A comprehensive score is calculated based on the key indicators, and the scheduling template with the highest comprehensive score is selected as the default configuration.
4. The adaptive scheduling method based on real-time traffic feature awareness according to claim 2, characterized in that, Based on the traffic pattern, a corresponding scheduling template is selected, including online feedback during the runtime phase, specifically including: The traffic is split into a small proportion of traffic and a large proportion of traffic, and the small proportion of traffic is allocated to the new scheduling template, while the large proportion of traffic is allocated to the old scheduling template. The measured indicators corresponding to the new scheduling template are continuously collected and compared with the benchmark indicators of the old scheduling template to calculate the template score. The decision to allocate full traffic to the new scheduling template is based on the numerical relationship between the template score and the score threshold.
5. The adaptive scheduling method based on real-time traffic feature awareness according to claim 3 or 4, characterized in that, Obtaining the protocol and / or content characteristics corresponding to the request to select a forwarding strategy, and performing differentiated forwarding of the request in conjunction with the dynamic weight, includes the following steps: By default, the request is forwarded to the node with the highest health score, and the configurable weight coefficient in the scheduling template is forcibly configured based on the goal of the protocol. And / or, The request is routed to a dedicated high-bandwidth egress link node pool, and the bandwidth utilization limit of the high-bandwidth egress link node pool is adjusted. And / or, The source IP of the request is subject to token bucket rate limiting control, and subsequent new requests are directed to a dedicated anti-scraping processing cluster for deep verification.
6. The adaptive scheduling method based on real-time traffic feature awareness according to claim 1, characterized in that, It also includes the following steps: Continuously monitor the status of each backend node under the corresponding traffic mode; When any preset abnormal condition is met within multiple consecutive monitoring windows, the circuit breaker and degradation mechanism is triggered. Executing the circuit breaker and degradation mechanism includes: pausing the forwarding of the request to the original target backend node or node pool, and enabling a backup processing scheme. The backup processing scheme includes switching the paused request to the backup node pool and starting a fast retry strategy or returning to the prompt interface and recording the event log. In each subsequent sliding window, the recovery status of the original target backend node is monitored, and the traffic weight is gradually restored when the recovery benchmark is met.
7. The adaptive scheduling method based on real-time traffic feature awareness according to claim 6, characterized in that, For the aforementioned traffic pattern, the following steps are also included: Configure a time series library, which stores full time series data of each traffic pattern change in historical time. Collect the change path corresponding to the current continuous traffic pattern, and match it with the historical path in the time series database based on time features and content features to output spatiotemporal prediction information, which includes predicted traffic pattern, predicted peak value, and predicted arrival time. A pre-scheduling strategy is generated and executed based on the spatiotemporal prediction information, while deviation calculation is performed on the spatiotemporal prediction information based on the real-time collected front-end traffic characteristics. Based on the result of the deviation calculation, it is determined whether to continue executing the pre-scheduling strategy.
8. The adaptive scheduling method based on real-time traffic feature awareness according to claim 7, characterized in that, The pre-scheduling strategy includes: Resource pre-positioning: Pre-lock a preset proportion of computing power resources in the backup node pool for the predicted peak value and gradually increase the dynamic weight of the corresponding backend node; Link pre-planning: Pre-establish connections to the backend node with the highest dynamic weight; Template preloading: The scheduling template corresponding to the predicted traffic pattern is preloaded and added to the local cache.
9. A distribution and forwarding server, characterized in that, For implementing the method as described in any one of claims 1-8, comprising: The traffic acquisition module is used to collect multi-dimensional real-time metrics of requests to obtain front-end traffic characteristics and back-end node status. The feature analysis engine is used to perform sliding window clustering based on the front-end traffic features to identify traffic patterns. It is also used to select the corresponding scheduling template based on the traffic patterns and calculate the node health score in combination with the back-end node status. Based on the node health score, it generates and adjusts the dynamic weights for node traffic allocation. Strategy Decision Center: Used to store the scheduling templates and generate the final scheduling strategy; The intelligent forwarding engine is used to obtain the protocol and / or content characteristics corresponding to the request to select a forwarding strategy, and to perform differentiated forwarding of the request in combination with the dynamic weight; Configure the management interface and support dynamic updates of scheduling parameters via API.
Citation Information
Patent Citations
Web cluster load balancing algorithm and system based on Nginx dynamic weighting
CN112019620A
Intelligent routing load balancing method and system based on AI
CN121509305A