Elastic resource beforehand early warning and scheduling method and system based on load prediction

By constructing a historical data collector, a predictive analysis engine, an elastic planner, and an intelligent scheduler, and realizing early warning and scheduling of elastic resources based on load prediction, the problems of slow expansion and low resource utilization in cloud resource management are solved, and the system response speed and service quality are improved.

CN120909742APending Publication Date: 2025-11-07INSPUR SOFTWARE TECH CO LTD

Patent Information

Application Number
CN202511453918.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing cloud resource management technologies suffer from problems such as slow expansion, low resource utilization, and lack of business awareness, making them unable to efficiently cope with the cyclical and trend-based changes in business traffic.

Method used

By constructing a historical data collector, a predictive analysis engine, an elastic planner, and an intelligent scheduler, we can achieve early warning and scheduling of elastic resources based on load prediction. This includes collecting multi-dimensional load data, training prediction models, generating load prediction reports, and executing elastic resource pre-expansion plans, combined with intelligent scheduling algorithms to optimize resource utilization.

Benefits of technology

It enables accurate prediction of future business load, eliminates the lag in capacity expansion, improves resource utilization and task scheduling efficiency, and ensures business continuity and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909742A_ABST
    Figure CN120909742A_ABST
Patent Text Reader

Abstract

The invention discloses an elastic resource beforehand early warning and scheduling method and system based on load prediction, and belongs to the technical field of cloud computing resource management and scheduling, and the method comprises the following steps: constructing a historical data collector, and collecting and preprocessing multi-dimensional time sequence load data of a business system in real time; constructing a prediction analysis engine, training a prediction model, accurately predicting future busy and idle time points and period trends of each service system, and generating a quantitative load prediction report; an elastic planner is constructed, and an elastic resource pre-expansion plan is automatically generated and executed before a business peak arrives based on a load prediction report and a preset strategy rule; constructing an intelligent scheduler, and executing an elastic capacity expansion action according to the elastic resource pre-capacity expansion plan; and outputting a complete elastic resource beforehand early warning and scheduling system based on load prediction. According to the invention, the response speed and service quality of the system are remarkably improved, and the method is suitable for resource management requirements in complex environments such as multi-cloud and mixed-cloud environments.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of cloud computing resource management and scheduling, in particular to an elastic resource early warning and scheduling method and system based on load prediction. BACKGROUND

[0002] With the popularity of cloud-native applications, microservice architecture makes business systems increasingly dynamic and complex. The access traffic of the business usually presents obvious peak and valley characteristics, such as diurnal regularity, weekly regularity or sudden traffic caused by specific activities. How to efficiently and economically allocate and manage computing resources (such as CPU, memory, container instances, virtual machines) to cope with such changing loads has become a core challenge of cloud resource management.

[0003] Current mainstream resource scheduling and management technologies, such as Kubernetes HPA or various cloud platform elastic scaling groups, mostly belong to "reactive scheduling", which is to monitor real-time resource indicators (such as CPU utilization, application QPS, etc.), and when the indicators exceed the preset threshold, the system triggers the expansion operation. This method has inherent defects: 1. Expansion lag. From detecting that the load is too high to completing the expansion, it takes a certain amount of time, and during this window period, the system may be in an overloaded state, resulting in increased response delay or even service unavailability.

[0004] 2. Low resource utilization. To avoid lag, a higher resource redundancy is usually set, resulting in resource waste during off-peak periods.

[0005] 3. Lack of business awareness. Based only on underlying resource indicators, it is impossible to understand the periodicity and trend of the business itself, and it is impossible to cope with regular and predictable traffic changes.

[0006] Although some advanced systems attempt to introduce simple prediction, the prediction granularity is often rough and cannot be accurate to specific business systems and specific time points in the future, and the prediction results are disconnected from the elastic action plan, failing to form a "prediction-planning-execution" closed-loop management.

[0007] Therefore, there is an urgent need in the art for a forward-looking resource management solution that can accurately predict future loads based on historical data and proactively schedule resources in advance, thereby fundamentally solving the inherent defects of "reactive scheduling". SUMMARY

[0008] The technical task of the present application is to provide a load prediction-based proactive elastic resource early warning and scheduling method and system to solve the problems of capacity expansion lag and low resource utilization in traditional reactive scheduling, significantly improve system response speed and service quality, and be suitable for resource management requirements in complex environments such as multi-cloud and hybrid cloud.

[0009] The technical solution adopted by the present application to solve its technical problems is: A load prediction-based proactive elastic resource early warning and scheduling method, the implementation of the method includes the following steps: Step 1, build a historical data collector (HDC) to collect multi-dimensional time series load data of the business system in real time and perform preprocessing; Step 2, build a predictive analytics engine (PAE) to train a prediction model, analyze historical data, accurately predict future busy time points and periodic trends of each business system, and generate a quantitative load prediction report; Step 3, build an elastic planner dispatcher (EPD) based on the load prediction report and pre-set strategy rules to automatically generate and execute elastic resource pre-expansion plans before the arrival of business peaks, realizing the transition from passive response to active planning; Step 4, build an intelligent scheduler dispatcher (ISD) to execute elastic expansion actions according to the elastic resource pre-expansion plan, combine load balancing and task priority algorithms, and realize fine scheduling and efficient utilization of resources; Step 5, output the complete load prediction-based proactive elastic resource early warning and scheduling system (PEROS).

[0010] The method realizes elastic resource allocation based on machine learning prediction models, predicts future business demand by analyzing historical load data, and realizes proactive early warning and automatic elastic expansion of resources. By building a historical data collector (HDC), a predictive analytics engine (PAE), an elastic planner (EPD), and an intelligent scheduler (ISD), the future trend of business load is accurately predicted, and resource expansion is automatically completed before the arrival of traffic peaks, fundamentally eliminating the inherent time lag of traditional reactive expansion and contraction, effectively avoiding service performance jitter and interruption risks, and ensuring business continuity and user experience.

[0011] Further, the step 1 is specifically implemented as follows: By constructing a historical data collector (HDC), historical load data of each business system is continuously collected at a fixed frequency (for example, once per minute), and the collection content includes multiple dimensions of performance indicators such as CPU usage, memory occupation, network I / O, disk I / O, request QPS, response delay, error rate, etc. Firstly, the number of nodes of the data collection cluster is calculated by the predicted peak data inflow rate, safety factor, and maximum sustainable processing throughput of a single node, and the calculation formula is: The number of nodes of the data collection cluster (nums) = CEILING ((predicted peak data inflow rate x safety factor) / maximum sustainable processing throughput of a single node); Then, the collected raw data is processed through a data cleaning process, including removing outliers (such as negative values or extreme values beyond a reasonable range), filling missing values (using linear interpolation or forward-backward value filling method), and performing timestamp alignment processing to ensure that all indicators are aligned at the same time granularity. The processed data is classified by business system identifier and indicator type and stored in a high-performance time series database to form a structured historical data set, providing high-quality input for subsequent prediction analysis.

[0012] Further, the step 2 is specifically implemented as follows: By constructing a prediction analysis engine (PAE), a prediction model is independently constructed and trained for each business system; different models are dynamically adapted according to the business load characteristics: For businesses with obvious periodicity (such as daily and weekly cycles), a three-parameter exponential smoothing algorithm is preferred, which optimizes the level (Level), trend (Trend), and seasonality (Seasonality) parameters to accurately capture periodic fluctuations. For non-periodic or obviously trended loads, ARIMA (Autoregressive Integrated Moving Average Model) or Prophet algorithm is used for modeling; input the latest historical data to predict the load value every 15 minutes within the future time period (such as 6 hours), and generate a structured prediction report; the report content includes prediction time interval, expected load peak and valley, confidence interval, and system busy state marker (such as "peak period" or "low peak period"), and is output in JSON or XML format for use by downstream modules.

[0013] Further, the ARIMA (autoregressive integrated moving average model) or Prophet algorithm is used for modeling. During the model training process, the first 80% of the historical data set is used as the training set, and the last 20% is used as the validation set. The model parameters are optimized by minimizing the root mean square error or the average absolute percentage error. After the training is completed, the engine automatically runs at a preset period (such as every hour).

[0014] Further, step 3, build an elastic planner (EPD) to receive the load prediction report output by the prediction analysis engine and match it with the built-in strategy rule library; The strategy rule library uses a configurable YAML or JSON format, allowing users to customize rules based on business SLA (service level agreement) and cost constraints. For example, the rule can be defined as: "If the predicted CPU load exceeds 80% in the next 2 hours, expand the container instance number to 150% of the current number 30 minutes in advance" or "If the predicted load is less than 30% in the next 3 hours, automatically scale down to 50% of the instance number to save resources"; after a successful match, the elastic planner generates a specific elastic resource pre-expansion plan, including the target business system, planned execution time, resource type (such as CPU, memory, instance number), target number, operation type (expansion / contraction), etc., and persists the plan to the database while adding it to the pending task queue.

[0015] Further, the elastic planner matches the load prediction report with its built-in strategy rule library. Rule matching is based on fuzzy logic or threshold judgment, supporting multi-condition combination (such as meeting CPU and memory conditions simultaneously).

[0016] Further, step 4 is implemented as follows: Build an intelligent scheduler (ISD) to interact with the underlying resource management platform by integrating Kubernetes API, cloud service provider SDK, or infrastructure orchestration tools; Before the plan is executed, the system will perform a pre-check, including resource quota verification, network connectivity testing, and permission authentication. When the plan execution time is reached (such as 30 minutes in advance), the elastic planner calls the corresponding API interface to perform the expansion operation, such as calling the Kubernetes API to modify the replicas field of Deployment or adjusting the desired instance number of the scaling group through the cloud platform interface. During the execution process, the system monitors the operation status in real time, and if the execution fails, it triggers an alarm and attempts to retry or rollback. After successful execution, the resource state is updated to the meta database, and the intelligent scheduler is notified that the resource is ready; When the predicted business peak time is coming, the intelligent scheduler will evenly distribute the incoming requests to all nodes including the newly added nodes, and the system will smoothly pass through the peak; afterwards, the system will automatically optimize the prediction model parameters by comparing the predicted load with the actual load, realizing closed-loop learning.

[0017] The application also claims to protect an elastic resource early warning and scheduling system based on load prediction, comprising: A historical data collector (HDC) is used to continuously collect historical load data (such as CPU usage, memory usage, request QPS, response delay, etc.) of each business system, forming a high-quality time series data set; A predictive analytics engine (PAE) is the "brain" of the system, which is used to connect the data collector, independently model and train the load data of each business system by using exponential smoothing algorithm (such as three-parameter exponential smoothing) and other time series models (such as ARIMA, Prophet), so as to predict the load level of the business system at a specific time point or time period in the future, and generate a structured load prediction report; the report clearly indicates the "busy time" and "idle time" in the future and the expected load level; An elastic planner dispatcher (EPD) is used to receive the load prediction report output by the predictive analytics engine, and automatically match the preset strategy, which is innovative in that it contains a strategy rule base, which defines different elastic actions under different prediction scenarios (for example, "predicting that the CPU load will exceed 80% in the next hour" -> "triggering the expansion action to increase the number of container instances to 20"); the elastic planner dispatcher generates a specific elastic resource pre-expansion plan according to the prediction report and matches the rules before the business peak arrives (such as 30 minutes in advance), and issues it to the execution layer; An intelligent scheduler dispatcher (ISD) is the "heart" of the system, which supports a dual-mode scheduling mechanism of "normal scheduling" and "post-expansion scheduling"; among them, "normal scheduling" is to use algorithms including load balancing, task priority, etc. to perform efficient task allocation; "post-expansion scheduling" refers to the fact that after the elastic planner dispatcher pre-expands the resources, the intelligent scheduler dispatcher can perceive the newly added resource nodes and automatically schedule new tasks to the new nodes, ensuring that the pre-expanded resources can be immediately fully utilized and avoiding resource idling; The system realizes elastic resource early warning and scheduling through the above method.

[0018] The application also claims a load prediction-based elastic resource early warning and scheduling device, comprising at least one memory and at least one processor. The at least one memory is used for storing machine readable programs. The at least one processor is used for calling the machine readable programs to realize the above method.

[0019] The application also claims a computer readable medium, which stores computer instructions, and the computer instructions can realize the above method when executed by a processor.

[0020] Compared with the prior art, the load prediction-based elastic resource early warning and scheduling method and system have the following beneficial effects: 1. Precise business load prediction is realized. By analyzing historical load data of each business system, a time series prediction algorithm is used to accurately predict future busy time points and periodic trends, thereby providing a reliable basis for resource planning.

[0021] 2. An early warning mechanism is established. The trigger condition of resource management is changed from a real-time index exceeding a threshold to a predicted index exceeding a threshold, and the response is changed from an in-process response to an early warning, thereby fundamentally eliminating the expansion lag.

[0022] 3. Resource elasticity is pre-expanded. According to the prediction result, a resource pre-expansion plan is automatically generated and executed before the arrival of a business peak, thereby ensuring that the system is ready before the arrival of a traffic peak.

[0023] 4. Resource utilization and task scheduling efficiency are improved. Under the premise of ensuring business performance, unnecessary resource redundancy is reduced and overall resource utilization is improved through accurate prediction and planning. In combination with a high-level scheduling algorithm, task execution efficiency is optimized. BRIEF DESCRIPTION OF DRAWINGS

[0024] Figure 1 It is an embodiment of the application to provide a load prediction-based elastic resource early warning and scheduling method architecture diagram. DETAILED DESCRIPTION

[0025] The application will be further described below in combination with specific embodiments.

[0026] The application embodiment provides a load prediction-based elastic resource early warning and scheduling method, and the implementation of the method includes the following steps: Step 1, a historical data collector (HDC) is constructed to collect multi-dimensional time series load data of a business system in real time and perform preprocessing; Step 2, build a predictive analytics engine (PAE), train a prediction model, analyze historical data, accurately predict the busy time points and periodic trends of each business system in the future, and generate a quantitative load prediction report; Step 3, build an elastic planner dispatcher (EPD), based on the load prediction report and pre-set strategy rules, automatically generate and execute an elastic resource pre-expansion plan before the arrival of business peak, realize the transition from passive response to active planning; Step 4, build an intelligent scheduler dispatcher (ISD), according to the elastic resource pre-expansion plan, execute the elastic expansion action, combine load balancing and task priority algorithm, realize fine scheduling and efficient utilization of resources; Step 5, output the complete proactive elastic resource orchestration system (PEROS) based on load prediction.

[0027] This method realizes elastic resource allocation based on machine learning prediction model, predicts future business demand by analyzing historical load data, and realizes proactive warning and automatic elastic expansion of resources. By building a historical data collector (HDC), a predictive analytics engine (PAE), an elastic planner (EPD), and an intelligent scheduler (ISD), the future trend of business load is accurately predicted, and resource expansion is automatically completed before the arrival of traffic peak, which fundamentally eliminates the inherent time lag of traditional reactive expansion and contraction, effectively avoids service performance jitter and interruption risk, and guarantees business continuity and user experience.

[0028] In combination with the accompanying Figure 1 The specific implementation process of the method is as follows: 1, build a historical data collector (HDC) to collect and preprocess historical data.

[0029] By building a historical data collector (HDC), historical load data of each business system is continuously collected at a fixed frequency (e.g., once per minute), including CPU usage, memory usage, network I / O, disk I / O, request QPS, response delay, error rate, and other multi-dimensional performance indicators. First, the number of nodes of the data collection cluster is calculated by the expected peak data inflow rate, safety factor, and maximum sustainable processing throughput of a single node, and the formula is: nums = CEILING ((expected peak data inflow rate x safety factor) / maximum sustainable processing throughput of a single node). Then, the collected raw data is processed through data cleaning processes, including removing outliers (such as negative values or extreme values outside the reasonable range), filling missing values (using linear interpolation or forward-backward value filling method), and timestamp alignment processing to ensure that all indicators are aligned at the same time granularity. The processed data is classified by business system identifier and indicator type and stored in a high-performance time series database to form a structured historical data set, providing high-quality input for subsequent prediction analysis.

[0030] 2. Build a prediction analysis engine (PAE) to train prediction models and generate load prediction reports.

[0031] By building a prediction analysis engine (PAE), a prediction model is built and trained for each business system independently. The engine can dynamically adapt different models based on business load characteristics: For businesses with obvious periodicity (such as daily and weekly cycles), the three-parameter exponential smoothing algorithm is preferred, which optimizes the level (Level), trend (Trend), and seasonality (Seasonality) parameters to accurately capture periodic fluctuations.

[0032] For non-periodic or trend-based loads, ARIMA (Autoregressive Integrated Moving Average Model) or Prophet algorithm is used for modeling. During model training, the first 80% of the historical data set is used as the training set and the last 20% is used as the validation set, and the model parameters are optimized by minimizing the root mean square error or average absolute percentage error. After training, the engine automatically runs at a pre-set period (e.g., every hour), inputs the latest historical data, predicts the load value every 15 minutes for a certain period (e.g., 6 hours), and generates a structured prediction report. The report includes the prediction time interval, expected load peak and valley, confidence interval, and system busy state markers (such as "peak period" or "low peak period"), and is output in JSON or XML format for use by downstream modules.

[0033] 3. Build an elastic planner (EPD) to develop an elastic resource pre-expansion plan based on the load prediction report.

[0034] By constructing an elastic planner (EPD), the load prediction report output by the prediction analysis engine is received and matched with the built-in policy rule library.

[0035] The policy rule library adopts a configurable YAML or JSON format, allowing users to customize rules based on business SLA (service level agreement) and cost constraints. For example, a rule can be defined as: "If the predicted CPU load exceeds 80% in the next 2 hours, scale out the container instances to 150% of the current number 30 minutes in advance" or "If the predicted load is less than 30% in the next 3 hours, automatically scale down to 50% of the instance number to save resources." Rule matching is based on fuzzy logic or threshold judgment, supporting multi-condition combination (such as meeting CPU and memory conditions simultaneously). After a successful match, the elastic planner generates a specific elastic resource pre-scaling plan, including target business system, planned execution time, resource type (such as CPU, memory, instance number), target number, operation type (scaling up / down), etc., and persists the plan to the database while adding it to the pending task queue.

[0036] 4、Construct an intelligent scheduler (ISD) to execute elastic scaling actions based on the elastic resource pre-scaling plan.

[0037] Build an intelligent scheduler (ISD) to interact with the underlying resource management platform by integrating Kubernetes API, cloud service provider SDK, or infrastructure orchestration tools.

[0038] Before the plan is executed, the system will perform pre-checks, including resource quota verification, network connectivity testing, and permission authentication. When the planned execution time arrives (e.g., 30 minutes in advance), the elastic planner calls the corresponding API interface to perform scaling operations, such as calling the Kubernetes API to modify the replicas field of Deployment or adjusting the desired instance number of the scaling group through the cloud platform interface. During execution, the system monitors the operation status in real time, and if the execution fails, it triggers an alarm and attempts to retry or rollback. After successful execution, the resource state is updated to the meta database, and the intelligent scheduler is notified that the resources are ready. When the predicted business peak time point is approaching, the intelligent scheduler will evenly distribute incoming requests to all nodes, including new nodes, and the system will smoothly pass through the peak. Afterward, the system can compare the predicted load with the actual load, automatically optimize the prediction model parameters, and achieve closed-loop learning.

[0039] 5、Output a complete load prediction-based elastic resource pre-warning and scheduling system (PEROS).

[0040] So far, through the above steps, combined with machine learning and large model technologies, an elastic resource early warning and scheduling system based on load prediction is constructed, which can accurately predict the future trend of business load and automatically complete resource expansion before the traffic peak arrives, fundamentally eliminating the inherent time lag of traditional reactive expansion and contraction, effectively avoiding service performance jitter and interruption risk, and ensuring business continuity and user experience. At the same time, through accurate prediction and on-demand pre-expansion, the system significantly reduces the resource redundancy commonly prepared to cope with sudden traffic, greatly improves resource utilization, and reduces operating costs.

[0041] The embodiment of the application also provides an elastic resource early warning and scheduling system based on load prediction, which comprises four core modules: A historical data collector (HDC) is a "sensory system" of the system, responsible for continuously collecting historical load data (such as CPU usage, memory usage, request QPS, response delay, etc.) of each business system, forming a high-quality time series data set.

[0042] A predictive analytics engine (PAE) is the "brain" of the system, responsible for connecting the data collector, modeling and training the load data of each business system independently by using exponential smoothing algorithms (such as three-parameter exponential smoothing) and other time series models (such as ARIMA, Prophet), predicting the load level of the business system at a specific time point or time period in the future, and generating a structured load prediction report. The report clearly indicates the "busy time" and "idle time" in the future and the expected load level.

[0043] An elastic planner dispatcher (EPD) is the "brainstem" of the system, responsible for receiving the load prediction report output by the predictive analytics engine and automatically matching the preset strategy. Its innovation lies in containing a strategy rule base, which defines elastic actions under different prediction scenarios (for example, "predicting that the CPU load will exceed 80% for the next one hour" -> "triggering the expansion action to increase the number of container instances to 20"). The elastic planner dispatcher matches the rules according to the prediction report and generates a specific elastic resource pre-expansion plan before the business peak arrives (such as 30 minutes in advance), and issues it to the execution layer.

[0044] Intelligent scheduler (ISD). It is the "heart" of the system, supporting the dual-mode scheduling mechanism of "normal scheduling" and "post-expansion scheduling". Among them, "normal scheduling" is to use load balancing, task priority and other algorithms for efficient task allocation; "post-expansion scheduling" refers to the fact that after the elastic planner expands resources, the intelligent scheduler can perceive the newly added resource nodes and automatically schedule new tasks to the new nodes first, ensuring that the expanded resources can be fully utilized immediately, avoiding resource idling.

[0045] The system realizes the pre-warning and scheduling of elastic resources through the load prediction-based pre-warning and scheduling method described in the above embodiments.

[0046] 1. The historical data collector (HDC) collects and pre-processes historical data.

[0047] By constructing the historical data collector (HDC), the historical load data of each business system is continuously collected at a fixed frequency (e.g. once per minute), and the collection content includes multi-dimensional performance indicators such as CPU usage, memory occupancy, network I / O, disk I / O, request QPS, response delay, error rate, etc. First, the number of nodes of the data collection cluster is calculated by using parameters such as predicted peak data inflow rate, safety factor, and maximum sustainable processing throughput per node, and the calculation formula is: the number of nodes of the data collection cluster (nums) = CEILING ((predicted peak data inflow rate x safety factor) / maximum sustainable processing throughput per node). Then, the collected raw data is processed through a data cleaning process, including removing outliers (such as negative values or extreme values outside the reasonable range), filling missing values (using linear interpolation or forward-backward value filling method), and performing timestamp alignment processing to ensure that all indicators are aligned at the same time granularity. The processed data is classified by business system identifier and indicator type and stored in a high-performance time series database to form a structured historical data set, providing high-quality input for subsequent prediction analysis.

[0048] 2. The prediction analysis engine (PAE) trains the prediction model and generates a load prediction report.

[0049] By constructing the prediction analysis engine (PAE), a prediction model is independently constructed and trained for each business system. The engine can dynamically adapt different models according to the characteristics of business load: For businesses with obvious periodicity (such as daily and weekly cycles), the three-parameter exponential smoothing algorithm is preferred, which optimizes the level (Level), trend (Trend), and seasonality (Seasonality) parameters to accurately capture periodic fluctuations.

[0050] For non-periodic or trend-based loads, ARIMA (AutoRegressive Integrated Moving Average) or Prophet algorithms are used for modeling. During model training, the first 80% of the historical dataset is used as the training set, and the last 20% is used as the validation set. Model parameters are optimized by minimizing the root mean square error or the mean absolute percentage error. After training, the engine automatically runs at a preset interval (e.g., every hour), inputs the latest historical data, predicts the load value at each 15-minute interval for a certain period (e.g., 6 hours), and generates a structured prediction report. The report includes the prediction time interval, expected load peak and valley, confidence interval, and system busy state label (e.g., "peak period" or "low period"), and is output in JSON or XML format for downstream modules.

[0051] 3. The elastic planner (EPD) formulates an elastic resource pre-scaling plan based on the load prediction report.

[0052] An elastic planner (EPD) is constructed to receive the load prediction report output by the prediction analysis engine and match it with its built-in policy rule library.

[0053] The policy rule library is in a configurable YAML or JSON format, allowing users to customize rules based on business SLA (Service Level Agreement) and cost constraints. For example, a rule can be defined as: "If the predicted CPU load exceeds 80% for the next 2 hours, scale up the container instance number to 150% of the current number 30 minutes in advance" or "If the predicted load is less than 30% for the next 3 hours, automatically scale down to 50% of the instance number to save resources." Rule matching is based on fuzzy logic or threshold judgment, supporting multiple condition combinations (such as CPU and memory conditions). After successful matching, the elastic planner generates a specific elastic resource pre-scaling plan, including target business system, planned execution time, resource type (such as CPU, memory, instance number), target number, operation type (scaling up / down), etc., and persists the plan to the database while adding it to the pending task queue.

[0054] 4. The intelligent scheduler (ISD) executes elastic scaling actions according to the elastic resource pre-scaling plan.

[0055] An intelligent scheduler (ISD) is constructed to interact with the underlying resource management platform by integrating Kubernetes API, cloud service provider SDK, or infrastructure orchestration tools.

[0056] Before the plan is executed, the system will perform pre-checking, including resource quota verification, network connectivity testing and permission authentication. When the plan execution time arrives (e.g. 30 minutes in advance), the elastic planner calls the corresponding API interface to perform the expansion operation, for example: calling the Kubernetes API to modify the replicas field of Deployment, or adjusting the desired instance number of the scaling group through the cloud platform interface. During the execution process, the system monitors the operation state in real time, and if the execution fails, it triggers an alarm and attempts to retry or rollback. After the execution is successful, the resource state is updated to the meta database, and the intelligent scheduler is notified that the resource is ready. When the predicted business peak time point is approaching, the intelligent scheduler evenly distributes the incoming requests to all nodes including the newly added nodes, and the system smoothly passes through the peak. After that, the system can compare the predicted load with the actual load, automatically optimize the prediction model parameters, and realize closed-loop learning.

[0057] The system collects multi-dimensional time series load data of the business system in real time through the historical data collector; the prediction analysis engine analyzes the historical data using algorithms such as exponential smoothing, accurately predicts the busy and idle time points and periodic trends of each business system in the future, and generates a quantitative load prediction report; the elastic planner automatically generates and executes an elastic resource pre-expansion plan before the business peak arrives according to the prediction results and pre-set strategy rules, realizing the transition from passive response to active planning; the intelligent scheduler combines load balancing and task priority algorithms to realize fine scheduling and efficient utilization of resources. The "prediction-planning-scheduling" integrated architecture is innovatively constructed, through pre-warning and elastic pre-expansion mechanism, the problems of expansion lag and low resource utilization in traditional reactive scheduling are completely solved, the system response speed and service quality are significantly improved, and the resource management demand in complex environments such as multi-cloud and hybrid cloud is suitable.

[0058] The embodiment of the application also provides an elastic resource pre-warning and scheduling device based on load prediction, which comprises at least one memory and at least one processor. The at least one memory is used for storing a machine readable program. The at least one processor is used for calling the machine readable program to realize the method of the elastic resource pre-warning and scheduling device based on load prediction described in the above embodiment.

[0059] The embodiment of the present application also provides a computer readable medium, wherein the computer readable medium stores computer instructions, and the computer instructions are executed by a processor to realize the method for pre-warning and scheduling of elastic resources based on load prediction. Specifically, a system or device provided with a storage medium can be provided, wherein the storage medium stores software program codes for realizing the functions of any one of the above embodiments, and a computer (or CPU or MPU) of the system or device reads and executes the program codes stored in the storage medium.

[0060] In this case, the program codes read from the storage medium can realize the functions of any one of the above embodiments, and thus the program codes and the storage medium storing the program codes constitute a part of the present application.

[0061] The storage medium for storing the program codes includes a floppy disk, a hard disk, a magneto-optical disk (e.g., CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a nonvolatile memory card, and a ROM. Alternatively, the program codes can be downloaded from a server computer through a communication network.

[0062] In addition, it should be understood that not only the program codes read by the computer can be executed, but also the operating system or the like operating on the computer can be caused to perform part or all of the actual operations based on the instructions of the program codes, so as to realize the functions of any one of the above embodiments.

[0063] In addition, it should be understood that the program codes read from the storage medium can be written into the memory provided in the expansion board inserted into the computer or the memory provided in the expansion unit connected to the computer, and then part or all of the actual operations can be performed based on the instructions of the program codes by the CPU or the like installed on the expansion board or the expansion unit, so as to realize the functions of any one of the above embodiments.

[0064] The present application has been described in detail by the above drawings and preferred embodiments, however, the present application is not limited to the disclosed embodiments, and those skilled in the art can know that the code auditing means in the above different embodiments can be combined to obtain more embodiments of the present application, and these embodiments are also within the protection scope of the present application.

Claims

1. A method for early warning and scheduling of elastic resources based on load prediction, characterized in that, The implementation of the method comprises the following steps: Step 1, build a historical data collector to collect multi-dimensional time series load data of the business system in real time and perform preprocessing; Step 2, build a prediction analysis engine, train the prediction model, analyze the historical data, accurately predict the busy time points and periodic trends of each business system in the future, and generate a quantitative load prediction report; Step 3, build an elastic planner, based on the load prediction report and pre-set strategy rules, automatically generate and execute the elastic resource pre-expansion plan before the arrival of the business peak; Step 4, build an intelligent scheduler, execute the elastic expansion action according to the elastic resource pre-expansion plan, combine load balancing and task priority algorithm to realize fine scheduling and efficient utilization of resources; Step 5, output the complete elastic resource early warning and scheduling system based on load prediction. 2.The load prediction based elastic resource early warning and scheduling method of claim 1, wherein, The step 1 is specifically implemented as follows: By building a historical data collector, the historical load data of each business system is continuously collected at a fixed frequency, and the collection content includes CPU usage, memory occupation, network I / O, disk I / O, request QPS, response delay, and error rate multi-dimensional performance indicators; Firstly, the number of nodes of the data collection cluster is calculated by predicting parameters including peak data inflow rate, safety factor, and maximum sustainable processing throughput per node, and the calculation formula is: Number of nodes of data collection cluster = CEILING ((predicted peak data inflow rate x safety factor) / maximum sustainable processing throughput per node); Then, the collected raw data is processed through a data cleaning process, including removing outliers, filling missing values, and performing timestamp alignment processing to ensure that all indicators are aligned at the same time granularity; The processed data is classified by business system identifier and index type and stored in a high-performance time series database to form a structured historical data set. 3.The load prediction based elastic resource early warning and scheduling method of claim 1, wherein, The step 2 is specifically implemented as follows: By building a prediction analysis engine, a prediction model is independently built and trained for each business system; different models are dynamically adapted according to business load characteristics: For businesses with obvious periodicity, the three-parameter exponential smoothing algorithm is preferred, which optimizes the level, trend, and seasonality parameters to accurately capture periodic fluctuations; For non-periodic or obviously trended loads, ARIMA or Prophet algorithm is used for modeling; Input the latest historical data to predict the load value every 15 minutes in the future time period and generate a structured prediction report; the report content includes prediction time interval, expected load peak and valley, confidence interval, and system busy state marker, and is output in JSON or XML format for use by downstream modules.

4. The method of claim 3, wherein, In the model training process, the first 80% of the historical data set is used as the training set and the last 20% is used as the validation set, and the model parameters are optimized by minimizing the root mean square error or mean absolute percentage error; After training, the engine automatically runs at a preset period.

5. The method of claim 1, wherein, The step 3, the elastic planner is used to receive the load prediction report output by the prediction analysis engine and match it with its built-in strategy rule library; The policy rule base adopts a configurable YAML or JSON format, allowing users to customize rules according to business SLA and cost constraints; after a successful match, the elastic planner generates a specific elastic resource pre-scaling plan, including the target business system, plan execution time, resource type, target quantity, operation type, and persists the plan to the database while adding it to the to-be-executed task queue.

6. The method of claim 5, wherein, The elastic planner matches the load prediction report with its built-in policy rule base, and the rule matching is based on fuzzy logic or threshold judgment, supporting multi-condition combination.

7. The method of claim 1, wherein, The step 4 is specifically implemented as follows: The intelligent scheduler interacts with the underlying resource management platform by integrating Kubernetes API, cloud service provider SDK, or infrastructure orchestration tools; Before plan execution, the system performs pre-checking, including resource quota verification, network connectivity testing, and permission authentication; At the plan execution time, the elastic planner calls the corresponding API interface to execute the scaling operation; during the execution process, the system monitors the operation state in real time, triggers an alarm and attempts to retry or rollback if the execution fails; after successful execution, the system updates the resource state to the meta-database and notifies the intelligent scheduler that the resource is ready; When the predicted business peak time is approaching, the intelligent scheduler evenly distributes incoming requests to all nodes, including the newly added nodes, and the system smoothly passes through the peak; after that, the system automatically optimizes the prediction model parameters by comparing the predicted load with the actual load, realizing closed-loop learning.

8. A load prediction based elastic resource early warning and scheduling system, characterized in that, It includes: a historical data collector for continuously collecting historical load data of each business system to form a time series data set; a prediction analysis engine for connecting the data collector, independently modeling and training the load data of each business system by using an exponential smoothing algorithm and other time series models, thereby predicting the load level of the business system at a specific time point or time period in the future, and generating a structured load prediction report; the report explicitly indicates the busy time and idle time in the future and the expected load magnitude; an elastic planner for receiving the load prediction report output by the prediction analysis engine, automatically matching the preset strategy, including a policy rule base that defines elastic actions under different prediction scenarios; the elastic planner generates a specific elastic resource pre-scaling plan according to the prediction report matching rules before the business peak arrives, and delivers it to the execution layer; an intelligent scheduler supporting a dual-mode scheduling mechanism of normal scheduling and post-scaling scheduling; among them, the normal scheduling is to use algorithms including load balancing and task priority for efficient task allocation; the post-scaling scheduling refers to the intelligent scheduler being able to perceive the newly added resource nodes after the elastic planner pre-scales the resources, and automatically scheduling new tasks to the new nodes, ensuring that the pre-scaled resources can be immediately fully utilized to avoid resource idling; The system realizes elastic resource pre-warning and scheduling through the method of any one of claims 1 to 7.

9. A load prediction based elastic resource early warning and scheduling device, characterized in that, It includes: at least one memory and at least one processor; the at least one memory is used to store machine-readable programs; The at least one processor is configured to invoke the machine readable program to implement the method of any one of claims 1 to 7.

10. A computer readable medium characterized by The computer readable medium stores computer instructions, and the computer instructions, when executed by a processor, can implement the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Self-adaptive control strategy based on pattern prediction

    CN108241526A

  • Shunting processing method and system, program product and readable storage medium

    CN118945115A

  • Container cloud elastic expansion and contraction method based on load prediction

    CN120216096A

  • Self-adaptive cloud management platform system based on intelligent resource scheduling and container arrangement

    CN120429116A

  • Dynamic resource allocation method and system under micro-service architecture

    CN120704900A

Cited By

  • Rapid containerization packaging method, system and device for generative AI application and medium

    CN121116500A

  • Desktop cloud AI dynamic resource scheduling method for multiple office loads

    CN121523904A

  • Elastic load balancing scheduling method adaptive to multiple nodes

    CN121691330A

  • Prejudgment method and device for subscription of cloud native database and medium

    CN122086713A