Resource allocation method and device of server
By building a cloud computing enhanced environment and utilizing hybrid load models and intelligent algorithms, the problem of low testing efficiency in cloud computing resource scheduling is solved, efficient dynamic resource allocation and anomaly detection are achieved, and the system stability and resource utilization are improved.
Patent Information
- Application Number
- CN202511232673.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-29
AI Technical Summary
Existing software testing for cloud computing resource scheduling is inefficient, time- and labor-intensive, and overly dependent on manual intervention, resulting in large errors and low accuracy.
By building a cloud computing enhanced environment, applying traffic and periodic fluctuations that meet preset requirements, using hybrid load models and fault injection technology to obtain test error values and key indicator values, and combining the Transformer converter model and multi-objective optimization algorithm, load prediction and scheduling strategy efficiency tests are carried out to achieve dynamic resource allocation.
It significantly shortens the anomaly discovery time and the average fault repair time, improves resource utilization and anomaly detection coverage, solves the limitations of a single tool in high-concurrency and complex scenarios, and improves testing efficiency and accuracy.
Smart Images

Figure CN120723484A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of cloud computing resource scheduling, and in particular to a method and apparatus for allocating server resources. Background Art
[0002] Cloud computing resource scheduling is a dynamic resource management technology based on intelligent algorithms. This technology can automatically allocate physical and virtualized resources (such as processors, memory, networks, etc.) by monitoring system load in real time, thereby optimizing resource utilization while ensuring service quality. For example, cloud computing resource scheduling technology can use the container orchestration platform to automatically increase computing resources during business peaks to ensure service stability; it can automatically recycle idle resources during business slack to reduce operating costs, thereby achieving fine-grained microservice resource scheduling, while optimizing communication efficiency between distributed nodes to ensure high-performance operation of the overall system.
[0003] To ensure the stability of dynamic resource scheduling, relevant software testing methods can be carried out by test engineers after completing the test case design and passing the expert review, strictly following the preset operating steps to perform the test, so as to compare the actual operation results of the system with the expected goals, thereby verifying the reliability of the software test.
[0004] However, the related software testing methods have low testing efficiency, high time and labor costs, and are overly dependent on manual intervention, resulting in large errors and low accuracy, which urgently need to be solved. Summary of the Invention
[0005] The present application provides a server resource allocation method and device to at least solve the technical problems in related technologies, namely, low software testing efficiency, high time and labor costs, and excessive reliance on manual intervention, resulting in large errors and low accuracy.
[0006] The present application provides a resource allocation method for a server, comprising the following steps: applying traffic and periodic fluctuations that meet preset requirements to a pre-built cloud computing basic environment to establish a corresponding cloud computing enhanced environment; obtaining a test error value and at least one key indicator value of the cloud computing enhanced environment, and judging whether the test error value and the at least one key indicator value meet corresponding preset verification requirements, so as to obtain corresponding load prediction test data and scheduling strategy efficiency test data when the test error value and the at least one key indicator value meet the preset verification requirements; performing elastic scaling tests and regional failure scenario simulations on the cloud computing enhanced environment to generate corresponding elastic scaling test data and cross-regional disaster recovery test data, so as to dynamically allocate resources for the target server based on the load prediction test data, the scheduling strategy efficiency test data, the elastic scaling test data and the cross-regional disaster recovery test data.
[0007] The present application also provides a resource allocation device for a server, including: an environment construction module, used to impose traffic and periodic fluctuations that meet preset requirements on a pre-built cloud computing basic environment to establish a corresponding cloud computing enhanced environment; a testing module, used to obtain a test error value and at least one key indicator value of the cloud computing enhanced environment, and determine whether the test error value and the at least one key indicator value meet the corresponding preset verification requirements, so as to obtain corresponding load prediction test data and scheduling strategy efficiency test data when the test error value and the at least one key indicator value meet the preset verification requirements; a resource allocation module, used to perform elastic scaling tests and regional failure scenario simulations on the cloud computing enhanced environment to generate corresponding elastic scaling test data and cross-regional disaster recovery test data, so as to dynamically allocate resources for the target server according to the load prediction test data, the scheduling strategy efficiency test data, the elastic scaling test data and the cross-regional disaster recovery test data.
[0008] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned server resource allocation methods when executing the computer program.
[0009] The present application also provides a non-volatile computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned server resource allocation methods are implemented.
[0010] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned server resource allocation methods when the computer program is executed by a processor.
[0011] Through this application, the traffic and periodic fluctuations that meet the preset requirements can be imposed on the pre-built cloud computing basic environment to establish a corresponding cloud computing enhanced environment; the test error value and at least one key indicator value of the cloud computing enhanced environment can be obtained, and it can be judged whether the test error value and at least one key indicator value meet the corresponding preset verification requirements, so as to obtain the corresponding load prediction test data and scheduling strategy efficiency test data when the test error value and at least one key indicator value meet the preset verification requirements; elastic scaling test and regional fault scenario simulation can be performed on the cloud computing enhanced environment to generate corresponding elastic scaling test data and cross-regional disaster recovery test data, so as to Load prediction test data, scheduling strategy efficiency test data, elastic scaling test data, and cross-regional disaster recovery test data are used to dynamically allocate resources for the target server. Therefore, it can solve the technical problems in related technologies, such as low software testing efficiency, high time and labor costs, and excessive reliance on manual intervention, resulting in large errors and low accuracy. By utilizing artificial intelligence prediction models, the time to discover anomalies and the average repair time of faults are greatly shortened, and a division of labor and collaborative testing strategy of stepped pressure and behavioral simulation is adopted to solve the limitations of a single tool in high-concurrency and complex scenarios, and improve resource utilization and anomaly detection coverage. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0013] Figure 1 A flowchart of a server resource allocation method provided according to an embodiment of the present application; Figure 2 A schematic diagram of the process of building a Locust and JMeter hybrid load model provided in one embodiment of the present application; Figure 3 A schematic diagram of a load prediction test process based on a converter model provided in one embodiment of the present application; Figure 4 A schematic diagram of an execution logic of a server resource allocation method provided in one embodiment of the present application; Figure 5 This is an example diagram of a resource allocation device for a server according to an embodiment of the present application.
[0014] Among them, 10-server resource allocation device, 100-environment construction module, 200-test module, 300-resource allocation module. DETAILED DESCRIPTION
[0015] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0016] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0017] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0018] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the resource allocation method of the server depends, the specific application environment architecture or specific hardware architecture is described here.
[0019] An embodiment of the present application provides a method for allocating resources of a server.
[0020] like Figure 1 FIG. 1 is a flow chart of a method for allocating resources of a server according to an embodiment of the present application, wherein the method for allocating resources of a server includes the following steps: In step S101 , traffic and period fluctuations that meet preset requirements are applied to a pre-built cloud computing basic environment to establish a corresponding cloud computing enhanced environment.
[0021] Those skilled in the art should understand that cloud computing resource scheduling refers to a core technical mechanism that dynamically allocates physical / virtual resources through intelligent strategies to optimize resource utilization, ensure quality of service (QoS), and reduce operating costs. It can achieve dynamic resource allocation and automatically adjust resource ratios based on real-time load changes. For example, it can automatically expand computing nodes under burst traffic and release idle resources during off-peak periods. It uses containerization technology (such as Kubernetes HPA (Horizontal Pod Autoscaler)) to achieve microservice-level elastic scaling and optimize cross-node communication efficiency.
[0022] Therefore, the embodiments of the present application can first construct a basic environment for cloud computing and perform environmental enhancement deployment operations on the basic environment, thereby establishing a corresponding cloud computing enhancement environment and providing reliable technical support for the execution of subsequent corresponding algorithm verification operations.
[0023] Optionally, in one embodiment of the present application, before applying traffic and periodic fluctuations that meet preset requirements to a pre-built cloud computing basic environment, it also includes: deploying a preset monitoring system automation management component and integrating a target log analysis stack to build a dynamic resource monitoring layer through the monitoring system automation management component and the target log analysis stack; constructing a hybrid load model that obtains the traffic characteristics and periodic fluctuation characteristics of the target server, and establishing an automated testing tool chain based on the hybrid load model and the preset fault injection mechanism; and constructing a cloud computing basic environment based on the automated testing tool chain and the dynamic resource monitoring layer.
[0024] It should be noted that the embodiments of the present application can build a cloud computing infrastructure environment by constructing a dynamic resource monitoring layer and an automated testing tool chain, as described below: 1. Dynamic resource monitoring layer: The embodiments of the present application can deploy monitoring system automated management components, such as Prometheus-Operator, to collect multi-dimensional indicators such as processor, memory, and network, and integrate the EFK (Elasticsearch Filebeat Kibana) log analysis stack to capture abnormal events in real time; 2. Automated testing tool chain: (1) Stress testing: The embodiment of the present application can construct a hybrid load model of Locust and JMeter to obtain traffic characteristics and periodic fluctuation characteristics of the server's burst traffic, etc. (2) Fault injection: The embodiments of the present application can support a topology-aware fault pattern library through the ChaosMesh enhanced version. In addition, in the embodiments of the present application, fault injection includes Pod-level fault injection, network fault injection, storage system fault injection, Kubernetes control plane fault injection, and mixed fault scenario injection, as described below: 1. Pod-level fault injection: (1) Pod termination: 1) pod-kill: Randomly terminates pods with specified labels, simulating a Kubernetes node crash scenario; 2) pod-failure: Makes the pod continuously unavailable to test the service's self-healing capability (the duration can be configured to 10-60 minutes).
[0025] (2) Resource limitations: 1) Processor pressure: Limit the number of processor cores available to the Pod to simulate resource competition scenarios; 2) Memory exhaustion: Test the memory recovery mechanism through OOM (Out Of Memory) injection.
[0026] 2. Network Fault Injection: (1) Basic network interference: 1) Delay injection: 50ms-5s controllable delay to simulate cross-region communication scenarios; 2) Packet loss rate: Set the packet loss rate from 1% to 30% to test the fault tolerance of weak networks; 3) Network partitioning: Isolate Pod communication within a specific namespace.
[0027] (2) Advanced network scenarios: 1) DNS (Domain Name System) pollution: tampering with the resolution results of a specific domain name; 2) Bandwidth limitation: Limit the Pod egress bandwidth (e.g., 10 Mbps).
[0028] 3. Storage system failure: (1) I / O (Input / Output) exception: 1) Delay injection: file read and write delays increase by 100ms-2s; 2) Error injection: Simulates a bad disk block and returns EIO (Input / Output Error).
[0029] (2) File system failure: 1) rm -rf attack: randomly delete files in the container; 2) Disk full: Use the dd command to quickly fill the storage space.
[0030] 4. Kubernetes control plane failure: (1) API (Application Programming Interface) Server Interference: 1) Simulate etcd connection timeout and test the controller retry logic; 2) Inject a 500 error response to verify client fault tolerance.
[0031] (2) Scheduler interference: 1) Node unschedulable mark; 2) Insufficient simulation resources lead to the Pending state.
[0032] 5. Mixed failure scenarios: (1) Cascading failure test: 1) Simultaneously inject network latency and Pod termination to verify the service degradation strategy; 2) Processor stress and disk I / O latency combined test.
[0033] (2) Timing control: 1) Periodic failures (e.g., a 5-minute network outage every hour); 2) Conditional triggering (automatically injecting memory leaks when processor pressure > 80%).
[0034] During the specific implementation process, all fault injections in the embodiments of the present application can be visually configured through Chaos Dashboard, and the impact indicators can be monitored in real time.
[0035] Therefore, the embodiments of the present application achieve full-stack observability and high-availability verification of the cloud computing environment through dynamic resource monitoring and intelligent fault injection, significantly improving system resilience, fault self-healing capabilities and operation and maintenance efficiency.
[0036] Optionally, in one embodiment of the present application, a hybrid load model for obtaining traffic characteristics and periodic fluctuation characteristics of a target server is constructed, including: inserting a ladder thread group plug-in based on a preset stepped stress configuration of a stress testing tool to generate corresponding stress test results through the ladder thread group plug-in; performing scenario modeling operations through the preset load testing tool to generate corresponding load test statistics; writing the stress test results to a target time series database through a preset backend listener, connecting the load test statistics to a preset visual monitoring dashboard, and performing aggregation calculations on multiple system performance indicators based on the stress test results and the load test statistics to obtain corresponding aggregate indicators; determining a stress testing mechanism and a load testing mechanism corresponding to the stress testing tool and the load testing tool, and constructing a hybrid load model based on the stress testing mechanism and the load testing mechanism; based on the aggregate indicators and the hybrid load model, detecting whether the CPU occupancy rate corresponding to the stress testing tool is greater than a preset occupancy threshold, and determining whether the load testing tool has detected a target type error; when the CPU occupancy rate corresponding to the stress testing tool is greater than the preset occupancy threshold, automatically capturing the stack log of the target server; when the load testing tool detects a target type error, correlating and analyzing the corresponding payment interface parameters.
[0037] Specifically, Figure 2 This is a flow chart for building a hybrid load model of Locust and JMeter for this application. Figure 2 As shown, the process of building the Locust and JMeter hybrid load model in this application is as follows: S201: Perform step-by-step stress configuration using the stress testing tool and insert the step-by-step thread group plug-in: JMeter step stress configuration, insert the Stepping Thread Group plugin (STG); S202: Load testing tools for complex scenario modeling: Locust complex scene modeling, the specific procedures are as follows: from locust import HttpUser, task, between import random class HybridUser(HttpUser): wait_time = between(1, 3) @task(3)# 70% probability of executing browsing def browse(self): categories = ["electronics", "books"] self.client.get(f" / product?cat={random.choice(categories)}") @task(1)# 30% probability of executing payment def checkout(self): self.client.post(" / order", json={"items": [101, 205]})# Dynamically construct order:ml-citation{ref="2,10" data="citationList"}; S203: Write the stress test results to the target time series database through the preset backend listener, and connect the load test statistics to the preset visual monitoring dashboard: Real-time data fusion: JMeter results (i.e., stress test results) are written to InfluxDB5 through the Backend Listener, and Locust statistics (i.e., load test statistics) are connected to the Grafana dashboard. S204: performing aggregate calculations on multiple system performance indicators (i.e., key indicators) such as throughput, error rate, and response time based on the stress test results and load test statistical data to obtain corresponding aggregate indicators; S205: Resource bottleneck detection: When JMeter's step-by-step pressure triggers processor pressure greater than 80%, stack logs are automatically captured; when Locust detects an HTTP 500 error, payment interface parameters are correlated and analyzed.
[0038] It can be understood that the embodiments of the present application can solve the limitations of a single tool in high-concurrency and complex scenarios through the division of labor and collaboration of "step pressure and behavioral simulation", and greatly improve the anomaly detection coverage through the hybrid load model, while also effectively reducing test resource consumption.
[0039] Optionally, in one embodiment of the present application, traffic and periodic fluctuations that meet preset requirements are applied to a pre-built cloud computing basic environment to establish a corresponding cloud computing enhanced environment, including: deploying a service grid through a preset container orchestration cluster in the cloud computing basic environment to containerize the cloud computing basic environment to obtain a corresponding containerized test environment; applying traffic and periodic fluctuations that meet preset requirements through a mixed load model in the containerized test environment to enhance the containerized test environment to obtain a cloud computing enhanced environment.
[0040] It should be noted that the embodiments of the present application can build a highly simulated containerized test environment based on the Kubernetes cluster, and achieve refined traffic management and service orchestration by deploying the Istio service grid.
[0041] Specifically, the embodiments of the present application can first utilize Istio's traffic mirroring and dynamic routing functions to isolate test traffic from production environment traffic, and at the same time simulate abnormal scenarios such as call delays, retries, and circuit breaking between services; secondly, the embodiments of the present application can use Istio's monitoring and tracing capabilities to collect service call chains, resource occupancy (processor / memory usage, network throughput) and other indicators in real time, providing full-link observability support for load testing. Afterwards, the embodiment of the present application can design a hybrid load model based on the burst traffic scenario and the long-period fluctuation scenario, as described below: 1. Burst traffic scenario: Construct a short-term impact model of 300% peak traffic (e.g., lasting 5 minutes). This traffic is injected in a step-by-step manner to simulate the impact of sudden requests in real business (such as flash sales and live streaming) on the system. Focus on verifying the cold start warm-up mechanism of the service (such as the rapid expansion efficiency of container instances, cache warm-up completion time, and the initialization response speed of dependent services) to ensure that the system can respond quickly to sudden traffic increases and avoid request timeouts.
[0042] 2. Long-term fluctuation scenario: In this embodiment, a 24-hour periodic load curve can be designed, using a sinusoidal wave to simulate the regular traffic fluctuations of daily business (such as the difference in traffic peaks between weekdays and nighttime), while also superimposing random pulse interference (such as short, small-amplitude traffic jitter) to simulate the uncertainty of user behavior. This model can be used to verify the system's adaptability to periodic loads (such as the smoothness of dynamic resource expansion and contraction and the efficiency of automatic adjustment of the database connection pool) and its tolerance to random interference during long-term operation. Therefore, the embodiments of the present application achieve high consistency between the test environment and the production environment by combining a containerized environment with a service grid, and at the same time improve the authenticity of the load test and the efficiency of problem location through refined traffic control and full-link monitoring; in addition, the hybrid load model in the embodiments of the present application covers sudden shocks and long-term fluctuation scenarios, which can comprehensively verify the system's cold start performance, elastic scalability and long-term stability, expose potential risks under extreme traffic in advance, and provide a reliable basis for resource allocation optimization and fault-tolerant mechanism design in the production environment.
[0043] In step S102, a test error value and at least one key indicator value of the cloud computing enhancement environment are obtained, and it is determined whether the test error value and at least one key indicator value meet the corresponding preset verification requirements, so as to obtain corresponding load prediction test data and scheduling strategy efficiency test data when the test error value and at least one key indicator value meet the preset verification requirements.
[0044] Furthermore, the embodiments of the present application also need to use the Transformer converter model to conduct load prediction tests on the cloud computing enhanced environment to output the average absolute percentage error of the prediction results; at the same time, a multi-objective optimization algorithm is used to verify the scheduling strategy and generate at least one key indicator value; thereafter, the embodiments of the present application can finally obtain load prediction test data and scheduling strategy efficiency test data by judging whether the test error value (such as the average absolute percentage error) and the key indicator value meet the respective preset verification standards.
[0045] Therefore, the embodiments of the present application achieve dual verification of the load prediction accuracy and scheduling strategy efficiency of the cloud computing environment by combining the converter model with the multi-objective optimization algorithm, making the test dimension more comprehensive; in addition, the embodiments of the present application quantify the evaluation results through preset verification standards, thereby providing a clear basis for subsequent optimization of the load prediction model and scheduling strategy, and improving the reliability and adaptability of cloud computing resource scheduling.
[0046] Optionally, in one embodiment of the present application, a test error value and at least one key indicator value of the cloud computing enhanced environment are obtained, and it is determined whether the test error value and at least one key indicator value meet the corresponding preset verification requirements, so that when the test error value and at least one key indicator value meet the preset verification requirements, corresponding load prediction test data and scheduling strategy efficiency test data are obtained, including: based on a pre-built converter model, performing a load prediction test on the cloud computing enhanced environment to calculate the test error value, and verifying whether the test error value is less than a preset prediction error threshold to obtain the load prediction test data; using a preset multi-objective optimization algorithm to perform scheduling strategy verification on the cloud computing enhanced environment to calculate at least one key indicator value, and verifying whether at least one key indicator value is less than or equal to the preset indicator threshold to obtain the scheduling strategy efficiency test data.
[0047] In the actual implementation process, the load prediction test of the cloud computing hardening environment using the converter model in the embodiment of the present application mainly includes two parts: model comparison and accuracy control, and real-time data stream processing mechanism, which are described in detail as follows: 1. Model comparison and accuracy control: In the embodiment of the present application, the Transformer model can be selected as the core prediction model, and compared with traditional time series models such as LSTM (Long Short-Term Memory) and ARIMA (Auto Regressive Integrated Moving Average), with a focus on verifying the advantages of the Transformer in capturing long-term time series dependencies and fitting nonlinear load patterns.
[0048] In the embodiment of the present application, the hard indicator of prediction accuracy can be set as a test error value, such as the mean absolute percentage error (MAPE) ≤ 12%, covering typical periodic loads (such as daily business peaks), burst loads (such as traffic pulses) and long-term trend load scenarios in cloud computing environments, ensuring that the model can still maintain high-precision predictions under complex load patterns.
[0049] 2. Real-time data stream processing mechanism: In the embodiments of the present application, the Flink stream processing framework can be used to implement real-time data access and calculation. By configuring sliding windows (such as 5-minute windows, 1-minute sliding steps) or session windows (dynamically dividing window boundaries based on load fluctuation characteristics), incremental processing and feature extraction are performed on real-time resource monitoring data (such as processor utilization, memory usage, and request volume). This ensures that the prediction model can dynamically update parameters based on the latest data and achieve near-real-time load prediction (latency controlled at the second level).
[0050] In the actual implementation process, the embodiment of the present application uses a multi-objective optimization algorithm to verify the scheduling strategy of the cloud computing reinforcement environment. The process mainly includes two parts: multi-objective optimization algorithm design and resource fragmentation control, which are described in detail as follows: 1. Multi-objective optimization algorithm design: The embodiments of the present application may use an improved multi-objective optimization algorithm (such as NSGA-III (Non-dominated Sorting Genetic Algorithm III, non-dominated sorting genetic algorithm III), MOEA / D (Multi-Objective Evolutionary Algorithm based on Decomposition, multi-objective evolutionary algorithm based on decomposition) to verify the scheduling strategy. The optimization objectives cover resource utilization (processor / memory usage), service response time, energy consumption cost and resource balance, and balance the priorities of different objectives through weight allocation (such as giving priority to ensuring response time during high-load periods and giving priority to improving resource utilization during low-load periods).
[0051] 2. Resource fragmentation control: The embodiment of the present application may use the resource fragmentation index (an indicator that measures the proportion of discrete idle resource blocks) as a core verification indicator, and requires that the index is no greater than 0.12; in addition, the embodiment of the present application may dynamically adjust the resource allocation granularity (such as virtual machine / container specification matching, cross-node resource aggregation) through an algorithm to reduce "small and scattered" idle resource blocks, while avoiding the problem of excessive single-point load caused by over-centralized allocation, ensuring that resource allocation achieves the optimal balance between compactness and balance. Therefore, the embodiments of the present application can ensure the high accuracy of load prediction through comparative verification of the converter model and the traditional model, combined with MAPE precision control, to provide a reliable basis for resource scheduling, and can support dynamic adaptation to load fluctuations through the real-time performance of Flink window calculations, thereby avoiding improper resource allocation due to prediction lags; in addition, the embodiments of the present application can also take into account resource efficiency and service quality through multi-objective optimization algorithms, and cooperate with resource fragmentation index control to reduce resource waste and operating costs, while ensuring service stability and improving the overall resource scheduling efficiency of the cloud computing environment.
[0052] Optionally, in one embodiment of the present application, based on a pre-built converter model, a load prediction test is performed on a cloud computing reinforcement environment to calculate a test error value, and verify whether the test error value is less than a preset prediction error threshold to obtain load prediction test data, including: obtaining historical load data, meteorological data, holiday marking data and target period load value of the target server, and constructing corresponding training data sets and test data sets according to the historical load data, meteorological data, holiday marking data and target period load value; normalizing the training data sets and test data sets to obtain standard training data sets and standard test data sets, and configuring multiple hyperparameters of the converter model, The converter model configured with multiple hyperparameters is trained using a standard training data set; the performance of the converter model is tested using a standard test data set to obtain a test error value of the converter model; it is determined whether the test error value is less than a preset prediction error threshold, wherein, when the test error value is less than the preset prediction error threshold, the trained converter model is controlled to generate a corresponding prediction curve at a target time; the prediction curve is distributed to the corresponding computing node of the target server to calculate the corresponding prediction deviation value, and it is determined whether the prediction deviation value is greater than the preset deviation threshold, wherein, when the prediction deviation value is greater than the preset deviation threshold, a dynamic expansion operation of computing resources is performed.
[0053] Specifically, Figure 3 Schematic diagram of the load prediction test process based on the converter model of this application. Figure 3 As shown, the load prediction test process based on the converter model of this application is as follows: S301: Construct training and test datasets and perform normalization: In the embodiment of the present application, the input features of the converter model include: historical load data (timestamp, power value), meteorological data (temperature, humidity), and holiday marks; the output target is the load value of the next 24 hours (sampling interval is 15 minutes); Thus, the embodiment of the present application can obtain the server's historical load data, meteorological data, holiday mark data, and target period load value (such as the load value for the next 24 hours) to construct corresponding training data sets and test data sets, and perform Z-Score normalization on the training data sets and test data sets to obtain standard training data sets and standard test data sets; S302: Configure the hyperparameters of the converter model and train the converter model using the normalized training dataset: In the specific implementation process, the embodiment of the present application mainly configures multiple hyperparameters of the converter model, such as the number of encoder layers, training cycle, batch size, and learning rate scheduling. As an achievable method, the number of encoder layers can be set to 3 layers (wherein the hidden layer dimension d_model=64 and the number of attention heads nhead=4); the training cycle can be set to a maximum of 200 epochs (early stopping threshold=10); the batch size is set to 256 (the number of gradient accumulation steps is set to 4); the loss function is set to HuberLoss (δ=1.0) to balance the advantages of MSE (Mean Squared Error) and MAE (Mean Absolute Error); learning rate scheduling: Cosine annealing (initial lr=5e-4, minimum lr=1e-5); S303: Ensure that the test error value (e.g., mean absolute percentage error) of the test data set is no greater than a preset prediction error threshold (e.g., 12%) to meet industry-level standards; S304: The trained converter model generates a corresponding prediction curve at the target time and distributes the prediction curve to the corresponding computing node of the target server to calculate the corresponding prediction deviation value and determine whether the prediction deviation value is greater than a preset deviation threshold. If the prediction deviation value is greater than the preset deviation threshold, a dynamic expansion operation of computing resources is performed: During the actual execution process, the converter model can generate a prediction curve at 0:00 every day, distribute it to each computing node through Flink broadcast status, and calculate the predicted deviation between the real-time window data and the predicted value. When the predicted deviation is greater than 15%, dynamic scheduling is triggered to expand computing resources.
[0054] Therefore, the embodiment of the present application can identify nonlinear correlations spanning more than 72 hours in historical load data through the self-attention mechanism, which greatly reduces the prediction error compared to the LSTM model. The multi-scale converter model of the embodiment of the present application effectively reduces the mean absolute percentage error by iteratively refining the predictions of different time granularities.
[0055] In addition, as an achievable method, the embodiment of the present application can also optimize and fine-tune the converter model based on the real-time prediction results. The specific process is as follows: Step 1: Collect server load test data in real time. Combined with the daily prediction curve generated by the Transformer model, calculate the deviation rate between the real-time window data and the predicted value. Build a deviation-feature association library to record the historical load, weather, and holiday characteristics corresponding to when the deviation rate exceeds 15%. Step 2: Set up a dynamic fine-tuning trigger mechanism. When the deviation rate of three consecutive 15-minute sampling intervals exceeds 15%, or when the cumulative deviation exceeds 15% for four or more periods in a single day, initiate model fine-tuning. Step 3: Extract a subset of abnormal features based on the deviation-feature association library and fine-tune the model using incremental training. Freeze the parameters of the first two encoder layers and only adjust the third encoder layer (the hidden layer dimension remains at 64, and the number of attention heads dynamically adapts to the abnormal feature dimension). Step 4: Optimize the fine-tuning strategy. The loss function uses weighted Huber Loss (the weight of samples with a deviation exceeding 15% is increased to 1.5). The initial learning rate is set to 1 / 5 of the original lr (1e-4). The training cycle is 1 / 4 of the original cycle (maximum 50 epochs, early stopping threshold 5). Step 5: If the fine-tuned model passes the test set validation (MAPE ≤ 10%), the original model is replaced and the Flink broadcast state is updated. Otherwise, the model is reverted to the most recently valid model and abnormal features are marked for data augmentation in the next round of training. Step 6: Fine-tune all features every 7 days, integrate weekly deviation data to optimize the attention weight distribution, and ensure that the model adapts to long-term load change trends.
[0056] It should be noted that in the above process, step 1 establishes the association between deviation and features to provide targeted basis for fine-tuning; the dynamic trigger mechanism in step 2 avoids frequent fine-tuning and ensures model stability; the incremental training and parameter freezing in step 3 balance the fine-tuning efficiency and model performance; the weighted loss function in step 4 focuses on high-deviation samples to improve fine-tuning accuracy; steps 5-6 form a closed loop of short-term correction and long-term optimization, which is logically coherent and can adapt to dynamic changes in load, innovatively realizing the adaptive optimization of the model.
[0057] Therefore, the embodiments of the present application improve the real-time adaptability of the model through a dynamic fine-tuning mechanism, focus on the optimization accuracy of high-deviation samples, enhance the reliability of server load prediction, and effectively ensure the accuracy of load test verification and the rationality of resource scheduling.
[0058] Optionally, in one embodiment of the present application, a preset multi-objective optimization algorithm is used to verify the scheduling strategy of the cloud computing enhanced environment to calculate at least one key indicator value, including: obtaining the number of default tasks, total number of tasks, unused resource blocks and total resources corresponding to the target server in the cloud computing enhanced environment; calculating the service level agreement default rate based on the number of default tasks and the total number of tasks, and calculating the resource fragmentation index using the unused resource blocks and the total resources; and determining at least one key indicator value based on the service level agreement default rate and the resource fragmentation index.
[0059] It should be noted that the embodiments of the present application can first construct a multi-objective optimization framework based on the NSGA-II algorithm. In the actual implementation process, the embodiments of the present application can use NSGA-II as the core optimization algorithm to search for the Pareto optimal solution set that simultaneously optimizes the SLA default rate and resource fragmentation index in the solution space through fast non-dominated sorting, congestion calculation, and elite retention strategy. The specific process is as follows: 1. When initializing the population, the embodiment of the present application can encode the key parameters of the scheduling strategy (such as resource allocation thresholds, task priority weights, migration trigger conditions, etc.) into chromosomes to ensure population diversity; 2. During the iteration process, crossover and mutation operations are used to generate offspring solutions. The parent and offspring solutions are then combined to perform non-dominated sorting, screening out individuals that perform better on both objective functions. Ultimately, a Pareto frontier solution set is formed. This solution set contains multiple non-dominated solutions, each corresponding to a set of scheduling policy parameters. These solutions can achieve different levels of trade-offs between the SLA (Service-Level Agreement) default rate and the resource fragmentation index (e.g., prioritizing solutions that guarantee the SLA or prioritizing solutions that reduce resource fragmentation). Secondly, the embodiments of the present application can perform SLA default rate quantification and control operations, as described below: 1. Objective function f1(x) = number of default tasks / total number of tasks, where default tasks refer to tasks that fail to meet the agreed indicators in the service level agreement (e.g., response time exceeds the threshold, availability is lower than the promised value, data transmission latency exceeds the limit, etc.); 2. During the test, task flows of different priorities were injected (high-priority tasks correspond to more stringent SLA indicators) to simulate real business scenarios. The optimized f1(x) was required to be less than 1% to ensure the stability of the service quality of the core business. Afterwards, the embodiment of the present application can perform accurate measurement and optimization operations of the resource fragmentation index, as described below: 1. Objective function f2(x) = unused resource blocks / total resources, where unused resource blocks refer to scattered resource units that cannot be effectively allocated (such as scattered memory pages in the physical machine that are not occupied by virtual machines, fragmented idle time of the processor core, and uncontiguously allocated disk blocks in the storage medium). 2. The test covers multiple resource types (such as processors, memory, and network bandwidth). By calculating the fragmentation index of different resource dimensions and weighted fusion, the optimized requirement is that f2(x) ≤ 0.15 to ensure the regularity of resource allocation and reduce the waste of small and scattered idle resources. Therefore, the embodiment of the present application generates a Pareto frontier solution set through the NSGA-II algorithm, and while strictly controlling the SLA default rate and resource fragmentation index, achieves a dynamic balance between service quality and resource efficiency, and adapts to diversified business needs; secondly, the embodiment of the present application greatly improves the abnormal response speed and fault recovery capability through the collaboration of prediction model and dynamic scheduling, and enhances system stability and reliability; in addition, the embodiment of the present application directly reduces resource waste by 40% through refined resource fragmentation control (processor / memory fragmentation rate is not greater than 0.12), reduces operating costs, and at the same time guarantees core business SLA through priority stratification, thereby improving user satisfaction.
[0060] In step S103, elastic scaling tests and regional failure scenario simulations are performed on the cloud computing enhanced environment to generate corresponding elastic scaling test data and cross-regional disaster recovery test data, so as to dynamically allocate resources for the target server based on the load prediction test data, scheduling strategy efficiency test data, elastic scaling test data and cross-regional disaster recovery test data.
[0061] Afterwards, the embodiment of the present application generates elastic scaling test data and cross-regional disaster recovery test data by performing full-link enhanced verification on the cloud computing enhanced environment; secondly, the embodiment of the present application can integrate the load prediction test data and scheduling strategy efficiency test data previously obtained, and combine the elastic scaling test data and cross-regional disaster recovery test data as data basis for dynamic allocation of server resources.
[0062] Therefore, the embodiments of the present application cover elastic scaling and cross-regional disaster recovery scenarios through full-link verification, and combine load prediction and scheduling efficiency data to achieve comprehensiveness of resource allocation basis and improve scientific decision-making; in addition, the embodiments of the present application can also support dynamic allocation through multi-dimensional test data linkage, thereby ensuring that server resources are more adaptable and stable when dealing with load fluctuations, fault disaster recovery and other scenarios, and optimizing the overall service quality.
[0063] Optionally, in one embodiment of the present application, elastic scaling testing and regional failure scenario simulation are performed on the cloud computing enhanced environment to generate corresponding elastic scaling test data and cross-regional disaster recovery test data, including: performing elastic scaling testing on the cloud computing enhanced environment to obtain the corresponding service startup time and network traffic switching delay, and verifying whether the service startup time and network traffic switching delay are less than the corresponding scaling test threshold to generate elastic scaling test data; simulating the regional failure scenario corresponding to the cloud computing enhanced environment to perform cross-regional disaster recovery testing on the cloud computing enhanced environment according to the regional failure scenario to obtain the cluster self-healing time in the regional failure scenario, and verifying whether the cluster self-healing time is less than the preset self-healing time threshold to obtain cross-regional disaster recovery test data.
[0064] It should be noted that the embodiments of this application mainly perform full-link verification and enhancement operations through elastic scaling testing and cross-region disaster recovery testing. The specific process is as follows: (1) Elastic scaling test: Preheating pool mechanism design: The embodiment of the present application can build a dynamic preheating pool, which automatically maintains a certain number of standby Java service instances (instance specifications match the mainstream business load characteristics) based on historical load peaks and real-time predicted traffic, and completes instance preheating by preloading core class libraries, initializing cache connections (such as database connection pools, Redis connections), etc. Startup time verification: For Java service instances in the preheated pool, simulate startup triggers in real business scenarios (such as capacity expansion triggered by traffic thresholds). Use tracking monitoring to monitor the entire process from "instance wake-up" to "receiving requests." The time requirement is no more than 10 seconds. At the same time, test the difference in time between cold start (non-preheated instances) and preheated start to verify the effectiveness of the preheating mechanism. Mesh traffic switching test: Dynamic traffic scheduling is implemented based on a service mesh (such as Istio). Traffic switching in a capacity expansion scenario (from existing instances to newly added preheated instances) is simulated. Full-link tracing tools (such as Jaeger) are used to measure request latency fluctuations during the traffic switching process. The switching delay (from the start of traffic migration to stable distribution) must be no more than 50ms, and there must be no request loss or timeout during the switching process.
[0065] (2) Cross-region disaster recovery test: AZ (Availability Zone)-level failure simulation: Through network isolation and power outages, we simulate sudden failures in the entire availability zone (such as data center power outages and network interruptions), triggering cross-zone disaster recovery mechanisms and focusing on testing the self-healing capabilities and data consistency of core components.
[0066] etcd cluster self-healing verification: Monitor the node status (number of surviving nodes, etc.) of the etcd cluster after an AZ failure, and verify whether the cluster can automatically complete operations such as data shard redistribution within 30 seconds and resume external services (confirmed through health check interfaces and data read and write tests). Verifying Redis cluster split-brain protection: When simulating an AZ network partition (inter-zone communication interruption), test the Redis cluster's split-brain protection mechanisms (such as the quorum mechanism and the min-replicas-to-write configuration) to verify whether dual-master node write conflicts occur in the partitioned state. After the partition is recovered, check whether the cluster can automatically merge data without loss to ensure data consistency. Therefore, the elastic scaling test of the embodiment of the present application can ensure that the service can quickly expand and respond to traffic bursts through the preheating pool and fast traffic switching mechanism, control the Java service startup delay to within 10 seconds and the traffic switching delay to within 50ms, and significantly improve the user experience and system elasticity; secondly, the cross-region disaster recovery test of the embodiment of the present application verifies the self-healing ability of the core components under extreme failures (etcd cluster self-healing within 30 seconds) and data consistency (Redis brain split protection), greatly improving the system's risk resistance in regional-level failures, reducing business interruption time and data loss. In addition, the full-link verification covers key scenarios from resource elasticity to disaster recovery, providing a quantitative basis for the high-availability design of cloud computing environments, and ensuring business continuity and stability.
[0067] Optionally, in one embodiment of the present application, an elastic scaling test is performed on a cloud computing enhanced environment to obtain a corresponding service startup time and network traffic switching delay, and verify whether the service startup time and network traffic switching delay are less than a corresponding scaling test threshold to generate elastic scaling test data, including: collecting multi-dimensional performance indicators of computing nodes in a target server, and generating expansion trigger conditions with time persistence and spatial correlation based on the multi-dimensional performance indicators and a preset historical load feature library to obtain an expansion decision instruction corresponding to the expansion trigger condition; selecting a matching instance specification from a preheated container resource pool according to the expansion decision instruction, and performing corresponding deployment and initialization operations using the matching instance specification to generate a new instance ready state signal; based on the new instance ready state signal, monitoring the service startup process of the target server to obtain a corresponding service startup time, and verifying whether the service startup time meets the preset startup requirements to obtain a service startup verification result; tracking the business traffic migration process of the target server through a traffic monitoring component to obtain a corresponding network traffic switching delay, and verifying whether the network traffic switching delay meets the preset compliance requirements to obtain a traffic switching verification result; generating elastic scaling test data based on the service startup verification result and the traffic switching verification result.
[0068] Specifically, the embodiments of the present application can generate elastic scaling test data through operations such as multi-dimensional trigger condition design, preheating resource pools and rapid deployment, and service startup and traffic switching verification, as described below: 1. Multi-dimensional trigger condition design: Real-time collection of multi-dimensional performance indicators such as processor utilization, memory usage, and network throughput of computing nodes. Combined with a preset historical load feature library (including load fluctuation patterns in different business scenarios), it generates expansion trigger conditions that have both time persistence (such as processor utilization greater than 70% and lasting for 2 minutes) and spatial correlation (multi-node load collaborative triggering). This ensures that the trigger timing accurately matches the actual load demand and outputs the corresponding expansion decision instructions. 2. Preheating the resource pool and rapid deployment: Based on the expansion decision instructions, the system selects instance specifications that match the current business load from the preheated container resource pool (such as preset JVM (Java Virtual Machine) parameters and preloaded core dependency containers). It performs lightweight deployment and initialization operations (skipping basic environment configuration and only activating business processes). It monitors the entire process from alarm triggering to the readiness of the new Pod, and requires that the expansion action take no more than 8 seconds. 3. Service startup and traffic switching verification: Track the service startup process of the new instance from initialization to receiving requests in real time, record the startup time and verify whether it meets the preset requirements; use the traffic monitoring component of the Nginx reverse proxy to track the migration process of business traffic from the original instance to the new instance, measure the switching delay (required to be less than 100ms), and ensure that no requests are lost during the switching process. Finally, combine the service startup verification results and traffic switching verification results to generate elastic scaling test data. Therefore, the embodiments of the present application can achieve rapid response to sudden load changes through precise triggering, rapid expansion and low-latency switching, thereby effectively ensuring service continuity and improving resource elasticity and user experience.
[0069] In addition, as an achievable method, the embodiments of the present application can also perform multi-level fault recovery testing and verification operations in a cloud computing environment in the following manner. The specific process is as follows: Step 1: Failure scenarios are simulated in stages. For data center-level failures, 40% of nodes are downgraded (increased by 15% / minute). For AZ-level failures, a combined network partition and power outage scenario is simulated. Step 2: Trigger the intelligent scheduling mechanism. Based on the K8s topology-aware scheduling and the dynamic weight of the service dependency graph and node health, prioritize the migration of critical services to pre-marked disaster recovery core nodes and record the migration completion time. Step 3: Perform data consistency checks. After the Redis cluster recovers, verify that the AOF (Append Only File) log is zero-loss and that the incremental hash fingerprints of the data before and after the failure match. During the self-healing process of the etcd cluster at the AZ level, cross-zone replicas are linked to accelerate synchronization. Step 4: Verify the disaster recovery and self-healing indicators. The etcd cluster self-healing time must be less than 30 seconds. The Redis cluster must implement split-brain protection through the AZ weight arbitration mechanism. There must be no data conflicts after the migration of key services. It should be noted that, in the embodiment of the present application, the above step 1 uses gradient downtime simulation to simulate the actual fault evolution process more closely, avoiding the extreme of instantaneous full downtime; the intelligent scheduling of step 2 combines service dependency and node health to improve migration accuracy and efficiency; the incremental hash check of step 3 supplements AOF log verification to enhance the credibility of data consistency; the AZ weight arbitration of step 4 allows Redis split-brain protection to adapt to cross-region scenarios, and the synchronization of etcd cross-region replicas accelerates the self-healing process, forming a closed loop of "fault simulation-intelligent migration-data verification-self-healing verification" as a whole. Therefore, the embodiments of the present application improve migration efficiency through gradient fault simulation and intelligent scheduling, and strengthen consistency through multi-dimensional data verification and cross-region collaborative self-healing, ensuring business continuity, significantly enhancing the disaster recovery reliability of the cloud computing environment, further ensuring dynamic allocation of server resources, and ensuring that the resource allocation system is more stable.
[0070] The following describes the execution logic of the resource allocation method of the server of the present application in conjunction with the accompanying drawings.
[0071] Figure 4 This is a schematic diagram of the execution logic of the resource allocation method of the server of this application. Figure 4 As shown, the execution process of the resource allocation method of the server of this application is as follows: S401: Build a dynamic resource monitoring layer and an automated testing tool chain to establish a cloud computing infrastructure environment based on the dynamic resource monitoring layer and the automated testing tool chain; S402: Build a containerized test environment corresponding to the cloud computing basic environment, and apply traffic and periodic fluctuations that meet preset requirements through a hybrid load model to strengthen the containerized test environment and obtain a cloud computing strengthened environment. S403: Perform load prediction testing and scheduling strategy verification on the cloud computing hardening environment to generate corresponding load prediction test data and scheduling strategy efficiency test data if the corresponding test or verification requirements are passed; S404: Perform elastic scaling test and cross-region disaster recovery test on the cloud computing enhanced environment to generate corresponding elastic scaling test data and cross-region disaster recovery test data if the corresponding test requirements are passed, and generate server allocation resources based on the load prediction test data, scheduling strategy efficiency test data, elastic scaling test data and cross-region disaster recovery test data.
[0072] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0073] An embodiment of the present application also provides a resource allocation device for a server.
[0074] like Figure 5 As shown, the resource allocation device 10 of the server includes: an environment construction module 100 , a testing module 200 and a resource allocation module 300 .
[0075] The environment building module 100 is used to impose traffic and period fluctuations that meet preset requirements on the pre-built cloud computing basic environment to establish a corresponding cloud computing enhanced environment.
[0076] The test module 200 is used to obtain the test error value and at least one key indicator value of the cloud computing enhancement environment, and to determine whether the test error value and at least one key indicator value meet the corresponding preset verification requirements, so as to obtain the corresponding load prediction test data and scheduling strategy efficiency test data when the test error value and at least one key indicator value meet the preset verification requirements.
[0077] The resource allocation module 300 is used to perform elastic scaling tests and regional failure scenario simulations on the cloud computing enhanced environment to generate corresponding elastic scaling test data and cross-regional disaster recovery test data, so as to dynamically allocate resources for the target server based on the load prediction test data, scheduling strategy efficiency test data, elastic scaling test data and cross-regional disaster recovery test data.
[0078] Optionally, in one embodiment of the present application, the resource allocation device 10 of the server further includes: a deployment module, a creation module and a construction module.
[0079] Among them, the deployment module is used to deploy the preset monitoring system automation management components before applying traffic and periodic fluctuations that meet preset requirements to the pre-built cloud computing basic environment, and integrate the target log analysis stack to build a dynamic resource monitoring layer through the monitoring system automation management components and the target log analysis stack.
[0080] A module is established to construct a hybrid load model that obtains the traffic characteristics and periodic fluctuation characteristics of the target server, so as to establish an automated testing tool chain based on the hybrid load model and the preset fault injection mechanism.
[0081] A building block for building a cloud computing infrastructure based on an automated testing tool chain and a dynamic resource monitoring layer.
[0082] Optionally, in one embodiment of the present application, the environment building module 100 includes: a containerized processing unit and an enhancement unit.
[0083] Among them, the containerized processing unit is used to deploy a service grid through a preset container orchestration cluster in the cloud computing basic environment to containerize the cloud computing basic environment and obtain a corresponding containerized test environment.
[0084] The strengthening unit is used to impose traffic and periodic fluctuations that meet preset requirements through a mixed load model in the containerized test environment to strengthen the containerized test environment and obtain a cloud computing strengthened environment.
[0085] Optionally, in one embodiment of the present application, the test module 200 includes: a first verification unit and a second verification unit.
[0086] Among them, the first verification unit is used to perform a load prediction test on the cloud computing enhancement environment based on a pre-built converter model to calculate a test error value and verify whether the test error value is less than a preset prediction error threshold to obtain load prediction test data.
[0087] The second verification unit is used to verify the scheduling strategy of the cloud computing enhancement environment using a preset multi-objective optimization algorithm to calculate at least one key indicator value and verify whether the at least one key indicator value is less than or equal to a preset indicator threshold to obtain scheduling strategy efficiency test data.
[0088] Optionally, in one embodiment of the present application, the resource allocation module 300 includes: an elastic scaling test unit and a scenario simulation unit.
[0089] Among them, the elastic scaling test unit is used to perform elastic scaling tests on the cloud computing enhanced environment to obtain the corresponding service startup time and network traffic switching delay, and verify whether the service startup time and network traffic switching delay are less than the corresponding scaling test threshold to generate elastic scaling test data.
[0090] The scenario simulation unit is used to simulate the regional failure scenario corresponding to the cloud computing enhanced environment, so as to perform cross-regional disaster recovery testing on the cloud computing enhanced environment according to the regional failure scenario, so as to obtain the cluster self-recovery time in the regional failure scenario, and verify whether the cluster self-recovery time is less than the preset self-recovery time threshold, so as to obtain cross-regional disaster recovery test data.
[0091] Optionally, in one embodiment of the present application, the establishment module includes: a generation unit, a scene modeling unit, an aggregation unit, a determination unit, a judgment unit, a capture unit and an association unit.
[0092] Among them, the generation unit is used to insert the ladder thread group plug-in based on the stepped stress configuration of the preset stress testing tool, so as to generate corresponding stress testing results through the ladder thread group plug-in.
[0093] The scenario modeling unit is used to perform scenario modeling operations through a preset load testing tool to generate corresponding load testing statistical data.
[0094] The aggregation unit is used to write the stress test results into the target time series database through the preset backend listener, connect the load test statistics to the preset visual monitoring dashboard, and perform aggregation calculations on multiple system performance indicators based on the stress test results and load test statistics to obtain corresponding aggregate indicators.
[0095] The determination unit is used to determine the stress testing mechanism and the load testing mechanism corresponding to the stress testing tool and the load testing tool, so as to construct a hybrid load model based on the stress testing mechanism and the load testing mechanism.
[0096] The judgment unit is used to detect whether the CPU occupancy rate corresponding to the stress testing tool is greater than a preset occupancy rate threshold based on the aggregation index and the mixed load model, and to determine whether the load testing tool detects a target type error.
[0097] The capture unit is used to automatically capture the stack log of the target server when the CPU occupancy rate corresponding to the stress test tool is greater than a preset occupancy rate threshold.
[0098] The association unit is used to associate and analyze the corresponding payment interface parameters when the load testing tool detects a target type error.
[0099] Optionally, in one embodiment of the present application, the first verification unit includes: a first acquisition subunit, a normalization subunit, a performance testing subunit, a control subunit and a distribution subunit.
[0100] Among them, the first acquisition subunit is used to obtain the historical load data, meteorological data, holiday marking data and target period load value of the target server, so as to construct corresponding training data set and test data set based on the historical load data, meteorological data, holiday marking data and target period load value.
[0101] The normalization subunit is used to normalize the training dataset and the test dataset to obtain a standard training dataset and a standard test dataset, configure multiple hyperparameters of the converter model, and train the converter model after configuring multiple hyperparameters using the standard training dataset.
[0102] The performance testing subunit is used to perform performance testing on the converter model using a standard test data set to obtain a test error value of the converter model.
[0103] The control subunit is used to determine whether the test error value is less than a preset prediction error threshold, wherein, when the test error value is less than the preset prediction error threshold, the trained converter model is controlled to generate a corresponding prediction curve at a target time.
[0104] The distribution sub-unit is used to distribute the prediction curve to the corresponding computing node of the target server to calculate the corresponding prediction deviation value and determine whether the prediction deviation value is greater than the preset deviation threshold. When the prediction deviation value is greater than the preset deviation threshold, the computing resources are dynamically expanded.
[0105] Optionally, in one embodiment of the present application, the second verification unit includes: a second acquisition subunit, a first calculation subunit and a second calculation subunit.
[0106] The second acquisition subunit is used to obtain the number of default tasks, the total number of tasks, the unused resource blocks and the total resources corresponding to the target server in the cloud computing enhancement environment.
[0107] The first calculation subunit is used to calculate the service level agreement default rate according to the number of default tasks and the total number of tasks, and calculate the resource fragmentation index using the unused resource blocks and the total resources.
[0108] The second calculation subunit is configured to determine at least one key indicator value based on the service level agreement default rate and the resource fragmentation index.
[0109] Optionally, in one embodiment of the present application, the elastic scaling test unit includes: a collection subunit, a selection subunit, a monitoring subunit, a tracking subunit and a test data generation subunit.
[0110] Among them, the collection subunit is used to collect multi-dimensional performance indicators of the computing nodes in the target server, and based on the multi-dimensional performance indicators and the preset historical load feature library, generate expansion trigger conditions with time continuity and spatial correlation to obtain expansion decision instructions corresponding to the expansion trigger conditions.
[0111] The selection subunit is used to select a matching instance specification from the preheated container resource pool according to the expansion decision instruction, so as to perform corresponding deployment and initialization operations using the matching instance specification to generate a new instance ready state signal.
[0112] The monitoring subunit is used to monitor the service startup process of the target server based on the new instance ready state signal to obtain the corresponding service startup time, and verify whether the service startup time meets the preset startup requirements to obtain a service startup verification result.
[0113] The tracking subunit is used to track the business traffic migration process of the target server through the traffic monitoring component to obtain the corresponding network traffic switching delay, and verify whether the network traffic switching delay meets the preset standard requirements to obtain the traffic switching verification result.
[0114] The test data generation subunit is used to generate elastic scaling test data based on the service startup verification results and traffic switching verification results.
[0115] For the description of the features in the embodiment corresponding to the resource allocation device of the server, please refer to the relevant description of the embodiment corresponding to the resource allocation method of the server, which will not be repeated here.
[0116] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned server resource allocation method embodiments.
[0117] An embodiment of the present application further provides a non-volatile computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned server resource allocation method embodiments when running.
[0118] In an exemplary embodiment, the non-volatile computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0119] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned server resource allocation method embodiments are implemented.
[0120] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned server resource allocation method embodiments are implemented.
[0121] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0122] The above is a detailed introduction to the resource allocation method, device, equipment and medium of a server provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
Claims
1. A method for allocating resources of a server, characterized in that: The following steps are involved: Apply traffic and period fluctuations that meet preset requirements to the pre-built cloud computing infrastructure environment to establish a corresponding cloud computing reinforcement environment; Obtaining a test error value and at least one key indicator value of the cloud computing hardening environment, and determining whether the test error value and the at least one key indicator value meet corresponding preset verification requirements, so as to obtain corresponding load prediction test data and scheduling strategy efficiency test data when the test error value and the at least one key indicator value meet the preset verification requirements; Perform elastic scaling tests and regional failure scenario simulations on the cloud computing enhanced environment to generate corresponding elastic scaling test data and cross-regional disaster recovery test data, so as to dynamically allocate resources for the target server based on the load prediction test data, the scheduling strategy efficiency test data, the elastic scaling test data, and the cross-regional disaster recovery test data.
2. The server resource allocation method according to claim 1, characterized in that: Before applying traffic and period fluctuations that meet preset requirements to the pre-built cloud computing infrastructure environment, it also includes: Deploy a preset monitoring system automation management component and integrate the target log analysis stack to build a dynamic resource monitoring layer through the monitoring system automation management component and the target log analysis stack; Constructing a hybrid load model that obtains the traffic characteristics and periodic fluctuation characteristics of the target server, and establishing an automated testing tool chain based on the hybrid load model and a preset fault injection mechanism; The cloud computing basic environment is constructed based on the automated testing tool chain and the dynamic resource monitoring layer.
3. The server resource allocation method according to claim 2, characterized in that: The aforementioned step of applying traffic and period fluctuations that meet preset requirements to the pre-built cloud computing infrastructure environment to establish a corresponding cloud computing reinforcement environment includes: Deploying a service grid in the cloud computing infrastructure environment through a preset container orchestration cluster to containerize the cloud computing infrastructure environment to obtain a corresponding containerized test environment; The mixed load model in the containerized test environment is used to apply traffic and periodic fluctuations that meet preset requirements to strengthen the containerized test environment and obtain the cloud computing strengthened environment.
4. The server resource allocation method according to claim 1, characterized in that: The obtaining of a test error value and at least one key indicator value of the cloud computing hardening environment, and determining whether the test error value and the at least one key indicator value meet corresponding preset verification requirements, so as to obtain corresponding load prediction test data and scheduling strategy efficiency test data when the test error value and the at least one key indicator value meet the preset verification requirements, includes: Based on the pre-built converter model, a load prediction test is performed on the cloud computing hardening environment to calculate the test error value, and verify whether the test error value is less than a preset prediction error threshold to obtain the load prediction test data; A preset multi-objective optimization algorithm is used to verify the scheduling strategy of the cloud computing enhancement environment to calculate the at least one key indicator value, and to verify whether the at least one key indicator value is less than or equal to a preset indicator threshold to obtain the scheduling strategy efficiency test data.
5. The server resource allocation method according to claim 1, characterized in that: The elastic scaling test and regional failure scenario simulation of the cloud computing hardening environment to generate corresponding elastic scaling test data and cross-region disaster recovery test data include: Performing an elastic scaling test on the cloud computing hardening environment to obtain corresponding service startup time and network traffic switching delay, and verifying whether the service startup time and the network traffic switching delay are less than corresponding scaling test thresholds to generate the elastic scaling test data; Simulate the regional failure scenario corresponding to the cloud computing enhanced environment to perform a cross-regional disaster recovery test on the cloud computing enhanced environment according to the regional failure scenario to obtain the cluster self-healing time in the regional failure scenario, and verify whether the cluster self-healing time is less than a preset self-healing time threshold to obtain the cross-regional disaster recovery test data.
6. The server resource allocation method according to claim 2, characterized in that: The constructing and obtaining a hybrid load model of the traffic characteristics and periodic fluctuation characteristics of the target server includes: Based on the preset stepped stress configuration of the stress testing tool, inserting a stepped thread group plug-in to generate corresponding stress testing results through the stepped thread group plug-in; Perform scenario modeling operations using preset load testing tools to generate corresponding load testing statistics; Writing the stress test results into a target time series database through a preset backend listener, accessing the load test statistics to a preset visual monitoring dashboard, and performing aggregate calculations on multiple system performance indicators based on the stress test results and the load test statistics to obtain corresponding aggregate indicators; Determining a stress testing mechanism and a load testing mechanism corresponding to the stress testing tool and the load testing tool, so as to construct the hybrid load model based on the stress testing mechanism and the load testing mechanism; Based on the aggregated indicator and the mixed load model, detecting whether the CPU occupancy rate corresponding to the stress testing tool is greater than a preset occupancy rate threshold, and determining whether the load testing tool detects a target type error; When the CPU occupancy rate corresponding to the stress testing tool is greater than the preset occupancy rate threshold, the stack log of the target server is automatically captured; When the load testing tool detects that the target type is wrong, the corresponding payment interface parameters are analyzed in association.
7. The server resource allocation method according to claim 4, characterized in that: The load prediction test is performed on the cloud computing hardening environment based on the pre-built converter model to calculate the test error value, and verify whether the test error value is less than a preset prediction error threshold to obtain the load prediction test data, including: Obtaining historical load data, meteorological data, holiday marking data, and target period load value of the target server, so as to construct corresponding training data sets and test data sets according to the historical load data, the meteorological data, the holiday marking data, and the target period load value; Normalizing the training dataset and the test dataset to obtain a standard training dataset and a standard test dataset, configuring multiple hyperparameters of the converter model, and training the converter model after configuring the multiple hyperparameters using the standard training dataset; Performing a performance test on the converter model using the standard test data set to obtain a test error value of the converter model; determining whether the test error value is less than the preset prediction error threshold, wherein, if the test error value is less than the preset prediction error threshold, controlling the trained converter model to generate a corresponding prediction curve at a target time; The prediction curve is distributed to the corresponding computing node of the target server to calculate the corresponding prediction deviation value, and determine whether the prediction deviation value is greater than the preset deviation threshold. When the prediction deviation value is greater than the preset deviation threshold, a dynamic expansion operation of computing resources is performed.
8. The server resource allocation method according to claim 4, characterized in that: The employing of a preset multi-objective optimization algorithm to perform scheduling strategy verification on the cloud computing enhancement environment to calculate the at least one key indicator value includes: Obtaining the number of default tasks, the total number of tasks, unused resource blocks, and the total resources corresponding to the target server in the cloud computing enhancement environment; Calculating a service level agreement default rate according to the number of defaulted tasks and the total number of tasks, and calculating a resource fragmentation index using the unused resource blocks and the total resources; The at least one key indicator value is determined based on the service level agreement breach rate and the resource fragmentation index.
9. The server resource allocation method according to claim 5, characterized in that: The elastic scaling test is performed on the cloud computing hardening environment to obtain corresponding service startup time and network traffic switching delay, and verifying whether the service startup time and the network traffic switching delay are less than corresponding scaling test thresholds to generate the elastic scaling test data, including: Collecting multidimensional performance indicators of computing nodes in the target server, and generating expansion trigger conditions with temporal persistence and spatial correlation based on the multidimensional performance indicators and a preset historical load feature library, so as to obtain expansion decision instructions corresponding to the expansion trigger conditions; Selecting a matching instance specification from the preheated container resource pool according to the expansion decision instruction, performing corresponding deployment and initialization operations using the matching instance specification, and generating a new instance ready state signal; Based on the new instance ready state signal, monitoring the service startup process of the target server to obtain a corresponding service startup time, and verifying whether the service startup time meets a preset startup requirement to obtain a service startup verification result; Tracking the service traffic migration process of the target server through a traffic monitoring component to obtain a corresponding network traffic switching delay, and verifying whether the network traffic switching delay meets a preset compliance requirement to obtain a traffic switching verification result; Based on the service startup verification result and the traffic switching verification result, the elastic scaling test data is generated.
10. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the server resource allocation method according to any one of claims 1 to 9 when executing the computer program.
Citation Information
Patent Citations
Resource scheduling method and device, equipment and storage medium
CN116257363A
Predictive elastic scaling method and system considering multi-dimensional load characteristics
CN118260089A
Multi-cluster distributed training-oriented disaster recovery drill and performance evaluation method and system
CN120104453A
Container cloud elastic expansion and contraction method based on load prediction
CN120216096A
High-availability testing method and system for task scheduling of computing power host management platform
CN120316014A
Cited By
Software testing method
CN120929386A