Resource scheduling method and system based on cloud environment
By acquiring real-time performance data and task queue information for trend analysis, resource demand prediction strategies and elastic scaling trigger conditions are generated. This solves the problem that traditional resource allocation schemes cannot adapt to the fluctuations of carbon satellite data processing tasks, and achieves efficient and adaptive resource scheduling, improving resource utilization and data processing stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NAT SATELLITE METEOROLOGICAL CENT
- Filing Date
- 2026-05-12
- Publication Date
- 2026-07-31
AI Technical Summary
Traditional cloud computing resource allocation schemes cannot adapt to the drastic fluctuations in carbon satellite data processing tasks, resulting in insufficient or idle resources, affecting data processing timeliness and computing costs, and making it difficult to maintain stable system operation.
By acquiring real-time performance data and task queue information, trend analysis is performed to generate resource demand prediction strategies and elastic scaling trigger conditions. Resource control instructions are formulated and executed, and resource allocation is monitored and adjusted to dynamically adapt to fluctuations in resource demand.
This improved resource utilization, reduced mission delays and resource waste, and ensured the stable operation and timeliness of carbon satellite data processing.
Smart Images

Figure CN122489210A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cloud computing resource management technology, and in particular to a resource scheduling method and system based on a cloud environment. Background Technology
[0002] In applications that integrate cloud computing and remote sensing satellite technology, carbon satellites, as core tools for monitoring global atmospheric carbon dioxide concentrations, place unique demands on the dynamic allocation of computing resources for their data processing tasks. These tasks are characterized by massive data volumes, complex processing procedures, and time sensitivity, leading to drastic fluctuations in resource demands.
[0003] Traditional resource allocation schemes generally adopt a static preset mode, which cannot adapt to the periodic changes and sudden increases in the workload. When carbon satellites pass overhead and generate massive amounts of observation data, fixed resource allocation is difficult to expand in a timely manner, resulting in backlog of task queues and processing delays, which directly affects the timeliness and value of environmental monitoring data; while during the off-season, the continuous idleness of surplus resources leads to a significant increase in computing costs.
[0004] Existing technical solutions lack a deep understanding of mission time patterns and workload characteristics, making them particularly ill-equipped to handle sudden changes in resource demands caused by temporary adjustments to carbon satellite observation plans. During peak mission periods, insufficient resource supply often leads to reduced processing efficiency, while resource redundancy during off-peak periods drives up maintenance costs. This dynamic imbalance between resource allocation and mission requirements makes it difficult for the system to maintain stable operation under the stringent time-sensitive requirements of carbon satellite data processing, severely hindering the timely acquisition, analysis, and application of environmental monitoring data. Summary of the Invention
[0005] This invention provides a cloud-based resource scheduling method and system that can dynamically adapt to fluctuations in resource demand, improve resource utilization, reduce task delays and resource waste, and ensure the stable operation of time-sensitive applications such as carbon satellite data processing.
[0006] To achieve the above objectives, in a first aspect, the present invention provides a resource scheduling method based on a cloud environment, comprising: acquiring real-time performance data and task queue information from the cloud environment to form raw feature data; performing trend analysis based on the raw feature data to obtain resource usage patterns and potential bottleneck prediction results; generating a resource demand prediction strategy file and a set of elastic scaling trigger conditions based on the resource usage patterns and potential bottleneck prediction results; wherein the set of elastic scaling trigger conditions includes time-period-based planned trigger rules and performance index threshold-based trigger rules; generating a resource configuration template based on the resource demand prediction strategy file, and generating an elastic strategy template with defined trigger rules based on the set of elastic scaling trigger conditions; formulating an elastic scaling execution plan and issuing resource control instructions based on the trigger rules in the elastic strategy template and the resource parameters in the resource configuration template; monitoring the execution status of the resource control instructions, identifying service endpoints with abnormal availability through a health check mechanism, and adjusting resource allocation based on resource prediction scheduling logic; simultaneously, feeding back the adjusted resource allocation status to the raw feature data acquisition stage.
[0007] Secondly, this invention provides a cloud-based resource scheduling system, based on the cloud-based resource scheduling method described above. The system includes: a raw feature data composition module, an acquisition module, a first generation module, a second generation module, a control command issuance module, and a resource allocation adjustment module. The raw feature data composition module acquires real-time performance data and task queue information from the cloud environment to form raw feature data. The acquisition module performs trend analysis based on the raw feature data to obtain resource usage patterns and potential bottleneck prediction results. The first generation module generates a resource demand prediction strategy file and a set of elastic scaling trigger conditions based on the resource usage patterns and potential bottleneck prediction results; wherein the elastic scaling trigger condition set includes time-period-based plan trigger rules and performance index threshold-based trigger rules. The second generation module generates a resource configuration template based on the resource demand prediction strategy file and an elastic strategy template with defined trigger rules based on the elastic scaling trigger condition set. The control command issuance module formulates an elastic scaling execution plan and issues resource control commands based on the trigger rules in the elastic strategy template and the resource parameters in the resource configuration template. The resource allocation adjustment module is used to monitor the execution status of the resource control commands, identify service endpoints with abnormal availability through a health check mechanism, and adjust resource allocation in combination with resource prediction and scheduling logic; at the same time, the adjusted resource allocation status is fed back to the original feature data collection stage.
[0008] Thirdly, the present invention provides an electronic device, comprising:
[0009] At least one processor; and
[0010] A memory that is communicatively connected to the at least one processor;
[0011] The memory stores instructions that can be executed by the at least one processor, which are then executed by the at least one processor to enable the at least one processor to perform the cloud-based resource scheduling method as described above.
[0012] Fourthly, the present invention provides a computer-readable storage medium including a computer program and instructions, which, when the computer program or the instructions are executed on a computer, cause the computer to perform the resource scheduling method based on a cloud environment as described above.
[0013] Compared with existing technologies, the cloud-based resource scheduling method and system of the present invention obtains real-time performance data and task queue information from the cloud environment, performs trend analysis to generate resource demand prediction strategies and elastic scaling conditions, formulates and executes resource control instructions, and monitors and adjusts resource allocation. This achieves efficient and adaptive resource scheduling, dynamically adapts to fluctuations in resource demand, improves resource utilization, reduces task delays and resource waste, and ensures the stable operation of time-sensitive applications such as carbon satellite data processing. Attached Figure Description
[0014] Figure 1 This is a flowchart illustrating a resource scheduling method based on a cloud environment according to Embodiment 1 of the present invention.
[0015] Figure 2 This is a schematic diagram of a cloud-based resource scheduling system according to Embodiment 2 of the present invention;
[0016] Figure 3 This is a schematic diagram of the structure of an electronic device according to Embodiment 3 of the present invention. Detailed Implementation
[0017] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0018] It should be noted that the terms "first," "second," "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0019] To facilitate understanding, the main implementation concepts of the various embodiments of the present invention will be briefly described first.
[0020] In traditional cloud-based carbon satellite data processing missions, the massive data volume, complex processes, and high timeliness requirements result in a lack of real-time matching between resource allocation and mission needs. This problem stems from the fact that traditional resource allocation methods use a pre-fixed model, which cannot be dynamically adjusted according to actual mission requirements, leading to an ineffective dynamic balance between resources and mission demands. Specifically, the system struggles to maintain stable operation when the workload fluctuates drastically, and insufficient or idle resources directly impact the timeliness of data processing and resource utilization efficiency, thus adversely affecting key system performance indicators.
[0021] For example, during each transit of the carbon satellite, a massive amount of observational data is generated in a short period of time that needs to be processed. If computing resources cannot be expanded in time, data processing tasks will back up, hindering subsequent analysis processes. Conversely, if resources are maintained at a high configuration for a long period, computing resources will be idle during off-peak periods, increasing the burden of operating costs. In this scenario, the sudden and cyclical changes in resource demand make it difficult for existing solutions to accurately adapt to mission characteristics and time patterns, thus continuously constraining system operating efficiency.
[0022] If these issues are not addressed, the system will be unable to respond quickly to changes in resource demands during peak task periods, causing data processing delays and hindering the timely acquisition of data value. Simultaneously, during periods of resource idleness, the waste of computing resources will lead to an imbalance in cost control. This mismatch between resources and task requirements will continuously reduce the overall operational efficiency of the system and may adversely affect global climate change research and environmental protection decision-making based on carbon satellite data.
[0023] In this regard, such as Figure 1 As shown, Embodiment 1 provides a resource scheduling method based on a cloud environment, including:
[0024] Step Y100: Obtain real-time performance data and task queue information from the cloud environment to form raw feature data.
[0025] In one implementation, step Y100 may further include:
[0026] Step Y110: Collect real-time performance metrics and task queue length data from the cloud environment through the cloud monitoring module as raw monitoring data;
[0027] Step Y120: Extract the observed data flow and corresponding time information reflecting the load characteristics from the original monitoring data;
[0028] Step Y130: Combine the observed data flow with the time information, and perform structured processing on the combined data to generate raw feature data with time-series attributes, which will serve as input for the subsequent trend analysis stage.
[0029] Specifically, the solution of this invention comprehensively collects system performance indicators and specific business indicators in the carbon satellite application cloud scenario through a cloud monitoring module. On the one hand, the cloud monitoring module monitors the conventional performance indicators of the cloud servers or containers supporting the operation of the carbon satellite application in real time, including CPU utilization, memory utilization, network I / O, disk I / O, etc. On the other hand, considering the special nature of carbon satellite data processing, the cloud monitoring module also adds monitoring and observation of specific business indicators such as task queue length, data throughput, and task execution latency, ensuring the breadth and consistency of the original monitoring data. Subsequently, the observation data traffic reflecting the load characteristics and its corresponding time information are accurately extracted from these original monitoring data. This screening step ensures the relevance of subsequent analysis. Finally, the extracted observation data traffic and time information are effectively combined and structured to generate original feature data with time-series attributes. This processing flow transforms scattered and heterogeneous original data into a unified, ordered dataset with time correlation, providing high-quality, directly usable input for the subsequent trend analysis stage and providing data support for the implementation of elastic scaling. In this way, the present invention overcomes the challenge of raw data being messy and difficult to use directly for analysis, laying a solid foundation for accurately identifying resource usage patterns and predicting potential bottlenecks, thereby improving the accuracy and foresight of the entire resource scheduling method.
[0030] In a specific example, consider a cloud environment deployed for carbon satellite data processing. When the carbon satellite passes overhead, it generates terabytes of raw data in a short period, requiring complex processing such as radiometric calibration, atmospheric correction, and carbon dioxide inversion, leading to a surge in computational and I / O demands. First, a cloud monitoring module continuously collects data from this cloud environment, including CPU utilization, memory usage, network I / O, disk I / O, and the backlog length, data throughput, and processing latency of individual inversion tasks for each processing node (e.g., virtual machines or containers). This collected data constitutes the raw monitoring data. Next, from this raw monitoring data, the system extracts metrics directly related to the carbon satellite data processing load, such as the CPU utilization percentage of each processing node, storage I / O throughput, and the number of pending batches in the current task queue, recording the precise timestamps corresponding to these metrics. For example, at time T1, when carbon satellite data begins to be transmitted back, the CPU utilization of processing node A is 85%, storage I / O is 800 MB / s, and the queue length is 1000. At time T2, the peak of data processing, the CPU utilization reaches 92%, storage I / O surges to 1500 MB / s, and the queue length increases to 1500. Finally, this extracted load data and time information are combined and structured, for example, stored in a time-series database, forming a series of data points arranged in chronological order. Each data point contains a timestamp and corresponding metrics such as CPU, storage I / O, and queue length, thereby generating raw feature data with time-series attributes, which can be directly used as input for subsequent trend analysis models.
[0031] Based on the above analysis, this invention can systematically collect, filter, and structure real-time performance data and carbon satellite-specific task queue information in the carbon satellite application cloud environment. This ensures that high-quality, highly relevant, and time-series-based input data can be obtained in the subsequent trend analysis stage, providing a solid data foundation for resource demand prediction and the generation of elastic scaling strategies. This significantly improves the accuracy and reliability of resource usage pattern identification and potential bottleneck prediction. This refined data preparation process effectively avoids analytical biases caused by low-quality or inconsistent formats of raw data, making prediction-based resource scheduling decisions more accurate and efficient, thereby optimizing cloud resource utilization and the stability of carbon satellite data processing applications.
[0032] Step Y200: Perform trend analysis based on the original feature data to obtain resource usage patterns and potential bottleneck prediction results.
[0033] In one implementation, step Y200 may further include:
[0034] Step Y210: Perform trend analysis processing on the original feature data to extract resource use feature information that reflects the periodicity of resource use and task density information that reflects the distribution of tasks.
[0035] Step Y220: Use the resource usage feature information and the task density information as input for prediction;
[0036] Step Y230: Analyze the input data using a prediction model to obtain the evolution trend of node computing load and storage read / write frequency;
[0037] Step Y240: Determine whether the node's computational load exceeds the bottleneck determination threshold based on the evolution trend;
[0038] Step Y250: If the node's computational load exceeds the bottleneck determination threshold, it is determined to be a peak resource usage, and the potential bottleneck prediction result is generated.
[0039] Step Y260: Summarize the resource usage characteristic information to form the resource usage pattern.
[0040] Specifically, based on the acquisition of raw feature data, this invention performs in-depth trend analysis on this data to first extract resource usage characteristic information reflecting the cyclical patterns of resource use and task density information reflecting task distribution. During this process, the cloud monitoring module utilizes data analysis mechanisms and machine learning models to deeply mine and predict trends in the collected carbon satellite application performance and business indicators, thereby identifying system performance bottlenecks and potential risks, and recording scenarios such as cyclical patterns of resource use, sudden increases in resource usage, and performance bottleneck points. This information serves as input to the prediction model, enabling it to more comprehensively understand the system's historical behavior and the load characteristics of carbon satellite data processing tasks. The prediction model then analyzes these inputs to obtain the evolution trends of node computing load and storage read / write frequency. By comparing the predicted node computing load with preset bottleneck judgment thresholds, potential resource occupancy peaks can be accurately identified, and potential bottleneck prediction results can be generated. Simultaneously, by summarizing resource usage characteristic information, a more refined and comprehensive resource usage pattern is formed. These steps work together to enable resource scheduling to no longer rely solely on simple historical trends, but to more accurately predict resource demand and potential bottlenecks based on a deep understanding of periodic patterns, task distribution, and future evolution trends, thus providing a more reliable basis for subsequent resource scheduling strategy formulation.
[0041] In one implementation, the prediction model described above is an end-to-end time series prediction model built on LSTM. The specific construction and training process of this model includes:
[0042] The network structure is based on LSTM and includes an input layer, at least two sequentially connected LSTM network layers, an activation layer following the LSTM network layers, a batch normalization layer, a dropout layer, and a fully connected output layer.
[0043] The input layer is used to receive a carbon satellite data processing load feature sequence with a length of M time units. Each time unit in the feature sequence contains indicators reflecting the load status, such as CPU utilization, memory utilization, storage I / O throughput, and task queue length at that moment.
[0044] The at least two LSTM network layers are used to extract long-term dependency features from the load feature sequence, wherein the first LSTM layer receives the input sequence and generates a hidden state sequence, and the second LSTM layer receives the output of the first LSTM layer to further capture deeper temporal dependencies.
[0045] The activation layer contains the ReLU activation function, which acts on the output of the second LSTM layer to introduce nonlinear mapping capabilities into the network, enabling the model to learn more complex load evolution patterns.
[0046] The batch normalization layer is used to normalize the intermediate network features output by the activation layer in order to accelerate model convergence and improve training stability.
[0047] The dropout layer is used to randomly filter out some neurons during training with a preset probability. The preset probability ranges from 0.3 to 0.5. By randomly dropping out neurons, the model is prevented from overfitting to the training data, thereby improving the generalization ability under different load modes in the carbon satellite data processing system.
[0048] The fully connected output layer is used to map the temporal features extracted from the LSTM network layer and processed as described above into a predicted output sequence for the next N time units. The predicted output sequence contains the predicted trend values of the node computing load and storage read / write frequency for each future time unit.
[0049] During model training, the preprocessed data is divided into training, validation, and test sets according to time sequence. Data from earlier time periods is used as the training set, data from middle time periods as the validation set, and data from later time periods as the test set to ensure that future information is not introduced into the model evaluation. Sample-label pairs are constructed based on the training set data, where each sample is a sequence of carbon satellite data processing load features over M consecutive time units, and the corresponding label is the actual load data over the following N consecutive time units. The LSTM prediction model is initialized with an input dimension of M, a hidden layer size of 128, and an output dimension of N. Training parameters are set, including at least a learning rate of 0.001, 200 iterations, and a mean squared error loss function. These parameters are determined through experimental tuning to balance prediction accuracy and training time cost in the carbon satellite data processing load prediction task. Mean squared error is used as the loss function, and the Adam adaptive optimization algorithm is used to update the model parameters. The normalized training set data is input into the LSTM prediction model in batches for multiple rounds of iterative training. In each iteration cycle, the gradient is calculated by the backpropagation algorithm and the weights and bias parameters of each network layer are optimized and updated.
[0050] During model training, a validation set is introduced to supervise and verify the training process: after each iteration of training, samples constructed from the validation set data in the same manner are input into the current state of the LSTM prediction model for forward inference, and the mean squared error between the predicted output and the true label is calculated to obtain the validation set loss. The preset stopping condition is as follows: when the validation set loss decreases by less than a preset threshold within 20 consecutive iterations, training is considered complete, and the current model parameters are saved; when the validation set loss does not meet the stopping condition, training is considered incomplete, and training continues. The preset threshold is set based on the magnitude of the validation set loss during training and is set to 1e-4 of the current validation set loss value.
[0051] In a specific example, a cloud environment for carbon satellite data processing runs multiple microservices responsible for radiometric calibration, atmospheric correction, and inversion calculations. Its resource requirements are closely related to the satellite transit cycle and data return schedule. The system continuously collects real-time performance data for each processing instance, such as CPU utilization, memory usage, network bandwidth, and storage I / O, as well as specific business indicators such as queue length of data batches to be processed and single-batch processing latency. By performing time-series analysis on these raw characteristic data (e.g., using a seasonal ARIMA model), the system can discover a clear periodic pattern in the resource usage of carbon satellite data processing tasks: computation and storage I / O peak daily between 2:00 AM and 7:00 AM during centralized satellite data return; and the demand for computing resources surges at the beginning of each month due to the concentrated generation of monthly global carbon concentration distribution reports. Simultaneously, statistical analysis of the task queues reveals the density distribution of different processing stages (such as calibration, correction, and inversion) at different times. These constitute resource usage characteristic information and task density information. The aforementioned periodic patterns of CPU utilization, resource surge trends during report generation, and density distribution data during processing phases were integrated as input to the prediction model. An LSTM-based deep learning model was used, based on these inputs, to predict the evolution trends of CPU load on each processing node and I / O throughput of the storage server over the next 24 hours. The preset bottleneck threshold was set at 90% CPU utilization or storage I / O throughput exceeding 80% of the cluster's total bandwidth. The system compared the predicted load trend with this threshold. If the prediction showed that a processing node's CPU utilization would reach 95% during the next round of carbon satellite data transmission at 4:00 AM tomorrow, exceeding the 90% bottleneck threshold, the system would determine that this period as a peak resource usage time and generate a potential bottleneck prediction, indicating that the node might experience a computing capacity bottleneck at 4:00 AM tomorrow. Ultimately, the system will summarize all extracted resource usage characteristics, such as summarizing that the carbon satellite data processing platform experiences a continuous 3-hour computing peak after each satellite transit, that the monthly report generation task leads to a surge in storage I / O, and that the overall load is low during routine monitoring periods when there is no data processing, thus forming a detailed resource usage pattern report.
[0052] Based on the above analysis, this invention can extract more refined periodic patterns of resource usage and task density information from raw feature data, and use this information as input to the prediction model. This allows the prediction model to more accurately capture the evolution trends of node computing load and storage read / write frequency. This meticulous analysis and prediction mechanism, combined with bottleneck determination thresholds, can accurately identify potential resource occupancy peaks and form more comprehensive and reliable resource usage patterns. This significantly improves the predictability and accuracy of resource scheduling for future load changes in carbon satellite data processing tasks, effectively avoiding resource waste or performance bottlenecks caused by inaccurate predictions, thereby improving the efficiency and stability of resource scheduling in the carbon satellite application cloud environment.
[0053] Step Y300: Based on the resource usage patterns and the potential bottleneck prediction results, generate a resource demand prediction strategy file and a set of elastic scaling trigger conditions; wherein, the set of elastic scaling trigger conditions includes planning trigger rules based on time periods and trigger rules based on performance index thresholds.
[0054] In one implementation, step Y300 may further include:
[0055] Step Y310: Obtain baseline state data from the resource usage pattern, analyze the node load and storage read / write frequency, and compare them with the performance trigger threshold;
[0056] Step Y320: If the comparison result indicates that the current resource configuration needs to be adjusted, then an initial configuration item containing computing resources and storage resources is generated.
[0057] Step Y330: Based on the initial configuration items and the task cycle characteristics of the target application, generate the resource demand prediction strategy file;
[0058] Step Y340: Based on the periodic characteristics in the resource usage pattern and the potential bottleneck prediction results, define a time-cycle-based plan triggering rule;
[0059] Step Y350: Based on the relationship between the baseline state data and the performance trigger threshold, define trigger rules based on the performance index threshold;
[0060] Step Y360: Combine the planned triggering rule with the triggering rule based on performance index threshold to form the elastic scaling triggering condition set.
[0061] The above steps are executed through the policy management module. This module is the core of elastic scaling. On one hand, it analyzes historical load data and resource usage from the cloud monitoring module, combines this with machine learning models to generate predicted resource requirements, and sets corresponding resource requirement prediction strategies. On the other hand, it allows users to flexibly define elastic scaling strategies based on different business scenarios and needs of the carbon satellite application. These strategies cover various triggering conditions: threshold triggers based on system resource utilization, such as automatic scaling when CPU utilization exceeds a defined threshold; planned triggers based on time periods, such as scaling up in advance before the daily data processing peak; and event triggers based on specific business events, such as initiating scaling up when a new batch of carbon satellite data arrives. Furthermore, the policy management module also has policy version control, priority settings, and conflict detection functions to ensure the coordinated operation of multiple policies and avoid system instability caused by policy conflicts. The predicted resource requirements cover indicators such as the number of CPU cores, memory capacity, storage capacity, recommended storage performance values, and network bandwidth.
[0062] Specifically, after acquiring baseline state data from resource usage patterns, the strategy management module analyzes the compute node load and storage read / write frequency, comparing them with preset performance trigger thresholds to determine if the current resource configuration needs adjustment. If the comparison indicates adjustment is necessary, an initial configuration item containing compute and storage resources is generated. Based on this, and considering the mission cycle characteristics of the carbon satellite application, the initial configuration item is refined and optimized to generate a resource demand prediction strategy file, ensuring that the resource configuration fully meets the specific needs of carbon satellite data reception, processing, and analysis at different time periods. Simultaneously, based on the periodic characteristics of resource usage patterns and potential bottleneck prediction results, time-cycle-based planned triggering rules are defined to enable early response to predictable load changes; based on the relationship between baseline state data and performance trigger thresholds, performance index threshold-based triggering rules are defined to ensure the system can respond promptly to sudden or unpredictable load changes. Finally, the planned triggering rules and performance index threshold-based triggering rules are combined to form a comprehensive and flexible set of elastic scaling triggering conditions.
[0063] In a specific example, suppose a carbon satellite data processing platform has a mission cycle characterized by data transmission from 2:00 AM to 4:00 AM every day, followed by preliminary processing within 3 hours after transmission, resulting in a surge in computing and storage resource demands; at the same time, there is a monthly report generation task every Monday morning from 9:00 AM to 12:00 PM, causing a periodic increase in computing resource demands.
[0064] First, the system continuously collects CPU utilization, memory utilization, and storage (such as object storage) read / write IOPS data from its compute nodes (e.g., virtual machine instances) via a cloud monitoring module. The system analyzes this historical data, identifying an average CPU utilization of 20%, memory utilization of 30%, and storage IOPS of 500 during non-task processing periods; these constitute baseline state data. Simultaneously, performance trigger thresholds are set: CPU utilization exceeding 75% for 5 minutes, or storage IOPS exceeding 2000 for 1 minute. When the system predicts an upcoming large volume of satellite data transmission, or when real-time monitoring detects that CPU utilization has reached 60% and continues to rise, the comparison indicates that the current resource configuration may be insufficient. At this point, the system generates an initial configuration, such as suggesting adding 2 CPU cores and 8GB of memory, and increasing storage IOPS to 3000. Based on the mission cycle characteristics of the carbon satellite application, the system generates a resource demand prediction strategy file according to the initial configuration items mentioned above. This file explicitly states that additional computing resources (e.g., adding two high-configuration instances) and storage resources (e.g., increasing storage IOPS to 3000) are needed daily from 2:00 AM to 7:00 AM. Historical data analysis reveals that the carbon satellite data processing platform generates monthly reports every Monday from 9:00 AM to 12:00 PM, causing a periodic increase in computing resource demand. Simultaneously, potential bottleneck prediction indicates that storage IOPS may reach a bottleneck in the early morning of the 1st of each month. Therefore, planned triggering rules are defined: automatically add two computing instances every Monday at 8:30 AM; automatically increase storage IOPS to 5000 at 0:00 AM on the 1st of each month. Based on baseline status data and performance trigger thresholds, triggering rules based on performance metric thresholds are defined: automatically add one computing instance when the CPU utilization of a computing node exceeds 75% for five consecutive minutes; automatically increase storage throughput when storage IOPS exceeds 2000 for one consecutive minute. Finally, the time-based planned triggering rules and performance metric threshold-based triggering rules are combined to form a set of elastic scaling triggering conditions. For example, this set might include: Rule 1: Time-triggered, every Monday at 8:30 AM, Action: Add 2 compute instances; Rule 2: Time-triggered, the 1st of each month at 0:00 AM, Action: Increase storage IOPS to 5000; Rule 3: Performance-triggered, CPU utilization > 75% (5 minutes), Action: Add 1 compute instance; Rule 4: Performance-triggered, storage IOPS > 2000 (1 minute), Action: Increase storage throughput.
[0065] This step refines the process of resource demand forecasting and generating elastic scaling trigger conditions, enabling the strategy management module to generate highly customized resource demand forecasting strategy files and elastic scaling trigger condition sets based on the periodic characteristics and sudden demands of carbon satellite data processing. By defining time-cycle-based planned triggering rules and performance index threshold-based triggering rules, and effectively combining the two, an elastic scaling mechanism is constructed that can proactively respond to predictable load changes and quickly respond to sudden performance bottlenecks. This significantly improves the intelligence level of resource scheduling, effectively avoids the problems of resource allocation lag or over-configuration, thereby optimizing the utilization efficiency of cloud resources, reducing operating costs, and ensuring the stability and availability of carbon satellite application services.
[0066] Step Y400: Generate a resource configuration template based on the resource demand prediction strategy file, and generate an elastic strategy template with defined trigger rules based on the elastic scaling trigger condition set.
[0067] In one implementation, step Y400 may further include:
[0068] Step Y410: Extract peak time information, duration and resource demand specifications from the resource demand forecasting strategy file;
[0069] Step Y420: Based on the resource requirement specifications, match the pre-established hierarchical and classified resource configuration templates to generate an instantiated resource configuration template suitable for the current resource requirements;
[0070] Step Y430: Extract the time-period-based planned triggering rules and corresponding execution action definitions from the elastic scaling triggering condition set;
[0071] Step Y440: Combine the plan triggering rule with the execution action definition to generate a time-based flexible strategy template;
[0072] Step Y450: Extract the triggering rules based on performance index thresholds and the corresponding execution action definitions from the elastic scaling triggering condition set;
[0073] Step Y460: Combine the triggering rule based on the performance indicator threshold with the execution action definition to generate an elastic strategy template based on the performance indicator threshold.
[0074] The above steps are executed through the template management module. This module primarily manages resource configuration template definitions and elastic scaling policy template definitions. Regarding resource configuration templates, it defines tiered and categorized basic computing, storage, and network resources required for the carbon satellite application's operation, based on the actual application scenarios of elastic scaling of resources involved in the carbon satellite application, according to the policy management and scheduler modules. Based on the characteristics of the carbon satellite application's computing resource requirements, it specifically defines parameters such as CPU core count, memory capacity, instance specifications, instance quantity range, CPU priority, memory priority, IO priority, network bandwidth, volume type, and security rules. Regarding elastic scaling policy templates, it defines various elastic scaling trigger conditions and execution actions for the policy engine to call and execute according to rules to cope with changes in resource requirements under different carbon satellite application scenarios. For example, based on the periodic characteristics of carbon satellite data reception, processing, and analysis, a time-based elastic scaling policy template is set up to automatically increase the number of computing resource instances during the daily satellite data reception period to ensure timely data reception and preliminary processing, while reducing resource instances during off-peak periods to lower operating costs. Furthermore, the elastic strategy template based on performance metric thresholds automatically triggers scaling up or down operations when the performance metrics of the cloud monitoring system, such as CPU utilization, memory usage, and IO, exceed the set thresholds. Meanwhile, to adapt to the constantly changing and optimized needs of carbon satellite applications, the template management module has comprehensive version management capabilities. Each time a template is modified or updated, the system automatically generates a new version number and records a version change log, allowing users to easily review historical versions, compare differences, and select the appropriate template version for application.
[0075] Specifically, when generating resource configuration templates, the system first precisely extracts peak time information, duration, and resource requirement specifications from the resource demand forecasting strategy file. Based on the extracted resource requirement specifications, the system intelligently matches pre-established hierarchical resource configuration templates, transforming abstract resource requirements into instantiated resource configuration templates containing parameters such as the number of CPU cores, memory capacity, instance specifications, storage volume type, and network bandwidth, ensuring the standardization and efficiency of resource configuration. When generating elasticity strategy templates, the system extracts time-period-based planned triggering rules and performance indicator threshold-based triggering rules from the elastic scaling triggering condition set, along with their corresponding execution action definitions. By combining planned triggering rules with execution action definitions, a time-based elasticity strategy template is generated to address predictable periodic load changes in carbon satellite data processing tasks; by combining performance indicator threshold-based triggering rules with execution action definitions, a performance-based elasticity strategy template is generated to address sudden load changes. In this way, the complex elastic scaling logic is decomposed into independent strategy templates that are easier to manage and execute.
[0076] In a specific example, consider a carbon satellite data processing platform that needs to handle peak data transmission times and sudden increases in CPU utilization every morning. First, in step Y410, the system parses the "peak time information" as "daily 2:00-7:00," the "duration" as "5 hours," and the "resource requirement specifications" as "8 vCPUs, 32GB memory, high IOPS storage" from the resource requirement prediction strategy file. Next, in step Y420, based on these resource requirement specifications, the system matches a hierarchical resource configuration template named "Carbon Satellite Data Processing - Computationally Intensive" from a pre-defined template library and instantiates it, generating a resource configuration template containing a specific virtual machine image, instance specifications, security rules, and storage volume type. Simultaneously, in step Y430, the system identifies the planned trigger rule "expand to 4 processing instances daily at 1:30" and its corresponding execution action definition "adjust the expected number of instances in the scaling group to 4" from the elastic scaling trigger condition set. In step Y440, these two elements are combined into a time-based elastic strategy template for a "carbon satellite data backhaul peak expansion strategy," version v2.1, with changelogs recording optimizations and adjustments to instance specifications. Furthermore, in step Y450, the system identifies a performance metric threshold trigger rule from the elastic scaling trigger condition set: "expand by one instance when CPU utilization exceeds 85% for 5 consecutive minutes," along with its "add one instance" execution action definition. In step Y460, these two elements are combined into a performance metric threshold-based elastic strategy template for a "CPU high load emergency expansion strategy." During elastic scaling operations, the system automatically selects the latest and compatible template version based on application requirements and strategies, ensuring the system always operates under optimal configuration.
[0077] This step combines the hierarchical resource configuration capabilities of the template management module with the ability to define multiple types of elastic policy templates, enabling the refined and automated generation of resource configuration templates and elastic policy templates. By accurately extracting resource requirement specifications and matching them with instantiated hierarchical resource configuration templates, it ensures that the generated resource configurations precisely meet the needs of carbon satellite data processing applications, avoiding over-allocation or under-allocation of resources. Simultaneously, by subdividing the elastic scaling trigger condition set into time-cycle-based planned trigger rules and performance indicator threshold-based trigger rules, and generating corresponding elastic policy templates for each, the system can more flexibly and intelligently respond to periodic load changes and sudden performance bottlenecks in carbon satellite application scenarios. The introduction of template version management further enhances the maintainability and reliability of the system during continuous requirement changes and optimization processes, enabling the entire resource scheduling method to continuously adapt to changes in business needs during long-term operation, thereby significantly improving cloud resource utilization, system stability, and response speed.
[0078] Step Y500: Based on the triggering rules in the elastic strategy template and the resource parameters in the resource configuration template, formulate an elastic scaling execution plan and issue resource control instructions.
[0079] In one implementation, step Y500 may further include:
[0080] Step Y510: Retrieve the elastic strategy template and the resource configuration template from the template management module, parse the triggering rules and resource parameters therein, and convert them into an executable instruction set; wherein, the triggering rules define the triggering conditions for elastic scaling;
[0081] Step Y520: Based on the triggering conditions in the instruction set, monitor the real-time performance data stream in the cloud environment and determine whether the real-time performance data stream meets the triggering conditions.
[0082] Step Y530: If the real-time performance data stream meets the triggering condition, then formulate an elastic scaling execution plan based on the resource parameters in the instruction set.
[0083] Step Y540: The elastic scaling execution plan is sent to the scheduler, which matches the corresponding scaling group identifier and issues resource control instructions.
[0084] The above steps are completed collaboratively by the execution engine and the scheduler module. The execution engine is primarily responsible for executing the elastic policy template and resource configuration template, and providing feedback on the execution process. When the cloud monitoring module detects that the system status meets the preset trigger conditions in the elastic policy template, it will transmit a trigger signal to the execution engine. Upon receiving the signal, the execution engine retrieves the corresponding elastic policy template and resource configuration template from the template management module. Through deep parsing of the template, it transforms the defined logical rules and parameters into a set of instructions that the computer can recognize and execute, and makes an elastic scaling execution plan accordingly. Based on the execution plan, the execution engine sends it to the scheduler. The scheduler, based on various scheduling strategies such as timed scheduling, periodic scheduling, and performance indicator threshold scheduling, matches the corresponding scaling group identifier and issues specific resource control instructions to the cloud platform to realize the scaling of cloud server instances or container POD instances within the scaling group. During execution, the execution engine monitors the operation progress and system status in real time. If, during scaling, it detects that instance creation is slow due to resource constraints on the cloud platform, the execution engine will adjust the scheduling strategy based on the resource prediction mechanism, prioritizing the acquisition of resources from other available resource pools to ensure the success rate and timeliness of the scaling operation. This approach, which combines template parsing, real-time monitoring, scheduling strategy selection, and execution monitoring, constitutes a closed-loop automated scheduling process.
[0085] In a specific instance, a carbon satellite data processing platform experiences peak data transmission and processing from 2:00 AM to 7:00 AM daily. The system is configured with a time-based elastic policy template, "Carbon Satellite Peak Expansion Policy," and a corresponding resource configuration template. When the preset trigger point of 1:55 AM is reached, the cloud monitoring module transmits a trigger signal to the execution engine. The execution engine retrieves the aforementioned policy and resource configuration templates from the template management module, parses them, and converts them into an instruction set: It checks the number of instances in the current scaling group; if it is less than the target value of 4, it executes expansion to 4 "Carbon Satellite Computing Instances" (specifications: 8 vCPUs, 32GB memory, high IOPS storage). The execution engine sends this execution plan to the scheduler, which matches the scaling group identifier "co2-process-asg" according to the plan and adjusts the desired capacity to 4 via the cloud platform API. During the expansion process, the execution engine continuously tracks the creation status of new instances and monitors system load to ensure that new instances successfully join and begin processing tasks. Similarly, for performance-triggered scenarios, when the CPU utilization of a certain processing instance is detected to exceed 85% for 5 consecutive minutes, the execution engine will call the "CPU High Load Emergency Expansion Template" to generate an execution plan to expand by 1 instance and hand it over to the scheduler for execution.
[0086] This step effectively solves the problem of directly implementing elastic scaling strategies in carbon satellite application cloud scenarios by refining the abstract elastic strategy template and resource configuration template, and combining it with dynamic monitoring of real-time performance data and various scheduling strategies of the scheduler. The execution engine's real-time monitoring of execution progress and its ability to predict and adjust resources improve operational robustness and success rate, significantly enhancing the system's response speed and accuracy to load changes. This ensures that carbon satellite data processing tasks receive sufficient computing resources during peak periods while avoiding resource waste during off-peak periods, thereby optimizing the availability and cost-effectiveness of cloud services.
[0087] Step Y600: Monitor the execution status of the resource control command, identify service endpoints with abnormal availability through a health check mechanism, and adjust resource allocation in combination with resource prediction and scheduling logic; at the same time, feed back the adjusted resource allocation status to the original feature data collection stage.
[0088] In one implementation, step Y600 may further include:
[0089] Step Y610: Monitor the execution status of the resource control command and perform availability detection on the service endpoint through a health check mechanism;
[0090] Step Y620: Based on the detection results, filter service endpoints with abnormal availability status and generate resource takeover or replacement plans for the abnormal endpoints.
[0091] Step Y630: Based on the resource prediction and scheduling logic, analyze the time cycle characteristics of task execution and the current load demand, and generate a resource pre-scheduling plan;
[0092] Step Y640: Combine the resource takeover or replacement scheme with the resource pre-adjustment scheme to adjust the computing capacity allocation and resource allocation status;
[0093] Step Y650 involves feeding back the adjusted resource allocation status to the cloud monitoring module to update the basic information of the original feature data collection process.
[0094] The above steps are completed collaboratively by the scheduler module and the execution engine module. After the resource control command is issued, the scheduler module immediately begins monitoring the command execution status and performs availability checks on the service endpoints using its built-in health check mechanism. This health check capability periodically checks the service endpoint availability using HTTP / TCP probes and also has the ability to customize test scripts to achieve deep availability testing of specific processing logic for the carbon satellite application. During execution, the execution engine module monitors the operation progress and system status in real time to ensure the efficiency and robustness of elastic scaling operations.
[0095] In one implementation, in steps Y610 to Y620, once a fault is confirmed based on the number of failed detections or abnormal status codes, the scheduler automatically triggers removal or isolation capabilities, removing or isolating the faulty instance from the service backend, and generating a resource takeover or replacement plan according to a preset strategy, initiating a new instance to take over traffic to minimize service interruption time. In step Y630, by analyzing the time-cycle characteristics of the carbon satellite data processing task and the current load requirements, a resource pre-tuning plan is generated to ensure that resources are ready in advance during sudden business surges. In steps Y640 to Y650, after the elastic scaling operation is completed, the execution engine performs a comprehensive analysis and feedback of the execution results, feeding back the operation results (such as successfully scaling up N instances or successfully releasing M instances) to the cloud monitoring module, enabling the cloud monitoring module to update system resource status information, providing accurate data for subsequent monitoring and policy triggering, forming a closed loop of continuous optimization.
[0096] In a specific instance, a carbon satellite data processing platform is performing daily data backhaul processing tasks at dawn. The system has issued a capacity expansion command based on a predictive strategy, increasing the number of processing instances. After the expansion command is issued, the scheduler immediately begins monitoring the execution status of the newly added processing instances. Through health checks, such as sending HTTP probe requests to the service ports of the processing instances every 10 seconds, it verifies whether they can respond normally to the distribution of data processing tasks. During one probe, the scheduler detects a newly started processing instance continuously returning error status codes, determining its availability to be abnormal. The scheduler immediately identifies this abnormal instance and generates a resource takeover or replacement plan: first, it removes the abnormal instance from the task distribution list, stops allocating new data batches to it, and simultaneously triggers the startup of a new processing instance to take over its unfinished tasks, ensuring that the ongoing carbon satellite data inversion calculations are not interrupted. At the same time, based on its resource prediction scheduling logic, the scheduler analyzes that the next batch of carbon satellite observation data will arrive in 30 minutes, and the current real-time load is trending upwards. The system generates a resource pre-adjustment plan, adding additional processing instances in advance. The execution engine integrates replacement solutions for abnormal instances and pre-tuning plans for handling new data, uniformly adjusting computing capacity to ensure sufficient total cluster capacity and monitoring the entire operation progress. Finally, after the elastic scaling operation is complete, the execution engine feeds back the adjusted resource allocation status (including the total number of currently processed instances and the configuration of each instance) to the cloud monitoring module. Based on this, the cloud monitoring module updates its internal resource topology and performance metric collection baseline, ensuring that subsequent trend analysis and resource prediction for the next round of data processing are based on the latest system status.
[0097] Based on the above analysis, this step achieves closed-loop management of the execution status of resource control commands by combining health checks of the scheduler module, multiple scheduling strategies, and the real-time monitoring and feedback capabilities of the execution engine. On the one hand, through proactive detection methods such as HTTP / TCP probes and custom probing scripts, it can promptly detect and automatically isolate and replace service endpoints with abnormal availability, ensuring the continuity of carbon satellite data processing services. On the other hand, combined with resource prediction scheduling logic, it can proactively perform resource preheating and expansion while recovering from anomalies. The execution engine's real-time monitoring of operation progress and comprehensive analysis and feedback of execution results constitute a closed loop of continuous optimization, ensuring that subsequent scheduling decisions are based on the latest and most accurate system information. This significantly improves the accuracy, robustness, and adaptability of resource scheduling in the carbon satellite application cloud environment, effectively reducing operation and maintenance costs and ensuring the stability of data processing tasks.
[0098] Example 2, Figure 2 This is a schematic diagram of a cloud-based resource scheduling system according to Embodiment 2 of the present invention, as shown below. Figure 2As shown, Embodiment 2 provides a cloud-based resource scheduling system based on the cloud-based resource scheduling method described in Embodiment 1. The system includes: a raw feature data composition module 201, an acquisition module 202, a first generation module 203, a second generation module 204, a control command issuance module 205, and a resource allocation adjustment module 206. The raw feature data composition module 201 acquires real-time performance data and task queue information from the cloud environment to form raw feature data. The acquisition module 202 performs trend analysis based on the raw feature data to obtain resource usage patterns and potential bottleneck prediction results. The first generation module 203 generates a resource demand prediction strategy file and a set of elastic scaling trigger conditions based on the resource usage patterns and potential bottleneck prediction results; wherein the elastic scaling trigger condition set includes time-cycle-based planned trigger rules and performance index threshold-based trigger rules. The second generation module 204 generates a resource configuration template based on the resource demand prediction strategy file and generates an elastic strategy template with defined trigger rules based on the elastic scaling trigger condition set. The control command issuance module 205 is used to formulate an elastic scaling execution plan and issue resource control commands based on the triggering rules in the elastic strategy template and the resource parameters in the resource configuration template. The resource allocation adjustment module 206 is used to monitor the execution status of the resource control commands, identify service endpoints with abnormal availability through a health check mechanism, and adjust resource allocation in conjunction with resource prediction and scheduling logic; at the same time, it feeds back the adjusted resource allocation status to the original feature data acquisition stage.
[0099] The various variations and specific examples of the cloud-based resource scheduling method provided in Embodiment 1 are also applicable to the cloud-based resource scheduling system provided in this embodiment. Through the foregoing detailed description of a cloud-based resource scheduling method, those skilled in the art can clearly understand the implementation method of the cloud-based resource scheduling system in this embodiment. Therefore, for the sake of brevity, it will not be described in detail here.
[0100] Example 3, Figure 3 This is a schematic diagram of the structure of an electronic device according to Embodiment 3 of the present invention, as shown below. Figure 3 As shown, Embodiment 3 also provides an electronic device 300, which may include a processor 301 and a memory 302.
[0101] Memory 302 is used to store programs. Memory 302 may include volatile memory, such as random-access memory (RAM), such as static random-access memory (SRAM), double data rate synchronous dynamic random-access memory (DDR SDRAM), etc.; memory may also include non-volatile memory, such as flash memory. Memory 302 is used to store computer programs (such as application programs, functional modules, etc. that implement the above methods), computer instructions, etc. The computer programs, computer instructions, etc., can be partitioned and stored in one or more memories 302. Furthermore, the computer programs, computer instructions, data, etc., can be accessed by processor 301.
[0102] The aforementioned computer programs and instructions can be stored in one or more partitions of memory 302. Furthermore, the aforementioned computer programs and instructions can be invoked by processor 301.
[0103] The processor 301 is configured to execute the computer program stored in the memory 302 to implement the various steps in the methods described in the above embodiments.
[0104] For details, please refer to the relevant descriptions in the preceding method embodiments.
[0105] The processor 301 and the memory 302 can be independent structures or integrated structures. When the processor 301 and the memory 302 are independent structures, the memory 302 and the processor 301 can be coupled together via bus 303.
[0106] The electronic device in this embodiment can execute the technical solution in the above method. Its specific implementation process and technical principle are the same, and will not be repeated here.
[0107] Example 4: Example 4 also provides a computer-readable storage medium including a computer program and instructions, which, when executed on a computer, cause the computer to perform the cloud-based resource scheduling method of any embodiment of the present invention.
[0108] Computer-readable storage media include various media that can store program code, such as USB flash drives, external hard drives, ROM, RAM, magnetic disks, or optical disks.
[0109] This embodiment also provides a computer program product, which includes: a computer program stored in a readable storage medium, at least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to cause the electronic device to perform the solution provided in any of the above embodiments.
[0110] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this invention disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.
[0111] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A resource scheduling method based on a cloud environment, characterized in that, include: Real-time performance data and task queue information are obtained from the cloud environment to form raw feature data; Based on the original feature data, trend analysis is performed to obtain resource usage patterns and potential bottleneck prediction results. Based on the resource usage patterns and the potential bottleneck prediction results, a resource demand prediction strategy file and a set of elastic scaling trigger conditions are generated; wherein, the set of elastic scaling trigger conditions includes planning trigger rules based on time periods and trigger rules based on performance index thresholds. A resource configuration template is generated based on the resource demand prediction strategy file, and an elastic strategy template with defined trigger rules is generated based on the elastic scaling trigger condition set. Based on the triggering rules in the elastic strategy template and the resource parameters in the resource configuration template, an elastic scaling execution plan is formulated and resource control instructions are issued. The system monitors the execution status of the resource control commands, identifies service endpoints with abnormal availability through a health check mechanism, and adjusts resource allocation based on resource prediction and scheduling logic. Simultaneously, it feeds back the adjusted resource allocation status to the original feature data collection stage.
2. The resource scheduling method based on a cloud environment as described in claim 1, characterized in that, The raw feature data, which consists of real-time performance data and task queue information obtained from the cloud environment, includes: The cloud monitoring module collects real-time performance metrics and task queue length data from the cloud environment as raw monitoring data. Extract the observed data flow and corresponding time information reflecting the load characteristics from the original monitoring data; The observed data flow is combined with the time information, and the combined data is structured to generate raw feature data with time-series attributes, which serves as input for the subsequent trend analysis stage.
3. The resource scheduling method based on a cloud environment as described in claim 1, characterized in that, The trend analysis based on the original feature data to obtain resource usage patterns and potential bottleneck prediction results includes: The original feature data is subjected to trend analysis to extract resource use feature information reflecting the periodicity of resource use and task density information reflecting the distribution of tasks. The resource usage characteristics and task density information are used as inputs for prediction. The input data is analyzed using a predictive model to obtain the evolution trend of node computing load and storage read / write frequency; Based on the evolution trend, determine whether the node's computing load exceeds the bottleneck determination threshold; If the computing load of the node exceeds the bottleneck determination threshold, it is determined to be a peak resource usage, and the potential bottleneck prediction result is generated. The resource usage characteristics information is summarized to form the resource usage pattern.
4. The resource scheduling method based on a cloud environment as described in claim 1, characterized in that, Based on resource usage patterns and the predicted potential bottlenecks, a resource demand forecasting strategy file and a set of elastic scaling triggering conditions are generated. The set of elastic scaling triggering conditions includes time-period-based planned triggering rules and performance indicator threshold-based triggering rules, including: Obtain baseline state data from the resource usage patterns, analyze and compute node load and storage read / write frequency, and compare them with performance trigger thresholds; If the comparison results indicate that the current resource configuration needs to be adjusted, an initial configuration item containing computing and storage resources will be generated. Based on the initial configuration items and the task cycle characteristics of the target application, the resource demand prediction strategy file is generated. Based on the periodic characteristics in the resource usage patterns and the potential bottleneck prediction results, a time-cycle-based plan triggering rule is defined. Based on the relationship between the baseline state data and the performance trigger threshold, a triggering rule based on the performance index threshold is defined; The planned triggering rules are combined with the triggering rules based on performance index thresholds to form the elastic scaling triggering condition set.
5. The resource scheduling method based on a cloud environment as described in claim 1, characterized in that, The process of generating a resource configuration template based on the resource demand forecasting strategy file and generating an elastic strategy template with defined trigger rules based on the elastic scaling trigger condition set includes: Extract peak time information, duration, and resource demand specifications from the resource demand forecasting strategy file. Based on the resource requirement specifications, match the pre-established hierarchical and classified resource configuration templates to generate an instantiated resource configuration template suitable for the current resource requirements; Extract time-period-based planned triggering rules and corresponding execution action definitions from the set of elastic scaling triggering conditions; The plan triggering rules are combined with the execution action definitions to generate a time-based flexible strategy template; Extract the triggering rules and corresponding execution action definitions based on performance index thresholds from the set of elastic scaling triggering conditions; The triggering rules based on performance metric thresholds are combined with the execution action definitions to generate a flexible strategy template based on performance metric thresholds.
6. The resource scheduling method based on a cloud environment as described in claim 1, characterized in that, The step of formulating an elastic scaling execution plan and issuing resource control instructions based on the triggering rules in the elastic policy template and the resource parameters in the resource configuration template includes: The elastic policy template and the resource configuration template are retrieved from the template management module, and the triggering rules and resource parameters therein are parsed and converted into an executable instruction set; wherein, the triggering rules define the triggering conditions for elastic scaling; Based on the triggering conditions in the instruction set, monitor the real-time performance data stream in the cloud environment and determine whether the real-time performance data stream meets the triggering conditions; If the real-time performance data stream meets the triggering condition, then an elastic scaling execution plan is formulated based on the resource parameters in the instruction set; The elastic scaling execution plan is sent to the scheduler, which matches the corresponding scaling group identifier and issues resource control instructions.
7. The resource scheduling method based on a cloud environment as described in claim 1, characterized in that, The execution status of the monitoring resource control commands is used to identify service endpoints with abnormal availability through a health check mechanism, and resource allocation is adjusted in combination with resource prediction and scheduling logic. Simultaneously, the adjusted resource allocation status will be fed back to the original feature data collection stage, including: Monitor the execution status of the resource control commands and perform availability detection on the service endpoints through a health check mechanism; Based on the detection results, filter service endpoints with abnormal availability status and generate resource takeover or replacement plans for the abnormal endpoints; Based on the resource prediction and scheduling logic, analyze the time cycle characteristics of task execution and the current load demand to generate a resource pre-scheduling scheme. By combining the resource takeover or replacement scheme with the resource pre-adjustment scheme, the computing capacity allocation and resource allocation status are adjusted; The adjusted resource allocation status is fed back to the cloud monitoring module to update the basic information of the original feature data collection process.
8. A cloud-based resource scheduling system, based on the cloud-based resource scheduling method as described in any one of claims 1-7, characterized in that, The system includes: The raw feature data composition module is used to obtain real-time performance data and task queue information from the cloud environment to form raw feature data. The module is used to perform trend analysis based on the original feature data to obtain resource usage patterns and potential bottleneck prediction results. The first generation module is used to generate a resource demand forecasting strategy file and a set of elastic scaling triggering conditions based on the resource usage patterns and the potential bottleneck prediction results; wherein, the set of elastic scaling triggering conditions includes planning triggering rules based on time periods and triggering rules based on performance index thresholds. The second generation module is used to generate a resource configuration template based on the resource demand prediction strategy file, and to generate an elastic strategy template with defined trigger rules based on the elastic scaling trigger condition set. The control command issuance module is used to formulate an elastic scaling execution plan and issue resource control commands based on the triggering rules in the elastic strategy template and the resource parameters in the resource configuration template. The resource allocation adjustment module is used to monitor the execution status of the resource control commands, identify service endpoints with abnormal availability through a health check mechanism, and adjust resource allocation in combination with resource prediction and scheduling logic; at the same time, the adjusted resource allocation status is fed back to the original feature data collection stage.
9. An electronic device, characterized in that, include: At least one processor; as well as A memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the cloud-based resource scheduling method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, It includes computer programs and instructions that, when run on a computer, cause the computer to perform the resource scheduling method based on a cloud environment as described in any one of claims 1-7.