Intelligent computing power and storage scheduling method and system of multi-service system

By combining real-time monitoring data feature extraction and timing prediction model, combined with reinforcement learning dynamic correction prediction results, a dynamic resource requirement table and a cross-service priority weight matrix are generated, and resource allocation and data migration is used to use hybrid scheduling algorithms and storage scheduling engines, the problem of low resource scheduling efficiency in multi-service systems is solved, and efficient and stable resource utilization is achieved.

CN120144260AActive Publication Date: 2025-06-13GOLDEN TIMES CULTURE COMM

Patent Information

Application Number
CN202510571228.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-13
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

In a multi-service system environment, it is difficult for the existing technology to accurately predict the future computing power and storage needs of the system, resulting in low resource scheduling efficiency and prone to resource overload or waste.

Method used

Real-time monitoring data feature extraction is used to generate structured feature vectors, combine the timing prediction model to predict future computing power and storage needs, and through reinforcement learning, dynamic correction of prediction results, dynamic resource requirements tables and cross-business priority weight matrix are generated, and resource allocation and data migration are used to use hybrid scheduling algorithms and storage scheduling engines.

Benefits of technology

It realizes intelligent scheduling of multi-service system resources, improves resource utilization, reduces system bottlenecks, ensures the efficiency of task execution and system stability, and continuously improves scheduling decisions through adaptive optimization strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144260A_ABST
    Figure CN120144260A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent computing power and storage scheduling method and system for a multi-service system, and belongs to the technical field of intelligent computing resource scheduling and storage management, and the method comprises the steps: collecting the real-time monitoring data of each node in the multi-service system, and carrying out the feature extraction; predicting a computing power demand value and a storage demand layering strategy in a future preset time window; dynamically correcting the prediction result, and generating a dynamic resource demand table and a cross-service priority weight matrix; generating a task allocation scheme through a hybrid scheduling algorithm; generating a storage data migration instruction through a storage scheduling engine; executing calculation node capacity expansion, task migration and storage data redistribution operation; and based on the executed latest state snapshot of the resource pool and the scheduling failure case library, updating a weight parameter of the time sequence prediction model and a scheduling strategy rule library to obtain an updated strategy version, and applying the updated strategy version to a next scheduling period. According to the invention, efficient scheduling of resources can be realized in a multi-service system environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of intelligent computing resource scheduling and storage management, and particularly relates to an intelligent computing power and storage scheduling method and system for a multi-service system. Background Art

[0002] With the continuous development of information technology, technologies such as cloud computing, big data, high-performance computing, and virtualization have gradually become the core components of modern computing architectures. Especially in multi-service systems, as computing tasks become increasingly complex and business requirements become more and more diverse, how to effectively manage system resources has become the key to improving overall performance and efficiency. Resource scheduling and management, as the basic components in these systems, their design and optimization directly affect the execution efficiency of tasks, the stability of the system, and the utilization rate of resources.

[0003] Multi-service systems usually run multiple types of tasks simultaneously. These tasks not only have different requirements for computing power but also show significant differences in the use of storage resources. Computationally intensive tasks often require high computing power, while storage-intensive tasks have high requirements for storage I / O throughput. In addition to the differences in task types, different tasks also have different priorities and resource consumption patterns, which makes the resource scheduling problem more complex. The system needs to be able to perceive these changes in real time and make corresponding adjustments to ensure that high-priority tasks are processed in a timely manner while avoiding low-priority tasks from occupying too many resources.

[0004] Currently, in common resource scheduling methods, many systems rely on static configurations or rule-based scheduling strategies. For example, computing nodes may be assigned to fixed tasks, and storage resources may also be allocated according to preset rules. However, these methods are often difficult to effectively handle complex business requirements and dynamically changing resource usage situations. When the system load changes, or the resource requirements of tasks exceed expectations, these static scheduling methods will lead to resource overload or waste, thereby affecting the overall performance and efficiency of the system. For example, excessive load on computing nodes may cause task execution delays, while uneven allocation of storage resources may lead to overuse or idleness of some storage media, wasting valuable hardware resources.

[0005] Therefore, how to accurately predict the future computing power and storage requirements of the system in order to achieve efficient resource scheduling in a multi-service system environment is an urgent problem to be solved at present. Summary of the Invention

[0006] In order to achieve efficient resource scheduling in a multi-service system environment, this application provides an intelligent computing power and storage scheduling method and system for a multi-service system.

[0007] In a first aspect, the present application provides an intelligent computing power and storage scheduling method for a multi-service system, adopting the following technical solutions: An intelligent computing power and storage scheduling method for a multi-service system, the method comprising: Collect real-time monitoring data of each node in the multi-service system; Extract features from the real-time monitoring data to generate a structured feature vector including business type labels, task priorities, and data access patterns, obtaining a real-time service feature data set; Based on the real-time service feature data set and historical load logs, predict the computing power demand value and storage demand stratification strategy within a preset future time window through a time series prediction model; According to the computing power demand value and storage demand stratification strategy, dynamically correct the prediction results in combination with a reinforcement learning algorithm to generate a dynamic resource demand table and a cross-service priority weight matrix; Initialize preset rules in the scheduling policy rule library based on the business type; Based on the preset rules and resource pool topology information, according to the dynamic resource demand table and cross-service priority weight matrix, generate a task allocation plan through a hybrid scheduling algorithm; Based on the storage demand stratification strategy, generate a storage data migration instruction through a storage scheduling engine; Send the task allocation plan and the storage data migration instruction to the virtualized resource orchestration module to perform operations such as computing node expansion, task migration, and storage data redistribution; Collect the resource pool status data after execution to generate a latest status snapshot of the resource pool; Based on the latest status snapshot of the resource pool and the scheduling failure case library, update the weight parameters of the time series prediction model and the scheduling policy rule library to obtain an updated policy version and apply it to the next scheduling cycle.

[0008] By adopting the above technical solutions, the intelligent scheduling of multi-service system resources is realized. By deeply combining real-time monitoring data with prediction algorithms, the system can dynamically adjust resource allocation, improve resource utilization, reduce system bottlenecks, and at the same time ensure the high efficiency of task execution and the stability of the system. In addition, through an adaptive optimization strategy, the system can continuously improve scheduling decisions, making the scheduling effect more accurate and efficient after long-term operation, thereby further improving the performance and reliability of the entire multi-service system.

[0009] Optionally, the step of dynamically correcting the prediction results in combination with a reinforcement learning algorithm according to the computing power demand value and storage demand stratification strategy to generate a dynamic resource demand table and a cross-service priority weight matrix includes: Obtain real-time resource pool status data, and receive the computing power demand value and storage demand stratification strategy within a future preset time window output by the time series prediction model; Merge the computing power demand value, storage demand stratification strategy with the real-time resource pool status data into a composite state vector, and perform normalization processing to generate a standardized state matrix; Define action space parameters according to the business type, and input the action space parameters into a pre-configured reinforcement learning agent; wherein, the action space parameters include resource allocation ratio adjustment instructions and storage migration trigger instructions; Based on the standardized state matrix and historical scheduling records, calculate resource allocation ratio adjustment instructions and storage migration trigger instructions through the reinforcement learning agent; wherein, the reinforcement learning agent evaluates the action value according to a preset reward function and updates the policy network parameters; According to the resource allocation ratio adjustment instructions, combined with a preset business feature library, generate a dynamic resource demand table; Based on the dynamic resource demand table and storage migration trigger instructions, construct a cross-business priority weight matrix.

[0010] By adopting the above technical solution, the dynamic resource scheduling and priority decision-making method based on reinforcement learning can automatically and intelligently handle the resource allocation problem in a multi-business system. By combining real-time resource pool status, time series prediction data, and reinforcement learning algorithms, the system can flexibly adapt to changes in business load and make efficient and accurate resource scheduling decisions.

[0011] Optionally, after the step of constructing a cross-business priority weight matrix based on the dynamic resource demand table and storage migration trigger instructions, it further includes: Input the dynamic resource demand table and cross-business priority weight matrix into a digital twin simulation environment to generate a simulated scheduling result; According to the error rate between the actual scheduling result and the simulated scheduling result, trigger the incremental learning module to update the policy network parameters of the reinforcement learning agent.

[0012] By adopting the above technical solution, the dynamic resource demand table and the cross-service priority weight matrix are input into the digital twin simulation environment for simulation, and the incremental learning module is triggered according to the error rate between the actual scheduling result and the simulated scheduling result to optimize the strategy, further strengthening the intelligent scheduling ability of the system in a multi-service environment. The use of digital twin simulation enables the system to test and verify various scheduling strategies in a virtual environment, thus avoiding unforeseen errors in the production environment. Through the incremental learning module, the system can continuously optimize the decision-making strategy according to new data during the actual operation process, ensuring that resource scheduling can always meet the requirements of high efficiency and flexibility in various complex business scenarios. The overall technical solution can significantly improve resource utilization, reduce task latency, enhance system stability and overall performance, and adapt to the changing multi-service environment.

[0013] Optionally, the preset rules in the scheduling strategy rule library initialized based on the service type include: When the service type is a real-time task, configure a preemptive scheduling rule and a storage bandwidth reservation threshold; or, when the service type is an offline task, configure an elastic resource pool allocation rule and a cold data storage migration strategy.

[0014] By adopting the above technical solution, combined with the different requirements of real-time tasks and offline tasks, customized scheduling rules are formulated for each task type. Real-time tasks adopt preemptive scheduling rules and bandwidth reservation strategies to ensure their priority execution under high load; offline tasks optimize resource utilization through elastic resource pools and cold data migration strategies. The system's scheduling strategy library is not only initialized according to business requirements at the initial stage, but also dynamically adjusted according to feedback information in each subsequent scheduling cycle to ensure that the strategy always meets the current resource and business requirements.

[0015] Optionally, based on the preset rules and resource pool topology information, the steps of generating a task allocation plan according to the dynamic resource demand table and the cross-service priority weight matrix through a hybrid scheduling algorithm include: Obtain the dynamic resource demand table, the cross-service priority weight matrix, the resource pool topology information, and the preset rules; Perform time window slicing and normalization processing on the dynamic resource demand table to generate a normalized resource demand time series slice; Construct a weighted directed graph based on the resource pool topology information to generate a topology graph structure file; Classify and weight the normalized resource demand time series slices according to the preset rules; Combine the cross-service priority weight matrix and the topology graph structure file to calculate the matching degree score between tasks and nodes, and generate a candidate node score table and a rule conflict flag list; Based on the candidate node scoring table and the rule conflict flag list, a final task allocation scheme is generated through a heuristic algorithm.

[0016] By adopting the above technical solution, the system can efficiently allocate tasks and resources under the circumstances of dynamic resource requirements and changing business priorities, optimize the system performance, ultimately achieve precise task scheduling and resource utilization, and ensure the efficient coordination and dynamic adaptation of multi-dimensional resources.

[0017] Optionally, the steps of generating a final task allocation scheme through a heuristic algorithm based on the candidate node scoring table and the rule conflict flag list include: Obtain the candidate node scoring table, the rule conflict flag list, and the resource pool topology information; According to the candidate node scoring table, screen the whitelist of assignable nodes for each task, and exclude the illegal node combinations in the rule conflict flag list; Construct an initial population set based on the whitelist of assignable nodes; wherein, each population individual uses chromosome encoding to represent the task allocation scheme; Based on the communication delay parameters and node load data in the candidate node scoring table and the resource pool topology information, construct a multi-objective fitness function; Dynamically adjust the weight coefficients of each optimization objective in the multi-objective fitness function according to the real-time resource pool status data; Perform a tournament selection operation on the initial population set to generate a parental population; Based on the communication path table in the resource pool topology information, perform a topology-aware crossover operation on the parental chromosomes in the parental population to generate an offspring population; Dynamically adjust the mutation rate according to the population diversity evaluation result, and perform probability mutation based on the candidate node scoring table on the offspring chromosomes in the offspring population; Perform local search optimization on the chromosomes with fitness scores higher than the preset threshold; Screen the individuals with the optimal fitness from the parental population and the offspring population to form an elite population; When the convergence condition is met, decode the chromosome with the highest fitness in the elite population into the final task allocation scheme.

[0018] By adopting the above technical solution, based on the combination of a heuristic algorithm and a genetic algorithm, the global optimization of the task allocation scheme is realized, the efficiency and quality of task allocation are improved, ultimately an optimal task allocation scheme can be generated, ensuring the efficient utilization of system resources, reducing the communication cost, and optimizing the load balance.

[0019] Optionally, based on the latest state snapshot of the resource pool and the scheduling failure case library, the steps of updating the weight parameters of the time series prediction model and the scheduling policy rule library to obtain an updated policy version include: Based on the latest state snapshot of the resource pool, obtain the node load rate, storage distribution status, and task execution progress; Perform dynamic normalization on the node load rate to generate a normalized load rate matrix; Perform topological modeling on the storage distribution status to generate a heat map of storage resource distribution; Perform time series alignment on the task execution progress to generate a task progress time series table; Based on the normalized load rate matrix and the task progress time series table, perform incremental training on the time series prediction model to update the model weight parameters; Dynamically adjust the storage constraint weight in the loss function of the time series prediction model according to the heat map of storage resource distribution; Combining the task progress time series table and the scheduling failure case library, detect the conflict rules in the scheduling policy rule library that conflict with the task execution progress; Perform conditional threshold recalibration and dynamic sorting of execution priorities on the conflict rules to generate an optimized scheduling policy rule library; Based on the heat map of storage resource distribution and the task execution progress, optimize the cold and hot data migration paths of the storage demand stratification policy to obtain an updated storage stratification configuration file; Package the updated time series prediction model, the optimized scheduling policy rule library, and the storage stratification configuration file into an updated policy version.

[0020] By adopting the above technical solution, the continuous optimization of the scheduling policy is ensured. The system's load, storage, and task scheduling policies are adjusted in real time in a data-driven manner, thereby improving the system resource utilization rate, reducing task execution latency, and ensuring the stable operation of the system.

[0021] In a second aspect, the present application provides an intelligent computing power and storage scheduling system for a multi-service system, adopting the following technical solution: An intelligent computing power and storage scheduling system for a multi-service system, the system includes: A data collection module for collecting real-time monitoring data of each node in the multi-service system; A feature extraction module for extracting features from the real-time monitoring data to generate a structured feature vector including service type labels, task priorities, and data access patterns, and obtaining a real-time service feature data set; A prediction module, configured to predict the computing power demand value and the storage requirement layering strategy within a preset future time window through a time series prediction model based on the real-time service feature dataset and the historical load log; A dynamic correction module, configured to dynamically correct the prediction result according to the computing power demand value and the storage requirement layering strategy, in combination with a reinforcement learning algorithm, to generate a dynamic resource demand table and a cross-service priority weight matrix; A rule initialization module, configured to initialize the preset rules in the scheduling policy rule library based on the service type; A task allocation module, configured to generate a task allocation plan through a hybrid scheduling algorithm based on the preset rules and the resource pool topology information, according to the dynamic resource demand table and the cross-service priority weight matrix; A storage data migration module, configured to generate a storage data migration instruction through a storage scheduling engine based on the storage requirement layering strategy; An execution module, configured to send the task allocation plan and the storage data migration instruction to a virtualized resource orchestration module to perform operations such as computing node expansion, task migration, and storage data redistribution; A status data collection module, configured to collect the resource pool status data after execution to generate a latest resource pool status snapshot; the latest resource pool status snapshot includes the node load rate, the storage distribution status, and the task execution progress; A policy version update module, configured to update the weight parameters of the time series prediction model and the scheduling policy rule library based on the latest resource pool status snapshot and the scheduling failure case library, to obtain an updated policy version and apply it to the next scheduling cycle.

[0022] In a third aspect, the present application provides a computer device, adopting the following technical solution: A computer device includes a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method as described in the first aspect.

[0023] In a fourth aspect, the present application provides a computer-readable storage medium, adopting the following technical solution: A computer-readable storage medium stores a computer program that can be loaded and executed by a processor to implement any of the methods in the first aspect.

[0024] In summary, the present application includes at least one of the following beneficial technical effects: By real-time monitoring of data and feature extraction, the system can dynamically predict computing power requirements and storage requirements, thereby generating accurate resource requirement tables and priority matrices. Using reinforcement learning to correct the prediction results and optimize task allocation and storage data migration strategies, through a hybrid scheduling algorithm and a storage scheduling engine, the system ensures that tasks and data can be intelligently allocated and migrated to improve resource utilization. Finally, based on real-time feedback data, the prediction model and scheduling strategy are updated to form a continuously optimized scheduling closed-loop, significantly improving the scheduling efficiency, resource utilization, and stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 FIG. is a first flowchart of an intelligent computing power and storage scheduling method for a multi-service system according to an embodiment of the present application.

[0026] Figure 2 FIG. is a second flowchart of an intelligent computing power and storage scheduling method for a multi-service system according to an embodiment of the present application.

[0027] Figure 3 FIG. is a third flowchart of an intelligent computing power and storage scheduling method for a multi-service system according to an embodiment of the present application.

[0028] Figure 4 FIG. is a fourth flowchart of an intelligent computing power and storage scheduling method for a multi-service system according to an embodiment of the present application.

[0029] Figure 5 FIG. is a fifth flowchart of an intelligent computing power and storage scheduling method for a multi-service system according to an embodiment of the present application.

[0030] Figure 6 FIG. is a sixth flowchart of an intelligent computing power and storage scheduling method for a multi-service system according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] In order to make the objectives, technical solutions, and advantages of the present application clearer, the following further describes the present application in detail with reference to the accompanying Figure 1-6 drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0032] An embodiment of the present application discloses an intelligent computing power and storage scheduling method for a multi-service system.

[0033] Referring to Figure 1 , an intelligent computing power and storage scheduling method for a multi-service system, the specific method includes: Step S101, collect the real-time monitoring data of each node in the multi-service system; Among them, in the multi-service system, each node (such as a computing node, a storage node) usually generates a large amount of monitoring data in real time. These data include CPU utilization, GPU utilization, storage I / O throughput, network latency, and task queue status, etc. These real-time monitoring data are an intuitive reflection of the system operation status and can provide key basis for subsequent prediction of computing power demand and storage demand.

[0034] For example, high CPU utilization may mean that the computing power of this node is approaching saturation and measures need to be taken for load balancing; while network latency may affect the efficiency of task execution, thus affecting the performance of the entire system. The collection of real-time monitoring data is usually achieved through monitoring software or hardware tools installed on each node of the system, which can ensure continuous acquisition of the system operation data. In this way, the system can respond to load changes in real time, thus effectively avoiding resource bottlenecks and task execution delays.

[0035] It can be understood that after collecting these data, the system can achieve comprehensive monitoring of the real-time status of each node, ensuring that the subsequent scheduling algorithm can make reasonable decisions based on the current load and demand. Through real-time data collection, the system can flexibly respond to changing business needs and ensure optimal allocation of resources.

[0036] Step S102, perform feature extraction on the real-time monitoring data to generate a structured feature vector including business type labels, task priorities, and data access patterns, and obtain a real-time business feature data set; Among them, the process of feature extraction is to analyze and process the collected real-time monitoring data to construct a data structure that can reflect the system working status and task characteristics. Feature extraction is not limited to extracting common metrics from the original monitoring data, such as CPU or GPU utilization, but also generating specific feature labels for each task according to business requirements. These labels may include task priorities, business types, data access patterns, etc. In this process, task priorities (such as high, medium, low) and data access patterns (such as frequent access, occasional access, cold data, etc.) become the core content of feature extraction. These features can help the system understand the resource requirements and operation status of each task in more detail.

[0037] In the embodiments of the present application, by extracting features from the monitoring data, the generated structured feature vectors will be able to clearly describe the specific requirements of each task or node. The real-time service feature data set composed of these feature vectors will provide accurate input data for the subsequent prediction model. For example, in a system, the feature vector of a certain task may include the following information: CPU utilization rate of 90%, GPU utilization rate of 30%, high task priority, and frequent data access pattern. This feature vector can accurately support subsequent resource scheduling.

[0038] Step S103: Based on the real-time service feature data set and the historical load log, predict the computing power demand value and the storage demand stratification strategy within a preset future time window through a time series prediction model. In one embodiment of the present application, the time series prediction model can adopt a long short-term memory network (LSTM). As a deep learning model specifically used to process time series data, based on the historical load log and the real-time service feature data set, LSTM can predict the computing power demand and storage demand of the system within a preset future time window. The core idea of the time series prediction model is to capture the potential time-dependent relationships in the data by learning from historical data, so as to predict future demands. This process requires using information from multiple dimensions including task type, priority, resource consumption, etc. to help the model make predictions. For example, if the computing requirements of certain tasks have periodic fluctuations in the past few hours, the LSTM model will predict the demands of similar tasks in the future based on this rule.

[0039] Specifically, when predicting the storage demand, the system will analyze the data access frequency and accordingly divide the storage demand into levels, such as hot data, warm data, and cold data. Hot data is accessed frequently and requires high-performance storage; cold data is accessed less frequently and can be stored on cheaper storage media. According to these stratification strategies, the system will allocate resources according to the performance characteristics of the storage media (such as SSD, HDD, object storage). By dividing the data into different levels, the system can optimize the use of storage resources and avoid over-allocation.

[0040] As an implementation manner of the time series prediction model, an LSTM model can be constructed using a deep learning framework. The LSTM model usually consists of an input layer, an LSTM layer, a fully connected layer, and an output layer. In the training stage, the Adam optimizer (learning rate = 0.001) and the MSE loss function are adopted, and an early stopping mechanism (patience = 10 epochs) is introduced to prevent over-training. During prediction, multi-step prediction results are recursively generated through a sliding window, and the MinMaxScaler inverse transformation is used to restore the actual dimension. This architecture achieves a prediction accuracy of RMSE ≤ 5% on the test set and supports dynamic loading of new data for incremental training to adapt to the drift of the power consumption pattern.

[0041] Step S104: According to the computing power demand value and the storage demand stratification strategy, combined with the reinforcement learning algorithm, dynamically correct the prediction results to generate a dynamic resource demand table and a cross-service priority weight matrix; Among them, the obtained prediction results are only the estimated values of future computing power demand and storage demand, but these estimated values still need further optimization and correction. Reinforcement learning is an algorithm that optimizes the decision-making process through a reward and punishment mechanism. In this step, the reinforcement learning algorithm will combine the computing power demand and the storage demand stratification strategy to dynamically adjust the resource demand for different tasks and services. For example, some high-priority tasks may require more computing power resources, while other low-priority tasks can be postponed when resources are tight.

[0042] Specifically, the generation of the dynamic resource demand table will be updated in real time according to information such as the priority, resource demand, and historical scheduling results of different tasks, and calculate the cross-service priority weight matrix. This matrix helps the system determine which tasks should be prioritized to obtain resources to ensure that high-priority tasks are not interfered with by low-priority tasks. At the same time, the cross-service priority weight matrix can handle the competition relationship between multiple tasks, and maximize the overall efficiency of the system by reasonably allocating computing resources and storage resources.

[0043] Step S105: Initialize the preset rules in the scheduling policy rule library based on the service type; Among them, initialize the scheduling policy rule library according to the service type, and set corresponding scheduling rules for different types of services. The scheduling policy rule library contains a variety of preset rules, which are set according to historical experience and business requirements.

[0044] For example, for real-time tasks or offline tasks, the system can process them according to different preset rules. Real-time tasks can be guaranteed priority through the configuration of preemptive scheduling rules even when system resources are insufficient, avoiding delays caused by problems such as insufficient bandwidth; while offline tasks can reduce the pressure on storage resources through the migration strategy of cold data storage, improving the overall resource utilization rate of the system. Through these rules, the system can maintain flexibility when processing multiple tasks and ensure that each task is executed on the most suitable resource node, thereby improving the processing ability and resource utilization rate of the system.

[0045] Step S106: Based on the preset rules and the resource pool topology information, according to the dynamic resource demand table and the cross-service priority weight matrix, generate a task allocation plan through a hybrid scheduling algorithm; Among them, the core of the hybrid scheduling algorithm lies in combining the advantages of different scheduling algorithms, comprehensively considering multiple factors such as computing power requirements, storage requirements, and resource pool topology, and formulating an optimal task allocation plan. The resource pool topology includes the physical location mapping of computing nodes and storage nodes, network bandwidth, the overhead of communication between nodes, etc. The system uses the hybrid scheduling algorithm to reasonably arrange the task allocation to ensure that tasks can be executed on the best nodes while avoiding excessive communication delays and resource contention.

[0046] Specifically, the hybrid scheduling algorithm optimizes task allocation by comprehensively considering multiple factors such as task priority, resource availability, network latency, etc., improves resource utilization and reduces system bottlenecks. In addition, the algorithm can be flexibly adjusted according to changes in the resource pool topology to ensure that resources can be most effectively allocated according to actual needs.

[0047] Step S107: Based on the storage requirement stratification strategy, generate storage data migration instructions through the storage scheduling engine; Among them, in terms of storage management, the system generates data migration instructions by using the storage scheduling engine according to the hierarchical strategy of storage requirements. The storage requirement stratification strategy divides the storage levels of data according to the data access frequency of tasks.

[0048] In the embodiments of the present application, hot data with high access frequency can be stored on high-performance SSD storage, while cold data is migrated to lower-cost HDD storage or object storage. In this way, the system can ensure the access speed of hot data and reduce unnecessary overhead of storage media. At the same time, the storage scheduling engine can also merge duplicate data blocks through the cross-business data fingerprint library, thereby reducing redundant storage and releasing storage space.

[0049] Step S108: Send the task allocation plan and storage data migration instructions to the virtualized resource orchestration module to perform operations such as computing node expansion, task migration, and storage data redistribution; Among them, the task allocation plan and storage data migration instructions are sent to the virtualized resource orchestration module to perform operations such as computing node expansion, task migration, and storage data redistribution. This process is that the system automatically expands computing nodes and storage nodes according to the current load situation to ensure that resources can meet task requirements. For example, when the load of a computing node is too high, the system may automatically expand new computing nodes or migrate some tasks from busy computing nodes to nodes with lower load.

[0050] In addition, the stored data migration instructions can ensure that data is migrated to the most suitable storage node, avoiding storage bottlenecks. The resource pool status data after execution will be collected to generate the latest resource pool status snapshot, which includes information such as the node load rate, storage distribution status, and task execution progress.

[0051] It can be understood that by dynamically adjusting computing resources and storage resources according to the changes in task load, the system can be ensured to flexibly respond to different workloads, avoiding over - utilization or shortage of resources. Through automatic expansion and task migration, the system can maintain efficient operation and reduce human intervention.

[0052] Step S109: Collect the resource pool status data after execution and generate the latest resource pool status snapshot; Among them, the latest resource pool status snapshot includes the node load rate, storage distribution status, and task execution progress.

[0053] Specifically, after the task migration and resource expansion operations are completed, the system will collect the resource pool status data after execution and generate the latest resource pool status snapshot including the node load rate, storage distribution status, and task execution progress.

[0054] Furthermore, through the analysis of these data, the system can evaluate the effect of scheduling decisions to ensure that resource allocation and task execution meet expectations. The node load rate can reflect the utilization of resources, the storage distribution status can reveal whether the storage resources are reasonably allocated, and the task execution progress can reflect whether the tasks are executed smoothly as planned.

[0055] Step S110: Based on the latest resource pool status snapshot and the scheduling failure case library, update the weight parameters of the time - series prediction model and the scheduling policy rule library, obtain the updated policy version, and apply it to the next scheduling cycle.

[0056] Among them, the policy version includes the version identifiers of the time - series prediction model weight file, the scheduling rule library entries, and the storage tiering policy configuration file. The policy version is synchronized to all scheduling nodes of the multi - business system through the distributed configuration center.

[0057] Specifically, the system will update the weight parameters of the time - series prediction model and optimize the scheduling policy rule library based on the latest resource pool status snapshot and the scheduling failure case library. This process continuously improves the prediction ability of the system through methods such as incremental training, enabling the system to improve the scheduling policy according to historical experience at the end of each scheduling cycle. For example, when a task migration fails, the system can analyze the reason for the failure and adjust the task migration rules according to the records in the scheduling failure case library to avoid similar situations from occurring again.

[0058] In the above embodiments, the intelligent scheduling of multi-service system resources is realized. By deeply integrating real-time monitoring data and prediction algorithms, the system can dynamically adjust resource allocation, improve resource utilization rate, reduce system bottlenecks, and at the same time ensure the high efficiency of task execution and the stability of the system. In addition, through the adaptive optimization strategy, the system can continuously improve the scheduling decision-making, making the scheduling effect more accurate and efficient after long-term operation, thus further enhancing the performance and reliability of the entire multi-service system.

[0059] Referring to Figure 2 , as an embodiment of step S104, the steps of dynamically correcting the prediction result according to the computing power demand value and the storage demand stratification strategy and generating a dynamic resource demand table and a cross-service priority weight matrix by combining the reinforcement learning algorithm include: Step S201, obtain the real-time resource pool status data, and receive the computing power demand value and the storage demand stratification strategy within a preset future time window output by the time series prediction model; Among them, the real-time resource pool status data includes the node load rate, the storage bandwidth occupancy rate, and the task queue length, which provide the resource utilization of each computing node and storage node in the system currently. The node load rate reflects the current load of the computing node, the storage bandwidth occupancy rate indicates the utilization degree of the storage device, and the task queue length can reveal the queuing situation of the current system tasks, which is of great significance for evaluating whether the resources are sufficient and whether the tasks are delayed.

[0060] Step S202, merge the computing power demand value, the storage demand stratification strategy, and the real-time resource pool status data into a composite state vector, and perform normalization processing to generate a standardized state matrix; Among them, the obtained computing power demand value, the storage demand stratification strategy, and the real-time resource pool status data are integrated to form a composite state vector. This composite state vector integrates information from multiple dimensions, such as future computing power requirements, storage level requirements, and the current load of nodes. Through normalization processing, it can ensure that data with different dimensions (such as load rate and demand value) have the same scale, thus avoiding biases in the learning process due to data scale differences. The standardized state matrix is to represent these composite state vectors in a consistent format for the reinforcement learning agent to process.

[0061] Step S203, define action space parameters according to the service type, and input the action space parameters into a pre-configured reinforcement learning agent; among them, the action space parameters include resource allocation ratio adjustment instructions and storage migration trigger instructions; Among them, the resource allocation ratio adjustment instruction is a continuous action space, which is used to represent the adjustment of the computing resource allocation ratio for each service type. For example, the system may allocate more resources to high-priority services according to the prediction of the current load and demand. The storage migration trigger instruction is a discrete action space, which defines the trigger thresholds for hot data and cold data migration, helping the system determine when to perform data migration and optimize the use of storage resources. These action space parameters will be input into a pre-configured reinforcement learning agent to guide the reinforcement learning agent on how to select the most appropriate actions (i.e., resource allocation adjustment and storage migration decisions).

[0062] Step S204: Based on the standardized state matrix and historical scheduling records, calculate the resource allocation ratio adjustment instruction and the storage migration trigger instruction through the reinforcement learning agent; among them, the reinforcement learning agent evaluates the action value according to a preset reward function and updates the policy network parameters; Among them, the reinforcement learning agent can adopt the Proximal Policy Optimization (PPO) algorithm to train the policy network. The input layer of the policy network fuses the storage tiering policy parameters, and the output layer maps to the action space. This reinforcement learning algorithm can efficiently and stably update the policy network, enabling the agent to make more accurate resource scheduling decisions in actual operations.

[0063] Specifically, the reinforcement learning agent calculates the resource allocation ratio adjustment instruction and the storage migration trigger instruction based on the standardized state matrix and historical scheduling records. The core of reinforcement learning is to guide the agent's decision-making through "rewards". The agent evaluates the value of each action according to the defined reward function and continuously updates its decision-making rules through the policy network.

[0064] In some embodiments, the reward function is evaluated based on the following aspects: resource utilization reward (based on the ratio of actual computing power to predicted computing power demand), SLA (Service Level Agreement) compliance reward (based on the on-time completion of tasks and the number of delays of high-priority tasks), and storage efficiency penalty (based on the amount of cross-rack data migration and migration time).

[0065] Step S205: Generate a dynamic resource demand table according to the resource allocation ratio adjustment instruction in combination with a preset service feature library; Among them, the system combines the resource allocation ratio adjustment instruction and a preset service feature library to generate a dynamic resource demand table. The service feature library contains feature information of various service types, such as task type, priority, resource consumption pattern, etc. The system uses these feature information to guide the dynamic allocation of resources, ensuring that different service types are properly supported during the resource scheduling process.

[0066] For example, compute-intensive tasks may be prioritized to obtain more computing power resources, while storage-intensive tasks will obtain more storage resources. The generated dynamic resource demand table will provide input data for the next cross-business priority decision.

[0067] Step S206: Based on the dynamic resource demand table and the storage migration trigger instruction, construct a cross-business priority weight matrix.

[0068] Among them, the priority relationship between business tasks can be determined by constructing a resource competition graph, and then priorities are assigned to each task. Specifically, tasks are used as the nodes of the graph, and the degree of overlap of resource requirements between tasks is used as the weight of the edges. The relationship between tasks can be modeled by a graph neural network (GNN). The system can extract the association features of tasks and calculate the initial priority scores of tasks based on these features. Then, the system will adjust the instruction and the initial priority score according to the resource allocation ratio, and calculate the normalized priority weight matrix to ensure that high-priority tasks obtain more resources.

[0069] It can be understood that through the generation of the priority matrix, the system can be ensured to allocate resources reasonably and process important tasks first. By modeling the relationship between tasks with a graph neural network, the accuracy and fairness of resource scheduling can be effectively improved, and resource competition and task latency can be reduced.

[0070] In the above implementation, the dynamic resource scheduling and priority decision method based on reinforcement learning can automatically and intelligently handle the resource allocation problem in a multi-business system. By combining the real-time resource pool status, time series prediction data, and reinforcement learning algorithm, the system can flexibly adapt to changes in business load and make efficient and accurate resource scheduling decisions.

[0071] Refer to Figure 3 As an implementation of step S206, after the step of constructing a cross-business priority weight matrix based on the dynamic resource demand table and the storage migration trigger instruction, it further includes: Step S301: Input the dynamic resource demand table and the cross-business priority weight matrix into the digital twin simulation environment to generate a simulated scheduling result; Among them, the system will verify the scheduling strategy through the digital twin simulation environment. The digital twin is a virtual model that can reflect the state and behavior of the system in the physical world. In the embodiments of the present application, the digital twin simulation environment can accurately simulate operations such as resource allocation, task scheduling, and storage management in the actual system. By inputting the dynamic resource demand table and the cross-business priority weight matrix, the system can generate a simulated scheduling result.

[0072] It is understandable that the core advantage of the digital twin simulation environment lies in its efficient simulation ability, which can simulate various complex resource scheduling scenarios in a virtual environment to predict various situations that may be encountered in the actual system. Through simulation scheduling, the system can test the effects of different scheduling decisions in actual operations, discover potential problems in a timely manner, and avoid direct operation tests in the actual production environment. The simulation results usually include key indicators such as resource usage, task completion time, and system load status.

[0073] Exemplarily, assume that the system tests two different resource allocation strategies in the simulation environment. One is to allocate more resources to high-priority services, and the other is to evenly allocate resources to all services. Through digital twin simulation, the system can observe which strategy can improve resource utilization and task completion speed in actual execution.

[0074] Step S302, according to the error rate between the actual scheduling result and the simulated scheduling result, trigger the incremental learning module to update the policy network parameters of the reinforcement learning agent.

[0075] Among them, the system judges whether the current scheduling strategy is optimal by comparing the error rate between the actual scheduling result and the simulated scheduling result. If the error between the actual scheduling result and the simulated scheduling result is large, it may indicate that the current scheduling strategy needs to be adjusted. At this time, the system will trigger the incremental learning module to update the policy network parameters of the reinforcement learning agent.

[0076] Specifically, incremental learning is a process of gradually optimizing the learning model based on existing knowledge without losing the knowledge learned before. Through incremental learning, the reinforcement learning agent can gradually optimize its decision-making strategy according to the newly obtained data during actual operation. In this way, the system can continuously self-optimize and improve the accuracy and effectiveness of decision-making.

[0077] Exemplarily, if the system tests a certain resource allocation strategy in the simulation environment, the simulation results show that high-priority tasks have received the expected resource allocation and are successfully completed. However, in actual scheduling, the system may experience delays in high-priority tasks due to some unforeseen factors. At this time, through the comparison of the error rate, the system will find the gap between the simulation and the actual results and trigger the incremental learning module to adjust the policy network parameters, so that the system can better adapt to the changes in actual operation.

[0078] In the above embodiments, the dynamic resource demand table and the cross-service priority weight matrix are input into the digital twin simulation environment for simulation, and the incremental learning module is triggered according to the error rate between the actual scheduling result and the simulated scheduling result to optimize the policy, further strengthening the intelligent scheduling ability of the system in a multi-service environment. The use of digital twin simulation enables the system to test and verify various scheduling policies in a virtual environment, thus avoiding unforeseen errors in the production environment. Through the incremental learning module, the system can continuously optimize the decision-making policy according to new data during actual operation, ensuring that resource scheduling can always meet the requirements of efficiency and flexibility in various complex business scenarios. The overall technical solution can significantly improve resource utilization, reduce task latency, enhance system stability and overall performance, and adapt to the changing multi-service environment.

[0079] As an implementation of step S105, the preset rules in the scheduling policy rule library initialized based on the service type include: when the service type is a real-time task, configure the preemptive scheduling rule and the storage bandwidth reservation threshold; or, when the service type is an offline task, configure the elastic resource pool allocation rule and the cold data storage migration strategy.

[0080] In the embodiments of the present application, for real-time tasks, due to their high requirements for response time, preemptive scheduling of resources may be involved. For example, when the system is processing multiple tasks, if the computational priority of a real-time task is higher than that of other tasks, the system will preempt and preferentially allocate computational resources. At the same time, the real-time task also has relatively strict requirements for bandwidth. Therefore, when initializing the scheduling policy, a bandwidth reservation threshold will be configured for the real-time task to ensure that its real-time performance will not be affected due to insufficient bandwidth during execution.

[0081] Different from real-time tasks, offline tasks have lower real-time requirements. Therefore, the elastic resource pool allocation rule can be used to more effectively utilize idle resources. Offline tasks can be processed during off-peak hours, and the data after task completion does not require an immediate response. The cold data storage migration strategy is to ensure that the data of offline tasks is reasonably stored and migrated in the system, reducing the occupancy of online storage and improving the resource utilization of the system at the same time.

[0082] In the above embodiments, combined with the different requirements of real-time tasks and offline tasks, customized scheduling rules are formulated for each task type. The real-time task adopts the preemptive scheduling rule and the bandwidth reservation strategy, ensuring its priority execution under high load; the offline task optimizes the resource utilization through the elastic resource pool and the cold data migration strategy. The scheduling policy library of the system is not only initialized according to business requirements at the beginning, but also dynamically adjusted according to feedback information in each subsequent scheduling cycle to ensure that the policy always meets the current resource and business requirements.

[0083] Reference Figure 4 , as an implementation of step S106, based on preset rules and resource pool topology information, according to the dynamic resource demand table and the cross-service priority weight matrix, the steps of generating a task allocation scheme through a hybrid scheduling algorithm include: Step S401, obtain the dynamic resource demand table, the cross-service priority weight matrix, the resource pool topology information, and the preset rules; Among them, the dynamic resource demand table contains the demand information of tasks for resources such as computing, storage, and bandwidth; the cross-service priority weight matrix reflects the priorities of different services and can help the system make reasonable resource allocation decisions among multiple tasks; the resource pool topology information describes the connection relationship between computing and storage resource nodes and can help optimize the data transmission path during task allocation; the preset rules include preemption strategies, storage affinity rules, and SLA constraints, etc. These rules stipulate the priorities and other constraints during task scheduling.

[0084] Step S402, perform time window slicing and normalization processing on the dynamic resource demand table to generate a normalized resource demand time series slice; Among them, by performing time slicing on the resource demand table, the resource demands of tasks are divided into multiple time windows, usually slices of a fixed duration. The demand values in each time window are normalized to ensure that data in different time periods can be uniformly compared and processed.

[0085] Exemplarily, assume that a task requires a large amount of bandwidth during some time periods and less demand during other time periods. Through time slicing, the resource demands of the task in different time periods can be clarified and normalized. For example, the bandwidth demand value is normalized to the interval [0, 1] for subsequent calculations and comparisons.

[0086] Step S403, construct a weighted directed graph based on the resource pool topology information to generate a topology graph structure file; Among them, a weighted directed graph is constructed through the topology information of the resource pool. The nodes of the graph represent different resource nodes, and the weights of the edges represent factors such as communication delay, remaining bandwidth rate, and storage migration time between nodes. This graph structure file will be used for subsequent task scheduling decisions.

[0087] Exemplarily, assume that there are computing nodes A, B and storage node C in the system. The communication delay between node A and node B is low and the bandwidth is large, so the weight of the edge between them is small. While the migration delay between node A and node C is high and the bandwidth is limited, the edge weight is large. Through graph construction, the resource transmission cost between nodes can be reflected.

[0088] It is understandable that the weighted directed graph provides a perspective of the network structure for subsequent scheduling decisions, enabling task scheduling to take into account the communication costs and migration costs between resource nodes, thereby achieving more refined resource allocation.

[0089] Step S404: Classify and weight the standardized resource demand time series slices according to preset rules; Among them, the standardized resource demand time series slices are classified and weighted according to preset rules. Each rule (such as preemption strategy, storage affinity, SLA constraint) will assign different weights according to different task requirements and business priorities, thereby affecting the task priority and resource allocation.

[0090] Exemplarily, for real-time tasks, the preemption strategy may be given a higher weight, while the storage affinity rule may have less impact on certain tasks. For offline tasks, the weight of the storage affinity rule may be higher to optimize the data migration path.

[0091] Step S405: Combine the cross-business priority weight matrix and the topology graph structure file, calculate the matching degree score between tasks and nodes, and generate a candidate node score table and a rule conflict flag list; Among them, by combining the cross-business priority weight matrix and the topology graph structure file, calculate the matching degree score between each task and the resource node. This score reflects the suitability of the task to be executed on a certain node, based on factors such as task priority, resource requirements, available resources of the node, and communication cost. At the same time, the rule conflict situation should also be marked. For example, when a task requests to preempt resources, but there is already a high-priority task on this node, a rule conflict may occur.

[0092] Exemplarily, if a high-priority real-time task needs to be executed on node A, but there is already a low-priority task running on node A, then the calculated matching degree score may be lower. At the same time, the system will mark the rule conflict to remind subsequent scheduling for adjustment.

[0093] Step S406: Based on the candidate node score table and the rule conflict flag list, generate a final task allocation plan through a heuristic algorithm.

[0094] Among them, based on the candidate node score table, use a heuristic algorithm to generate a preliminary task allocation plan. The heuristic algorithm usually selects the best task allocation plan according to factors such as task priority, node fitness, and resource utilization rate. Specifically, the heuristic algorithm can be an improved genetic algorithm, and its chromosome encoding includes task ID, target node ID, and resource quota.

[0095] Exemplarily, assume that the system has multiple nodes A, B, and C, as well as multiple tasks T1, T2, and T3. The heuristic algorithm will, based on the matching degree scores of each task with the nodes, preferentially allocate task T1 to node A, while tasks T2 and T3 may be allocated to nodes B and C.

[0096] In the above embodiments, the system can efficiently allocate tasks and resources under the circumstances of dynamic resource requirements and changing business priorities, optimize the system performance, and ultimately achieve precise task scheduling and resource utilization, ensuring the efficient coordination and dynamic adaptation of multi-dimensional resources.

[0097] Referring to Figure 5 , as an embodiment of step S406, the steps of generating a final task allocation scheme through a heuristic algorithm based on the candidate node score table and the rule conflict flag list include: Step S501, obtain the candidate node score table, the rule conflict flag list, and the resource pool topology information; Among them, the candidate node score table contains the matching degree scores between each task and the resource nodes, which helps to determine whether a task is suitable for execution on a specific node. The rule conflict flag list lists all possible node combinations with rule conflicts. For example, the nodes where conflicts may occur between the preemption strategy and the load balancing rule. The resource pool topology information is used to describe the layout of each node in the resource pool, including information such as communication delay and bandwidth between nodes, which is used for subsequent calculations and decisions.

[0098] Exemplarily, assume there is task T1, and the candidate resource nodes include nodes A, B, and C. The matching degrees of T1 to A, B, and C in the candidate node score table are 0.9, 0.7, and 0.8 respectively. In the rule conflict flag list, the combination of T1 and node A has no conflict, while there is a conflict between T1 and B (because node B has been occupied by other high-priority tasks), and the resource pool topology information provides the communication delay and bandwidth between each node.

[0099] Step S502, screen the white list of assignable nodes for each task according to the candidate node score table, and exclude the illegal node combinations in the rule conflict flag list; Among them, the system screens the white list of assignable nodes for each task based on the candidate node score table, and excludes those node combinations with rule conflicts. The rule conflict flag list will indicate which node combinations are unavailable. Usually, these combinations are due to task conflicts, resource constraints, or priority issues.

[0100] Exemplarily, for task T1, assume its candidate nodes are nodes A, B, and C. If there is a conflict in the combination of T1 and node B, then B will be excluded from the white list of assignable nodes. Finally, the assignable nodes for T1 are A and C.

[0101] Step S503: Construct an initial population set based on the whitelist of assignable nodes; among them, each population individual uses chromosome encoding to represent the task assignment scheme. Specifically, the system constructs an initial population set, and each population individual represents a task assignment scheme. The chromosome encoding of each population individual contains information such as task ID, target node ID, and resource quota, representing the mapping of tasks to nodes and resource allocation.

[0102] Exemplarily, assume that the system has tasks T1 and T2, and nodes A, B, and C. In the initial population, there may be a population individual representing that task T1 is assigned to node A and task T2 is assigned to node B, and another population individual representing that task T1 is assigned to node C and task T2 is assigned to node A, etc.

[0103] Step S504: Construct a multi-objective fitness function based on the communication delay parameters and node load data in the candidate node score table and resource pool topology information. Among them, the fitness function is the key to measuring the quality of the task assignment scheme. In this step, the system constructs a multi-objective fitness function to evaluate the effect of each task assignment scheme by combining multiple factors such as candidate node scores, communication delays, and node loads.

[0104] Specifically, multiple objectives may include: the matching degree between tasks and nodes (calculated through the candidate node score table), communication cost (calculating the data transmission delay between resource nodes), and load balancing (evaluating whether the load is uniform based on the resource occupancy of each node).

[0105] Step S505: Dynamically adjust the weight coefficients of each optimization objective in the multi-objective fitness function according to the real-time resource pool status data. Among them, according to the status of the real-time resource pool (such as node load, bandwidth, communication delay, etc.), the weights of each objective in the fitness function are dynamically adjusted. For example, when the system load is high, load balancing may obtain a higher weight, and when the network delay is low, the weight of communication cost may be low. Exemplarily, if the load of node A is very high, the system may increase the weight of load balancing to encourage more tasks to be assigned to other nodes with lower loads.

[0106] Step S506: Perform a tournament selection operation on the initial population set to generate a parental population. Among them, the tournament selection operation is a common selection method in genetic algorithms. The system randomly selects several individuals from the initial population for competition and selects individuals with higher fitness as the parental population to ensure that excellent individuals can be passed on to the next generation during the genetic process. Exemplarily, assuming there are multiple task assignment schemes in the initial population, the system randomly selects several schemes, calculates their fitness, and then selects the individual with the highest fitness as the parent.

[0107] Step S507: Based on the communication path table in the resource pool topology information, perform a topology-aware crossover operation on the parental chromosomes in the parental population to generate an offspring population. Among them, the topology-aware crossover operation is to perform a crossover operation on the parental chromosomes in genetic algorithms to generate new offspring individuals. The crossover operation is based on the communication paths in the resource pool topology information to ensure that the task assignment scheme after crossover is more optimized on the network, reducing communication latency and resource conflicts.

[0108] Exemplarily, if the parental chromosomes represent that task T1 is assigned to node A and task T2 is assigned to node B respectively, after crossover, it may generate a scheme where task T1 is assigned to node B and task T2 is assigned to node A, optimized based on communication latency.

[0109] Step S508: Dynamically adjust the mutation rate according to the population diversity evaluation result, and perform probability mutation on the offspring chromosomes in the offspring population based on the candidate node score table. Among them, the mutation operation is an important operation in genetic algorithms, which is used to maintain the diversity of the population. According to the population diversity evaluation, the system dynamically adjusts the mutation rate to ensure that new genes are introduced to a certain extent, thereby avoiding the algorithm falling into a local optimal solution. The mutation operation mutates the offspring individuals based on the candidate node score table. For example, if task T1 is currently assigned to node A, the mutation operation may assign task T1 to node B with a higher matching degree.

[0110] Step S509: Perform local search optimization on the chromosomes with fitness scores higher than the preset threshold. Specifically, when the fitness scores of some offspring individuals are relatively high, the system will perform local search optimization on these individuals. The local search optimization further improves the fitness by refining the task assignment scheme. For example, for chromosomes with higher fitness, local search may further reduce communication latency and optimize load balancing by adjusting the assignment nodes of some tasks.

[0111] Step S510: Screen out the individuals with the optimal fitness from the parental population and the offspring population to form an elite population. Among them, the elite strategy retains the individual with the highest fitness and keeps it in the elite population. In this way, the best task allocation scheme will not be lost due to genetic operations, which helps to accelerate the convergence of the algorithm.

[0112] Step S511, when the convergence condition is met, decode the chromosome with the highest fitness in the elite population into the final task allocation scheme.

[0113] Among them, when the genetic algorithm converges (that is, the fitness improvement in the elite population is lower than the threshold in N consecutive iterations, or the maximum number of iterations is reached), the system decodes the individual with the highest fitness in the elite population into the final task allocation scheme. After multiple generations of optimization, the individual with the highest fitness in the elite population represents the optimal task allocation scheme.

[0114] In the above implementation, based on the combination of the heuristic algorithm and the genetic algorithm, the global optimization of the task allocation scheme is realized, the efficiency and quality of task allocation are improved, and finally the optimal task allocation scheme can be generated, ensuring the efficient use of system resources, reducing the communication cost, and achieving load balancing optimization.

[0115] Refer to Figure 6 , as an implementation of step S110, based on the latest state snapshot of the resource pool and the scheduling failure case library, the steps of updating the weight parameters of the time series prediction model and the scheduling policy rule library to obtain the updated policy version include: Step S601, according to the latest state snapshot of the resource pool, obtain the node load rate, storage distribution status, and task execution progress; Exemplarily, assume there are three nodes (A, B, C), where the load rate of node A is 85%, node B is 60%, and node C is 92%. The storage resource distribution status shows that the storage resource utilization rate of node A is 50%, node B is 70%, and node C is almost at full load. The task execution schedule shows that the progress of task T1 on node A is 40% and the progress of task T2 on node B is 60%.

[0116] Step S602, perform dynamic normalization on the node load rate to generate a normalized load rate matrix; Specifically, the dynamic normalization of the node load rate ensures the comparability of different node loads. Through normalization methods (such as using the sliding window method or other standardization techniques), the load rates of each node are converted into standardized values, enabling the unified comparison of loads between different nodes. The normalized load rate matrix provides a unified reference for subsequent scheduling optimization. For example, assume the load rates of nodes A, B, and C are 85%, 60%, and 92% respectively. Through normalization, the system may convert the load rates to 0.85, 0.60, and 0.92, which can be directly used to compare the high and low of node loads.

[0117] Step S603: Perform topological modeling on the storage distribution status to generate a heat map of the storage resource distribution; Among them, the distribution of storage resources has an important impact on task scheduling and data migration. Through topological modeling, the system can visually represent the usage of storage resources in the resource pool and generate a heat map of the storage resource distribution. The heat map indicates the usage intensity of storage resources through the depth of color, facilitating the quick identification of bottlenecks in storage resources.

[0118] Exemplarily, in the heat map of the storage resource distribution, the storage utilization rate of node A is 50% and is displayed in green; the utilization rate of node B is 70% and is displayed in yellow; node C is almost fully loaded (90%) and is displayed in red. Through the heat map, the distribution of storage resources can be intuitively seen, facilitating subsequent optimization decisions.

[0119] Step S604: Align the task execution progress in time series to generate a task progress time series table; Among them, the execution progress of tasks is dynamically changing. The task progress time series table can uniformly represent the execution progress of each task at different time points by aligning the progress data of multiple tasks in time series. The system can then evaluate the status of different tasks on the same time axis and perform cross-task resource scheduling.

[0120] Exemplarily, the execution progress of task T1 is 30%, 50%, and 80% at time points t1, t2, and t3 respectively, and the progress of task T2 at time points t1, t2, and t3 is 20%, 40%, and 60%. Through time series alignment, the generated task progress time series table will contain the execution progress data of these tasks at each time point.

[0121] Step S605: Based on the standardized load rate matrix and the task progress time series table, perform incremental training on the time series prediction model to update the model weight parameters; Among them, through the standardized load rate matrix and the task progress time series table, the system performs incremental training on the existing time series prediction model to update the weight parameters of the model. Incremental training is a method of continuously optimizing the existing model based on new data, which can more accurately predict the future resource requirements of tasks and the system load.

[0122] Exemplarily, assume that at time point t1, the load of node A is 0.85 and the progress of task T1 is 30%. The model will update the weight parameters through incremental training to more accurately predict the execution progress of tasks and the node load at the next time point.

[0123] Step S606: Dynamically adjust the storage constraint weight in the loss function of the time series prediction model according to the heat map of the storage resource distribution; Specifically, the state of storage resources directly affects the storage requirements of tasks and the data migration path. Therefore, during the training process of the time series prediction model, based on the storage resource distribution heat map, the storage constraint weight in the model loss function is dynamically adjusted. For example, when the utilization rate of a certain storage node is too high, the model will increase the corresponding storage constraint weight to reduce the load on that node.

[0124] Step S607: Combine the task progress time series table and the scheduling failure case library to detect the conflict rules in the scheduling policy rule library that conflict with the task execution progress. Among them, through the task progress time series table and the scheduling failure case library, the system can detect the conflict rules in the existing scheduling policy rule library that do not match the task execution progress. These conflict rules may be caused by some tasks not being executed as expected or unreasonable resource allocation.

[0125] Step S608: Perform conditional threshold recalibration and dynamic sorting of execution priorities on the conflict rules to generate an optimized scheduling policy rule library. Specifically, when scheduling rule conflicts are detected, the system will perform conditional threshold recalibration on the relevant rules and dynamically adjust the priorities of the rules. This adjustment process ensures that the scheduling rules can maintain optimal execution under changing system states.

[0126] For example, assume a rule stipulates that task T1 must be executed on a node with a load not exceeding 80%. However, in reality, the load of node A exceeds 80%. The system will adjust the threshold of this rule or change the priority of this rule so that task T1 can be executed on time.

[0127] Step S609: Optimize the cold and hot data migration paths of the storage demand stratification strategy based on the storage resource distribution heat map and the task execution progress to obtain an updated storage stratification configuration file. Among them, according to the storage resource distribution heat map and the task execution progress, the system optimizes the stratification strategy of storage requirements and reasonably adjusts the migration paths of cold and hot data. For example, the system can migrate the hot data related to high-progress tasks to storage nodes with lower loads.

[0128] Step S610: Package the updated time series prediction model, the optimized scheduling policy rule library, and the storage stratification configuration file into an updated policy version.

[0129] Among them, this updated policy version will be synchronized to all scheduling nodes of the system to ensure that the system can uniformly execute the updated policy.

[0130] In the above embodiments, the continuous optimization of the scheduling policy is ensured. The system's load, storage, and task scheduling policies are adjusted in real time in a data-driven manner, thereby improving the system resource utilization rate, reducing task execution latency, and ensuring the stable operation of the system.

[0131] The embodiments of the present application also disclose an intelligent computing power and storage scheduling system for a multi-service system.

[0132] An intelligent computing power and storage scheduling system for a multi-service system specifically includes: A data collection module for collecting real-time monitoring data of each node in the multi-service system; A feature extraction module for extracting features from the real-time monitoring data, generating a structured feature vector including business type tags, task priorities, and data access patterns, and obtaining a real-time service feature data set; A prediction module for predicting the computing power demand value and storage demand stratification strategy within a preset time window in the future based on the real-time service feature data set and historical load logs through a time series prediction model; A dynamic correction module for dynamically correcting the prediction results according to the computing power demand value and storage demand stratification strategy, combining with a reinforcement learning algorithm, and generating a dynamic resource demand table and a cross-service priority weight matrix; A rule initialization module for initializing preset rules in the scheduling policy rule library based on business types; A task allocation module for generating a task allocation plan through a hybrid scheduling algorithm based on the preset rules and resource pool topology information, according to the dynamic resource demand table and the cross-service priority weight matrix; A storage data migration module for generating a storage data migration instruction through a storage scheduling engine based on the storage demand stratification strategy; An execution module for sending the task allocation plan and the storage data migration instruction to the virtualized resource orchestration module to perform operations such as computing node expansion, task migration, and storage data redistribution; A status data collection module for collecting the resource pool status data after execution and generating a latest status snapshot of the resource pool; the latest status snapshot of the resource pool includes node load rates, storage distribution status, and task execution progress; A policy version update module for updating the weight parameters of the time series prediction model and the scheduling policy rule library based on the latest status snapshot of the resource pool and the scheduling failure case library, obtaining an updated policy version and applying it to the next scheduling cycle.

[0133] The intelligent computing power and storage scheduling system for a multi-service system in the embodiments of the present application can implement any of the above scheduling methods, and the specific working processes of each module in the scheduling system can refer to the corresponding processes in the above method embodiments.

[0134] In several embodiments provided by the present application, it should be understood that the provided methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for example, the division of a certain module is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0135] The embodiments of the present application also disclose a computer device.

[0136] The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements an intelligent computing power and storage scheduling method of a multi-service system as described above.

[0137] The embodiments of the present application also disclose a computer-readable storage medium.

[0138] The computer-readable storage medium stores a computer program that can be loaded and executed by a processor to implement any one of the intelligent computing power and storage scheduling methods of a multi-service system as described above.

[0139] Among them, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component; the program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0140] It should be noted that in the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0141] The above are all the preferred embodiments of the present application. The protection scope of the present application is not limited by this. Any feature disclosed in this specification (including the abstract and drawings), unless specifically described, can be replaced by other equivalent or similar-purpose alternative features. That is, unless specifically described, each feature is only an example of a series of equivalent or similar features.

Claims

1. An intelligent computing power and storage scheduling method for a multi-business system, characterized in that: The method comprises: Collect real-time monitoring data of each node in the multi-service system; Extracting features from the real-time monitoring data to generate a structured feature vector including a business type label, a task priority, and a data access mode, thereby obtaining a real-time business feature data set; Based on the real-time business feature data set and historical load logs, the computing power demand value and storage demand tiering strategy within the future preset time window are predicted through the time series prediction model; According to the computing power demand value and storage demand stratification strategy, the prediction results are dynamically corrected in combination with the reinforcement learning algorithm to generate a dynamic resource demand table and a cross-business priority weight matrix; Initialize preset rules in the scheduling strategy rule base based on the business type; Based on the preset rules and resource pool topology information, according to the dynamic resource demand table and the cross-service priority weight matrix, a task allocation scheme is generated through a hybrid scheduling algorithm; Based on the storage demand tiering strategy, a storage data migration instruction is generated by a storage scheduling engine; Send the task allocation plan and storage data migration instructions to the virtualization resource orchestration module to perform computing node expansion, task migration and storage data redistribution operations; Collect resource pool status data after execution and generate the latest status snapshot of the resource pool; Based on the latest state snapshot of the resource pool and the scheduling failure case library, the weight parameters of the time series prediction model and the scheduling strategy rule library are updated to obtain an updated strategy version and apply it to the next scheduling cycle.

2. The intelligent computing power and storage scheduling method for a multi-business system according to claim 1, characterized in that: According to the computing power demand value and storage demand stratification strategy, the prediction results are dynamically corrected in combination with the reinforcement learning algorithm to generate a dynamic resource demand table and a cross-business priority weight matrix, including: Obtain real-time resource pool status data, and receive computing power demand values ​​and storage demand tiering strategies within future preset time windows output by the time series prediction model; The computing power demand value, storage demand tiering strategy and real-time resource pool status data are combined into a composite state vector, and normalized to generate a standardized state matrix; Defining action space parameters according to the business type, and inputting the action space parameters into a pre-configured reinforcement learning agent; wherein the action space parameters include resource allocation ratio adjustment instructions and storage migration trigger instructions; Based on the standardized state matrix and historical scheduling records, the resource allocation ratio adjustment instruction and the storage migration trigger instruction are calculated by the reinforcement learning agent; wherein the reinforcement learning agent evaluates the action value according to the preset reward function and updates the policy network parameters; Generate a dynamic resource demand table according to the resource allocation ratio adjustment instruction and in combination with a preset service feature library; Based on the dynamic resource requirement table and the storage migration trigger instruction, a cross-business priority weight matrix is ​​constructed.

3. The intelligent computing power and storage scheduling method for a multi-business system according to claim 2, characterized in that: After the step of constructing a cross-business priority weight matrix based on the dynamic resource requirement table and the storage migration trigger instruction, the method further includes: Input the dynamic resource demand table and the cross-business priority weight matrix into the digital twin simulation environment to generate a simulation scheduling result; According to the error rate between the actual scheduling result and the simulated scheduling result, the incremental learning module is triggered to update the policy network parameters of the reinforcement learning agent.

4. The intelligent computing power and storage scheduling method for a multi-business system according to claim 1, characterized in that: The preset rules in the service type-based initialization scheduling policy rule base include: When the service type is a real-time task, a preemptive scheduling rule and a storage bandwidth reservation threshold are configured; or, when the service type is an offline task, an elastic resource pool allocation rule and a cold data storage migration strategy are configured.

5. The intelligent computing power and storage scheduling method for a multi-business system according to claim 1, characterized in that: Based on the preset rules and resource pool topology information, according to the dynamic resource demand table and the cross-service priority weight matrix, the step of generating a task allocation solution through a hybrid scheduling algorithm includes: Obtain dynamic resource demand tables, cross-business priority weight matrices, resource pool topology information, and preset rules; Performing time window segmentation and normalization processing on the dynamic resource demand table to generate standardized resource demand time series slices; Construct a weighted directed graph based on the resource pool topology information and generate a topology structure file; Classify and weight the standardized resource demand time series slices according to preset rules; Combine the cross-business priority weight matrix and the topology structure file to calculate the matching score between the task and the node, and generate a candidate node score table and a rule conflict mark list; Based on the candidate node scoring table and the rule conflict mark list, a final task allocation solution is generated through a heuristic algorithm.

6. The intelligent computing power and storage scheduling method for a multi-business system according to claim 5, characterized in that: Based on the candidate node scoring table and the rule conflict mark list, the steps of generating a final task allocation solution by a heuristic algorithm include: Obtain candidate node scoring table, rule conflict marker list and resource pool topology information; Screening a whitelist of assignable nodes for each task according to the candidate node scoring table, and excluding illegal node combinations in the rule conflict mark list; An initial population set is constructed based on the allocatable node whitelist; wherein each population individual uses chromosome encoding to represent a task allocation scheme; Constructing a multi-objective fitness function based on the candidate node scoring table and the communication delay parameters and node load data in the resource pool topology information; Dynamically adjust the weight coefficient of each optimization objective in the multi-objective fitness function according to the real-time resource pool status data; Performing a tournament selection operation on the initial population set to generate a parent population; Based on the communication path table in the resource pool topology information, a topology-aware crossover operation is performed on the parent chromosomes in the parent population to generate a child population; Dynamically adjusting the mutation rate according to the population diversity evaluation result, and performing a probability mutation based on the candidate node scoring table on the offspring chromosomes in the offspring population; Perform local search optimization on chromosomes with fitness scores above a preset threshold; Select individuals with the best fitness from the parent population and the offspring population to form an elite population; When the convergence condition is met, the chromosome with the highest fitness in the elite population is decoded as the final task allocation solution.

7. The intelligent computing power and storage scheduling method for a multi-business system according to any one of claims 1 to 6, characterized in that: Based on the latest state snapshot of the resource pool and the scheduling failure case library, the weight parameters of the time series prediction model and the scheduling strategy rule library are updated to obtain an updated strategy version, including: According to the latest status snapshot of the resource pool, the node load rate, storage distribution status and task execution progress are obtained; Dynamically normalizing the node load rate to generate a standardized load rate matrix; Performing topological modeling on the storage distribution state to generate a storage resource distribution heat map; Performing time series alignment on the task execution progress to generate a task progress time series table; Based on the standardized load rate matrix and the task progress time series table, incrementally training the time series prediction model and updating the model weight parameters; Dynamically adjusting the storage constraint weight in the loss function of the time series prediction model according to the storage resource distribution heat map; In combination with the task progress time sequence table and the scheduling failure case library, detecting conflicting rules in the scheduling strategy rule library that are inconsistent with the task execution progress; Recalibrate the condition thresholds of the conflicting rules and dynamically sort the execution priorities to generate an optimized scheduling strategy rule base; Based on the storage resource distribution heat map and the task execution progress, the storage demand tiering strategy is optimized for hot and cold data migration paths to obtain an updated storage tiering configuration file; The updated time series prediction model, the optimized scheduling strategy rule base and the storage tiering configuration file are encapsulated into an updated strategy version.

8. An intelligent computing power and storage scheduling system for a multi-business system, characterized in that: The system comprises: Data collection module, used to collect real-time monitoring data of each node in the multi-service system; A feature extraction module is used to extract features from the real-time monitoring data, generate a structured feature vector including a business type label, a task priority, and a data access mode, and obtain a real-time business feature data set; A prediction module is used to predict the computing power demand value and storage demand tiering strategy within a future preset time window through a time series prediction model based on the real-time business feature data set and historical load log; A dynamic correction module is used to dynamically correct the prediction results according to the computing power demand value and storage demand stratification strategy in combination with the reinforcement learning algorithm to generate a dynamic resource demand table and a cross-business priority weight matrix; A rule initialization module is used to initialize preset rules in the scheduling strategy rule base based on the business type; A task allocation module, configured to generate a task allocation scheme through a hybrid scheduling algorithm based on the preset rules and resource pool topology information, according to the dynamic resource demand table and the cross-service priority weight matrix; A storage data migration module, used to generate storage data migration instructions through a storage scheduling engine based on the storage demand tiering strategy; An execution module, used to send the task allocation plan and storage data migration instructions to the virtualization resource orchestration module to perform computing node expansion, task migration and storage data redistribution operations; A status data collection module is used to collect the resource pool status data after execution and generate a latest status snapshot of the resource pool; the latest status snapshot of the resource pool includes the node load rate, storage distribution status and task execution progress; The strategy version update module is used to update the weight parameters of the time series prediction model and the scheduling strategy rule library based on the latest status snapshot of the resource pool and the scheduling failure case library, obtain the updated strategy version and apply it to the next scheduling cycle.

9. A computer device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program.

10. A computer-readable storage medium, characterized in that: A computer program is stored which can be loaded by a processor and execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Cloud game resource dynamic allocation and management system

    CN119105864A

  • Dynamic scheduling method for cloud computing resource pool

    CN119537025A

  • An efficient method for dynamic allocation of cloud computing resources

    CN119761745A

  • Computing resource configuration methods and apparatuses

    US20240054020A1

Cited By

  • Method and device for manual capacity expansion and auditing of computing power of intelligent computing center cloud platform

    CN120378310A

  • AI server data processing optimization system and method based on distributed heterogeneous computing

    CN120407210A

  • Multi-application resource allocation method and device

    CN120429125A

  • Software development resource scheduling method and system based on big data

    CN120469814A

  • Distributed computing power scheduling method and system based on AIGC

    CN120508404A