PCFarm resource scheduling method and system based on dynamic load prediction

By adopting dynamic load prediction and multi-objective optimization methods in PCFarm resource scheduling, the problem of difficult to accurately obtain the dependency relationship between tasks and load propagation in the existing technology is solved, and more efficient resource scheduling and task migration effects are achieved.

CN120104355AInactive Publication Date: 2025-06-06SHENZHEN ZHIAO TECH CO LTD

Patent Information

Application Number
CN202510592783.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing resource scheduling methods are difficult to accurately obtain dependencies between tasks, resulting in the inability to fully consider the order of tasks and resource competition between tasks during task scheduling, which affects task execution efficiency and overall performance. Traditional load prediction methods ignore the impact of cluster topology on load propagation, resulting in inaccurate prediction results.

Method used

Using the PCFarm resource scheduling method based on dynamic load prediction, a feature extractor is constructed through multi-modal feature fusion to obtain dependencies between tasks; modeling cluster topology, predicting the load propagation effect between nodes, building a spatiotemporal joint prediction framework, integrating time series and spatial topology information; proposing multi-objective optimization functions, maximizing resource utilization, minimizing the average waiting time of the task, and dynamically adjusting node power consumption; designing a prediction-driven elastic scaling and migration cost model, and optimizing resource scheduling scheme.

Benefits of technology

It achieves more accurate acquisition of dependencies and load propagation between tasks, improves the accuracy of load prediction and resource scheduling, reduces the risk of task migration, and improves the success rate of task migration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104355A_ABST
    Figure CN120104355A_ABST
Patent Text Reader

Abstract

The invention discloses a PCFarm resource scheduling method and system based on dynamic load prediction, and relates to the technical field of resource scheduling, and the method comprises the steps: constructing a feature extractor based on multi-modal feature fusion, mapping original data into a high-dimensional feature vector, and obtaining a dependency relationship between tasks; modeling a cluster topology, predicting a load propagation effect between nodes, constructing a spatio-temporal joint prediction framework, and fusing a time sequence and spatial topology information; proposing a multi-objective optimization function; training a scheduling strategy generator, and generating an optimal scheduling scheme based on the current cluster state; designing the elastic expansion and contraction of the prediction drive; constructing a migration cost model, and quantifying the influence of task migration on the performance; the overall load balance degree of the cluster is calculated through the global controller, and a coarse-grained migration instruction is generated. By arranging the load prediction module and the resource scheduling module, resource allocation is dynamically adjusted according to the real-time state and task requirements of the cluster, and a better resource scheduling effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of resource scheduling, and in particular to a PCFarm resource scheduling method and system based on dynamic load prediction. Background Art

[0002] In large-scale computing scenarios such as cloud computing and high-performance computing, PCFarm (personal computer farm or large-scale computing cluster) is widely used to handle various complex computing tasks, such as scientific computing, big data analysis, deep learning model training, etc. As the scale of computing tasks continues to expand and diversify, how to efficiently schedule resources in PCFarm to meet the performance requirements of different tasks while improving resource utilization, reducing energy consumption and task waiting time has become an urgent problem to be solved.

[0003] Existing resource scheduling methods often have difficulty accurately obtaining the dependencies between tasks, resulting in the inability to fully consider the order of tasks and resource competition when scheduling tasks, thus affecting the execution efficiency and overall performance of tasks. Traditional load prediction methods usually only consider time series information, while ignoring the impact of cluster topology on load propagation, resulting in inaccurate prediction results and failure to reflect cluster load changes in a timely manner. Summary of the invention

[0004] In order to solve the above technical problems, a PCFarm resource scheduling method and system based on dynamic load prediction are provided. This technical solution solves the problems raised in the above background technology.

[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is: The PCFarm resource scheduling method based on dynamic load prediction includes: Based on multimodal feature fusion, a feature extractor is constructed to map the original data into a high-dimensional feature vector to obtain the dependency relationship between tasks; Model the cluster topology, predict the load propagation effect between nodes, build a spatiotemporal joint prediction framework, and integrate time series and spatial topology information; A multi-objective optimization function is proposed to maximize resource utilization, minimize the average waiting time of tasks, and dynamically adjust node power consumption; The training scheduling strategy generator generates the optimal scheduling plan based on the current cluster status, introduces a delayed reward mechanism, and rewards long-term tasks in stages; Design prediction-driven elastic scaling to predict load peaks and use transfer learning to transfer historical scaling experience to new task scenarios. Build a migration cost model to quantify the impact of task migration on performance, simulate at least one migration path before migration, and select the solution with the lowest cost; The global controller is used to calculate the overall load balance of the cluster and generate coarse-grained migration instructions. Each node performs fine-grained task adjustments based on local load and neighbor status. Dynamically adjust the connection between nodes based on load distribution to determine whether the load in each area is higher than the preset threshold. If so, temporarily increase the bandwidth between nodes in the area. If not, no output is made.

[0006] Preferably, the cluster topology modeling, predicting the load propagation effect between nodes, building a spatiotemporal joint prediction framework, and integrating time series and spatial topology information specifically include: Obtain information about each node in the cluster, output the network topology as the connection relationship between nodes, output the CPU, memory, and storage as the hardware configuration of the node, and output the load and temperature as the operating status of the node; The topological structure of the cluster is represented as a graph, where nodes represent the nodes in the cluster and edges represent the connections between nodes; Analyze the propagation path of the load between nodes through network traffic to obtain propagation influencing factors, which include the hardware configuration of the node, the network bandwidth and the connection relationship between the nodes; According to the physical mechanism of load propagation, a propagation model based on physical model is constructed, and the fluid dynamics model is used to simulate the propagation process of load in the network; Collect historical load propagation data and divide the data into training and validation sets; Use the training set to train the propagation model, optimize the model parameters, use the validation set to validate the trained model, evaluate the model's prediction performance, and adjust and optimize the model based on the validation results; Extracting time series features from the load data of the node, wherein the time series features include mean, variance and trend; Extracting spatial features from the topological structure of the cluster, wherein the spatial features include degree centrality and betweenness centrality of the nodes; The spatial features are integrated with the time series features to build a spatiotemporal joint model. The spatiotemporal convolutional network is used to take the time series and spatial topology information as the input of the model to jointly predict the node load. A fusion layer is set in the model to fuse the outputs of the time series model and the spatial topology model, and the attention mechanism is used to perform weighted fusion based on the importance of each feature. The mean, variance and trend characteristics of the time series are spliced ​​with the spatial characteristics of the degree centrality and betweenness centrality of the nodes and input into a unified model; Use the time series model to predict the short-term changes in node load, use the spatial topology model to predict the long-term trend of node load, and perform a weighted fusion of the two results at the decision layer; Design the output layer to map the fused features into the prediction results of the node load, and output various forms of prediction results based on task requirements.

[0007] Preferably, the construction of a migration cost model, quantifying the impact of task migration on performance, simulating at least one migration path before migration, and selecting the solution with the lowest cost specifically includes: The migration cost is divided into data transmission cost, computing resource cost, service interruption cost, dependency cost and energy consumption cost; The ratio of the required transmission data volume to the current network bandwidth and the transmission delay are summed to obtain the data transmission cost; The CPU occupancy ratio, memory occupancy ratio and storage occupancy ratio are summed to obtain the computing resource cost; The product of the interruption time and the loss of revenue per unit time is output as the service interruption cost; Add up the delay increase of each dependent task and calculate the dependency cost; The product of migration power consumption and unit electricity price is output as energy consumption cost; Take the weighted sum of the costs of each dimension to get the comprehensive migration cost; Generate at least one migration path based on system topology, task dependencies, and resource status, and execute the migration operation in the simulation environment, recording the cost of each stage; Sort each migration path by comprehensive cost and select the path with the lowest cost.

[0008] Furthermore, a PCFarm resource scheduling system based on dynamic load prediction is proposed, which is used to implement the PCFarm resource scheduling method based on dynamic load prediction as described above, including: Load prediction module, which is used to build a feature extractor based on multimodal feature fusion, map raw data into high-dimensional feature vectors, obtain dependencies between tasks, model cluster topology, predict load propagation effects between nodes, build a spatiotemporal joint prediction framework, integrate time series and spatial topology information, propose a multi-objective optimization function, maximize resource utilization, minimize the average waiting time of tasks, dynamically adjust node power consumption, train a scheduling strategy generator, generate an optimal scheduling plan based on the current cluster status, introduce a delayed reward mechanism, reward long-term tasks in stages, design prediction-driven elastic scaling, predict load peaks, and use transfer learning to migrate historical scaling experience to new task scenarios; The resource scheduling module is used to build a migration cost model, quantify the impact of task migration on performance, simulate at least one migration path before migration, select the plan with the lowest cost, use the global controller to calculate the overall load balance of the cluster, and generate coarse-grained migration instructions. Each node performs fine-grained task adjustments based on local load and neighbor status, dynamically adjusts connections between nodes based on load distribution, and determines whether the load in each area is higher than the preset threshold.

[0009] Compared with the prior art, the present invention has the following beneficial effects: It can more accurately obtain the dependencies between tasks, integrate time series and spatial topology information, more comprehensively consider the factors of load propagation, improve the accuracy of load prediction, propose a multi-objective optimization function, comprehensively consider multiple goals such as resource utilization, average task waiting time and node power consumption, and more comprehensively optimize the resource scheduling plan to achieve better resource scheduling effects. It can also construct a migration cost model, quantify the impact of task migration on performance, reduce the risk of task migration, and improve the success rate of task migration. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 It is a flow chart of the PCFarm resource scheduling method based on dynamic load prediction of the present invention; Figure 2 A flow chart of a method for constructing a feature extractor according to the present invention; Figure 3 It is a flow chart of the method for predicting load propagation effect between nodes of the present invention; Figure 4 A flow chart of a method for proposing a multi-objective optimization function of the present invention; Figure 5 It is a flow chart of the training scheduling strategy generator method of the present invention; Figure 6 A flow chart of the method for constructing a migration cost model of the present invention; Figure 7 This is a flow chart of the method for calculating the overall load balancing degree of a cluster using a global controller according to the present invention. DETAILED DESCRIPTION

[0011] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are only examples, and those skilled in the art may think of other obvious variations.

[0012] Reference Figure 1 As shown, the PCFarm resource scheduling method based on dynamic load prediction includes: Based on multimodal feature fusion, a feature extractor is constructed to map the original data into a high-dimensional feature vector to obtain the dependency relationship between tasks; Model the cluster topology, predict the load propagation effect between nodes, build a spatiotemporal joint prediction framework, and integrate time series and spatial topology information; A multi-objective optimization function is proposed to maximize resource utilization, minimize the average waiting time of tasks, and dynamically adjust node power consumption; The training scheduling strategy generator generates the optimal scheduling plan based on the current cluster status, introduces a delayed reward mechanism, and rewards long-term tasks in stages; Design prediction-driven elastic scaling to predict load peaks and use transfer learning to transfer historical scaling experience to new task scenarios. Build a migration cost model to quantify the impact of task migration on performance, simulate at least one migration path before migration, and select the solution with the lowest cost; The global controller is used to calculate the overall load balance of the cluster and generate coarse-grained migration instructions. Each node performs fine-grained task adjustments based on local load and neighbor status. Dynamically adjust the connection between nodes based on load distribution to determine whether the load in each area is higher than the preset threshold. If so, temporarily increase the bandwidth between nodes in the area. If not, no output is made.

[0013] Reference Figure 2 As shown in the figure, based on multimodal feature fusion, a feature extractor is constructed to map the original data into a high-dimensional feature vector, and the dependencies between tasks are obtained, including: Acquire visual modality, language modality, speech modality, and sensor modality data; Perform word segmentation to divide the text into words or subword units, and use word embedding vectorization method to convert the text into numerical vectors; Normalize the data to scale it to the range of [0,1] and perform data augmentation to increase data diversity and the generalization ability of the model. Integrate the feature extraction method for each modality into a feature extractor, which receives the raw data of each modality as input and outputs the corresponding feature vector; Selecting a multimodal feature fusion method based on task requirements and data characteristics, wherein the multimodal feature fusion method includes feature-level fusion, decision-level fusion, hybrid-level fusion and model-level fusion; Feature-level fusion maintains the independence of each modality data and uses the complementarity between the modalities to connect the features extracted from each modality into a single high-dimensional feature vector; Decision-level fusion processes and makes decisions on each modality separately, and integrates single-modality decisions into the final decision; Hybrid-level fusion improves the limitations of feature-level fusion and decision-level fusion by combining the output of early fusion and single modality prediction; Model-level fusion obtains the joint feature representation of each modality, and achieves fusion by building a specific fusion model, combining at least two layers of networks with a long short-term memory recursive neural network model, and dealing with multimodal fusion problems at the discourse level based on the relationship between discourses; Based on the fused high-dimensional feature vector, a model is constructed to represent the dependencies between tasks. Tasks are regarded as nodes in a graph neural network, and the dependencies between tasks are regarded as edges in the graph. The graph structure is learned and inferred through the graph neural network to obtain the dependencies between tasks.

[0014] The visual modality uses convolutional neural networks to extract spatial features. The language modality uses pre-trained language models to extract semantic features. The speech modality extracts temporal features through time-frequency analysis combined with long short-term memory recursive neural networks. The sensor modality designs a timing network or manually extracts the mean and variance of statistical features.

[0015] Reference Figure 3 As shown in the figure, cluster topology is modeled, load propagation effects between nodes are predicted, a spatiotemporal joint prediction framework is constructed, and time series and spatial topology information are integrated. Specifically, the following are included: Obtain information about each node in the cluster, output the network topology as the connection relationship between nodes, output the CPU, memory, and storage as the hardware configuration of the node, and output the load and temperature as the operating status of the node; The topological structure of the cluster is represented as a graph, where nodes represent the nodes in the cluster and edges represent the connections between nodes; Analyze the propagation path of the load between nodes through network traffic to obtain propagation influencing factors, which include the hardware configuration of the node, the network bandwidth and the connection relationship between the nodes; According to the physical mechanism of load propagation, a propagation model based on physical model is constructed, and the fluid dynamics model is used to simulate the propagation process of load in the network; Collect historical load propagation data and divide the data into training and validation sets; Use the training set to train the propagation model, optimize the model parameters, use the validation set to validate the trained model, evaluate the model's prediction performance, and adjust and optimize the model based on the validation results; Extracting time series features from the load data of the node, wherein the time series features include mean, variance and trend; Extracting spatial features from the topological structure of the cluster, wherein the spatial features include degree centrality and betweenness centrality of the nodes; The spatial features are integrated with the time series features to build a spatiotemporal joint model. The spatiotemporal convolutional network is used to take the time series and spatial topology information as the input of the model to jointly predict the node load. A fusion layer is set in the model to fuse the outputs of the time series model and the spatial topology model, and the attention mechanism is used to perform weighted fusion based on the importance of each feature. The mean, variance and trend characteristics of the time series are spliced ​​with the spatial characteristics of the degree centrality and betweenness centrality of the nodes and input into a unified model; Use the time series model to predict the short-term changes in node load, use the spatial topology model to predict the long-term trend of node load, and perform a weighted fusion of the two results at the decision layer; Design the output layer to map the fused features into the prediction results of the node load, and output various forms of prediction results based on task requirements.

[0016] The prediction result mapping supports multiple output forms: time series: load curve for the next hour; threshold alarm: trigger an alarm when the predicted load exceeds the threshold; resource scheduling suggestion: adjust node resource allocation according to the load prediction results.

[0017] Reference Figure 4 As shown in the figure, a multi-objective optimization function is proposed to maximize resource utilization, minimize the average waiting time of tasks, and dynamically adjust node power consumption. Specifically, it includes: Based on historical load data, preset weighted coefficients for resource utilization, average task waiting time, and node power consumption; Construct a multi-objective optimization function, and use the weighted sum method to weight the resource utilization, average task waiting time and node power consumption to obtain the multi-objective optimization function; Based on real-time monitoring data, the deviation between the current system state and the target state is calculated, and the node power consumption is adjusted based on the deviation to make the system state close to the target state.

[0018] The multi-objective optimization function is: , In the formula, is a multi-objective optimization function, are the weighted coefficients of resource utilization, average task waiting time and node power consumption, and they all sum to 1. is the resource usage of node i, is the total resource amount of node i, is the total number of nodes, is the waiting time of task j, is the total number of tasks, is the power consumption of node i.

[0019] Reference Figure 5 As shown in the figure, the training scheduling strategy generator generates the optimal scheduling plan based on the current cluster status, introduces a delayed reward mechanism, and rewards long-term tasks in stages. Specifically, it includes: Through the monitoring agent deployed on each node of the cluster, the node's resource usage data, network traffic data and task queue information are collected in real time, and the task priority and estimated execution time information are obtained from the task scheduling system; The cluster scheduling problem is modeled as a Markov decision process, where the state is the current state vector of the cluster, the action is the scheduling plan output by the scheduling strategy generator, and the reward is the reward value calculated based on the execution effect of the scheduling plan; Select a deep neural network algorithm to build a scheduling strategy generator and initialize the model parameters; Generate training data by running the cluster in simulation, select a scheduling plan based on the current state at each time step, execute the plan and observe the next state and reward value, and form a training sample with the state, action, reward and next state; Use the generated training data to train the scheduling strategy generator. During the training process, update the model parameters so that the model selects the scheduling solution that obtains the maximum cumulative reward in each state. Design a reward function that takes both short-term and long-term benefits into account. The reward function is recorded as the weighted sum of task completion time, resource utilization, and task failure rate indicators. For long-term tasks, the execution process is divided into at least two stages, and corresponding rewards are set for each stage. At the end of each stage, the stage reward is calculated based on the task execution status of that stage and added to the cumulative reward; The delayed reward is decayed using an exponential decay function so that the rewards farther away from the current time step have less impact on the cumulative reward.

[0020] Aggregate the original data into a state vector, the dimensions of which include: resource utilization of each node, CPU utilization, memory utilization; network traffic characteristics, average bandwidth, maximum delay; task queue characteristics, such as the number of high-priority tasks and the number of low-priority tasks.

[0021] Reference Figure 6 As shown in the figure, a migration cost model is constructed to quantify the impact of task migration on performance. At least one migration path is simulated before migration, and the solution with the lowest cost is selected. Specifically, the following are included: The migration cost is divided into data transmission cost, computing resource cost, service interruption cost, dependency cost and energy consumption cost; The ratio of the required transmission data volume to the current network bandwidth and the transmission delay are summed to obtain the data transmission cost; The CPU occupancy ratio, memory occupancy ratio and storage occupancy ratio are summed to obtain the computing resource cost; The product of the interruption time and the loss of revenue per unit time is output as the service interruption cost; Add up the delay increase of each dependent task and calculate the dependency cost; The product of migration power consumption and unit electricity price is output as energy consumption cost; Take the weighted sum of the costs of each dimension to get the comprehensive migration cost; Generate at least one migration path based on system topology, task dependencies, and resource status, and execute the migration operation in the simulation environment, recording the cost of each stage; Sort each migration path by comprehensive cost and select the path with the lowest cost.

[0022] The migration cost is divided into five core dimensions: data transmission cost reflects the network overhead of data migration, computing resource cost evaluates target resource consumption, service interruption cost measures business loss during migration, dependency cost quantifies the impact of dependent task delays, and energy consumption cost evaluates the energy consumption cost during the migration process.

[0023] Reference Figure 7 As shown in the figure, the global controller is used to calculate the overall load balance of the cluster and generate coarse-grained migration instructions. Each node performs fine-grained task adjustments based on the local load and neighbor status, including: The global controller collects the global status of the cluster, calculates the load balance, generates migration instructions, and the local nodes execute the migration instructions and make fine-grained adjustments based on the local load and neighbor status. The global controller collects resource usage of each node, aggregates node-level data into a cluster-level view, and builds a global state matrix; The load balancing index is defined as the variance of resource utilization of each node; According to the load balancing degree, nodes with low resource utilization are selected as candidate targets, and a coarse-grained migration instruction is generated, wherein the coarse-grained migration instruction includes a source node, a target node, a task ID list, and a migration priority; Monitor its own resource usage in real time, update load status, exchange load information with neighboring nodes, and build a local topology view; Analyze the load after migration to determine whether there is local overload or idle resources. If so, reschedule the migration task. If not, do not output.

[0024] The global controller is responsible for the overall load balancing of the cluster. It adjusts the large-scale task distribution through coarse-grained migration instructions. The local node is responsible for executing the migration instructions and making fine-grained adjustments based on the local load and neighbor status to ensure local load balancing. After executing the migration task, the local node feeds back the migration results to the global controller, including whether the migration is successful, the migration time, resource changes, etc. The global controller updates the global state matrix based on the feedback results, recalculates the load balancing degree, and generates new migration instructions when necessary. Further, based on the same inventive concept as the above PCFarm resource scheduling method based on dynamic load prediction, the present solution also proposes a PCFarm resource scheduling system based on dynamic load prediction, including: Load prediction module, which is used to build a feature extractor based on multimodal feature fusion, map raw data into high-dimensional feature vectors, obtain dependencies between tasks, model cluster topology, predict load propagation effects between nodes, build a spatiotemporal joint prediction framework, integrate time series and spatial topology information, propose a multi-objective optimization function, maximize resource utilization, minimize the average waiting time of tasks, dynamically adjust node power consumption, train a scheduling strategy generator, generate an optimal scheduling plan based on the current cluster status, introduce a delayed reward mechanism, reward long-term tasks in stages, design prediction-driven elastic scaling, predict load peaks, and use transfer learning to migrate historical scaling experience to new task scenarios; The resource scheduling module is used to build a migration cost model, quantify the impact of task migration on performance, simulate at least one migration path before migration, select the plan with the lowest cost, use the global controller to calculate the overall load balance of the cluster, and generate coarse-grained migration instructions. Each node performs fine-grained task adjustments based on local load and neighbor status, dynamically adjusts connections between nodes based on load distribution, and determines whether the load in each area is higher than the preset threshold.

[0025] Furthermore, the present solution also proposes a computer-readable storage medium on which a computer-readable program is stored. When the computer-readable program is called, the above-mentioned PCFarm resource scheduling method based on dynamic load prediction is executed.

[0026] It is understandable that the storage medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; an optical medium, such as a DVD; or a semiconductor medium, such as a solid state drive (SSD).

[0027] To sum up, the advantages of the present invention are: it can more accurately obtain the dependency between tasks, integrate time series and spatial topology information, can more comprehensively consider the factors of load propagation, improve the accuracy of load prediction, propose a multi-objective optimization function, comprehensively consider multiple objectives such as resource utilization, average task waiting time and node power consumption, and can more comprehensively optimize the resource scheduling scheme to achieve better resource scheduling effects. It can construct a migration cost model, quantify the impact of task migration on performance, reduce the risk of task migration, and improve the success rate of task migration.

[0028] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions only describe the principles of the present invention. The present invention may be subject to various changes and improvements without departing from the spirit and scope of the present invention. These changes and improvements fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the attached claims and their equivalents.

Claims

1. PCFarm resource scheduling method based on dynamic load prediction, characterized in that: include: Based on multimodal feature fusion, a feature extractor is constructed to map the original data into a high-dimensional feature vector to obtain the dependency relationship between tasks; Model the cluster topology, predict the load propagation effect between nodes, build a spatiotemporal joint prediction framework, and integrate time series and spatial topology information; A multi-objective optimization function is proposed to maximize resource utilization, minimize the average waiting time of tasks, and dynamically adjust node power consumption; The training scheduling strategy generator generates the optimal scheduling plan based on the current cluster status, introduces a delayed reward mechanism, and rewards long-term tasks in stages; Design prediction-driven elastic scaling to predict load peaks and use transfer learning to transfer historical scaling experience to new task scenarios. Build a migration cost model to quantify the impact of task migration on performance, simulate at least one migration path before migration, and select the solution with the lowest cost; The global controller is used to calculate the overall load balance of the cluster and generate coarse-grained migration instructions. Each node performs fine-grained task adjustments based on local load and neighbor status. Dynamically adjust the connection between nodes based on load distribution to determine whether the load in each area is higher than the preset threshold. If so, temporarily increase the bandwidth between nodes in the area. If not, no output is made.

2. The PCFarm resource scheduling method based on dynamic load prediction according to claim 1 is characterized in that: The method of constructing a feature extractor based on multimodal feature fusion, mapping the original data into a high-dimensional feature vector, and obtaining the dependency relationship between tasks specifically includes: Acquire visual modality, language modality, speech modality, and sensor modality data; Perform word segmentation to divide the text into words or subword units, and use word embedding vectorization method to convert the text into numerical vectors; Normalize the data to scale it to the range of [0,1] and perform data augmentation to increase data diversity and the generalization ability of the model. Integrate the feature extraction method for each modality into a feature extractor, which receives the raw data of each modality as input and outputs the corresponding feature vector; Selecting a multimodal feature fusion method based on task requirements and data characteristics, wherein the multimodal feature fusion method includes feature-level fusion, decision-level fusion, hybrid-level fusion and model-level fusion; Feature-level fusion maintains the independence of each modality data and uses the complementarity between the modalities to connect the features extracted from each modality into a single high-dimensional feature vector; Decision-level fusion processes and makes decisions on each modality separately, and integrates single-modality decisions into the final decision; Hybrid-level fusion improves the limitations of feature-level fusion and decision-level fusion by combining the output of early fusion and single modality prediction; Model-level fusion obtains the joint feature representation of each modality, and achieves fusion by building a specific fusion model, combining at least two layers of networks with a long short-term memory recursive neural network model, and dealing with multimodal fusion problems at the discourse level based on the relationship between discourses; Based on the fused high-dimensional feature vector, a model is constructed to represent the dependencies between tasks. Tasks are regarded as nodes in a graph neural network, and the dependencies between tasks are regarded as edges in the graph. The graph structure is learned and inferred through the graph neural network to obtain the dependencies between tasks.

3. The PCFarm resource scheduling method based on dynamic load prediction according to claim 2 is characterized in that: The cluster topology modeling, prediction of the load propagation effect between nodes, construction of a spatiotemporal joint prediction framework, and integration of time series and spatial topology information specifically include: Obtain information about each node in the cluster, output the network topology as the connection relationship between nodes, output the CPU, memory, and storage as the hardware configuration of the node, and output the load and temperature as the operating status of the node; The topological structure of the cluster is represented as a graph, where nodes represent the nodes in the cluster and edges represent the connections between nodes; Analyze the propagation path of the load between nodes through network traffic to obtain propagation influencing factors, which include the hardware configuration of the node, the network bandwidth and the connection relationship between the nodes; According to the physical mechanism of load propagation, a propagation model based on physical model is constructed, and the fluid dynamics model is used to simulate the propagation process of load in the network; Collect historical load propagation data and divide the data into training and validation sets; Use the training set to train the propagation model, optimize the model parameters, use the validation set to validate the trained model, evaluate the model's prediction performance, and adjust and optimize the model based on the validation results; Extracting time series features from the load data of the node, wherein the time series features include mean, variance and trend; Extracting spatial features from the topological structure of the cluster, wherein the spatial features include degree centrality and betweenness centrality of the nodes; The spatial features are integrated with the time series features to build a spatiotemporal joint model. The spatiotemporal convolutional network is used to take the time series and spatial topology information as the input of the model to jointly predict the node load. A fusion layer is set in the model to fuse the outputs of the time series model and the spatial topology model, and the attention mechanism is used to perform weighted fusion based on the importance of each feature. The mean, variance and trend characteristics of the time series are spliced ​​with the spatial characteristics of the degree centrality and betweenness centrality of the nodes and input into a unified model; Use the time series model to predict the short-term changes in node load, use the spatial topology model to predict the long-term trend of node load, and perform a weighted fusion of the two results at the decision layer; Design the output layer to map the fused features into the prediction results of the node load, and output various forms of prediction results based on task requirements.

4. The PCFarm resource scheduling method based on dynamic load prediction according to claim 3 is characterized in that: The proposed multi-objective optimization function maximizes resource utilization, minimizes the average waiting time of tasks, and dynamically adjusts node power consumption, specifically including: Based on historical load data, preset weighted coefficients for resource utilization, average task waiting time, and node power consumption; Construct a multi-objective optimization function, and use the weighted sum method to weight the resource utilization, average task waiting time and node power consumption to obtain the multi-objective optimization function; Based on real-time monitoring data, the deviation between the current system state and the target state is calculated, and the node power consumption is adjusted based on the deviation to make the system state close to the target state.

5. The PCFarm resource scheduling method based on dynamic load prediction according to claim 4 is characterized in that: The training scheduling strategy generator generates the optimal scheduling plan based on the current cluster status, introduces a delayed reward mechanism, and rewards long-term tasks in stages, including: Through the monitoring agent deployed on each node of the cluster, the node's resource usage data, network traffic data and task queue information are collected in real time, and the task priority and estimated execution time information are obtained from the task scheduling system; The cluster scheduling problem is modeled as a Markov decision process, where the state is the current state vector of the cluster, the action is the scheduling plan output by the scheduling strategy generator, and the reward is the reward value calculated based on the execution effect of the scheduling plan; Select a deep neural network algorithm to build a scheduling strategy generator and initialize the model parameters; Generate training data by running the cluster in simulation, select a scheduling plan based on the current state at each time step, execute the plan and observe the next state and reward value, and form a training sample with the state, action, reward and next state; Use the generated training data to train the scheduling strategy generator. During the training process, update the model parameters so that the model selects the scheduling solution that obtains the maximum cumulative reward in each state. Design a reward function that takes both short-term and long-term benefits into account. The reward function is recorded as the weighted sum of task completion time, resource utilization, and task failure rate indicators. For long-term tasks, the execution process is divided into at least two stages, and corresponding rewards are set for each stage. At the end of each stage, the stage reward is calculated based on the task execution status of that stage and added to the cumulative reward; The delayed reward is decayed using an exponential decay function so that the rewards farther away from the current time step have less impact on the cumulative reward.

6. The PCFarm resource scheduling method based on dynamic load prediction according to claim 5 is characterized in that: The construction of the migration cost model, quantifying the impact of task migration on performance, simulating at least one migration path before migration, and selecting the solution with the minimum cost specifically includes: The migration cost is divided into data transmission cost, computing resource cost, service interruption cost, dependency cost and energy consumption cost; The ratio of the required transmission data volume to the current network bandwidth and the transmission delay are summed to obtain the data transmission cost; The CPU occupancy ratio, memory occupancy ratio and storage occupancy ratio are summed to obtain the computing resource cost; The product of the interruption time and the loss of revenue per unit time is output as the service interruption cost; Add up the delay increase of each dependent task and calculate the dependency cost; The product of migration power consumption and unit electricity price is output as energy consumption cost; Take the weighted sum of the costs of each dimension to get the comprehensive migration cost; Generate at least one migration path based on system topology, task dependencies, and resource status, and execute the migration operation in the simulation environment, recording the cost of each stage; Sort each migration path by comprehensive cost and select the path with the lowest cost.

7. The PCFarm resource scheduling method based on dynamic load prediction according to claim 6 is characterized in that: The global controller is used to calculate the overall load balance of the cluster, generate coarse-grained migration instructions, and each node performs fine-grained task adjustment based on the local load and neighbor status, specifically including: The global controller collects the global status of the cluster, calculates the load balance, generates migration instructions, and the local nodes execute the migration instructions and make fine-grained adjustments based on the local load and neighbor status. The global controller collects resource usage of each node, aggregates node-level data into a cluster-level view, and builds a global state matrix; The load balancing index is defined as the variance of resource utilization of each node; According to the load balancing degree, nodes with low resource utilization are selected as candidate targets, and a coarse-grained migration instruction is generated, wherein the coarse-grained migration instruction includes a source node, a target node, a task ID list, and a migration priority; Monitor its own resource usage in real time, update load status, exchange load information with neighboring nodes, and build a local topology view; Analyze the load after migration to determine whether there is local overload or idle resources. If so, reschedule the migration task. If not, do not output.

8. A PCFarm resource scheduling system based on dynamic load prediction, used to implement the PCFarm resource scheduling method based on dynamic load prediction as described in any one of claims 1 to 7, characterized in that: include: Load prediction module, which is used to build a feature extractor based on multimodal feature fusion, map raw data into high-dimensional feature vectors, obtain dependencies between tasks, model cluster topology, predict load propagation effects between nodes, build a spatiotemporal joint prediction framework, integrate time series and spatial topology information, propose a multi-objective optimization function, maximize resource utilization, minimize the average waiting time of tasks, dynamically adjust node power consumption, train a scheduling strategy generator, generate an optimal scheduling plan based on the current cluster status, introduce a delayed reward mechanism, reward long-term tasks in stages, design prediction-driven elastic scaling, predict load peaks, and use transfer learning to migrate historical scaling experience to new task scenarios; The resource scheduling module is used to build a migration cost model, quantify the impact of task migration on performance, simulate at least one migration path before migration, select the plan with the lowest cost, use the global controller to calculate the overall load balance of the cluster, and generate coarse-grained migration instructions. Each node performs fine-grained task adjustments based on local load and neighbor status, dynamically adjusts connections between nodes based on load distribution, and determines whether the load in each area is higher than the preset threshold.

Citation Information

Patent Citations

  • Large model reasoning scheduling method based on off-network computing power server

    CN119537032A

  • AI-based big data distributed computing task automatic optimization method and system

    CN119576507A

  • Resource based virtual computing instance scheduling

    US20180095776A1

  • Multi-policy intelligent scheduling method and apparatus oriented to heterogeneous computing power

    US20240111586A1

Cited By

  • Intelligent resource scheduling method and system based on dynamic data consanguinity map

    CN120407208A

  • Super-computing resource intelligent allocation method and system based on AI load monitoring

    CN120429121A

  • Elastic resource adjustment method in distributed data parallel training

    CN120448040A

  • Data resource migration risk prediction method and system, terminal equipment and storage medium

    CN120448161A

  • Dynamic load distribution method and device based on vehicle body controller, equipment and medium

    CN120481887A