Intelligent Scheduling Optimization Method and System for Heterogeneous Metadata Based on Reinforcement Learning Algorithm

By applying intelligent scheduling optimization method based on reinforcement learning algorithms in heterogeneous metadata systems, dynamic analysis and multi-dimensional factor considerations of scheduling optimization in the existing technology are solved, and efficient scheduling strategy optimization and system performance improvement are achieved.

CN119719782BActive Publication Date: 2025-05-27北京科杰科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510218161.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-27
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The prior art lacks the ability to analyze real-time metadata in the scheduling optimization of heterogeneous metadata systems, and cannot respond to system state changes in time. In addition, traditional methods have limitations in feature extraction and optimization mechanisms, and it is difficult to comprehensively consider multi-dimensional factors.

Method used

Using the intelligent scheduling optimization method of heterogeneous metadata based on reinforcement learning algorithm, we can obtain and fuse multi-source heterogeneous metadata features by building a multi-level dynamic feature extraction mechanism and an adaptive dimensionality reduction algorithm, train the scheduling strategy model of the dual-channel neural network structure, and design a multi-objective evaluation function and a hierarchical optimization mechanism to achieve continuous optimization of the scheduling strategy.

Benefits of technology

It significantly improves scheduling efficiency and system performance, ensures the rational use of resources, and can continuously optimize scheduling strategies based on real-time monitoring data and scheduling effect data, and realizes adaptive adjustment and continuous improvement of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719782B_ABST
    Figure CN119719782B_ABST
Patent Text Reader

Abstract

The present invention provides a heterogeneous metadata intelligent scheduling optimization method and system based on a reinforcement learning algorithm, relating to the technical field of heterogeneous data. It includes constructing a multi-level dynamic feature extraction mechanism, using a self-attention mechanism to decompose the features of multi-source heterogeneous metadata, generating a set of feature matrices, and performing feature fusion through a transformer encoder. Training a dual-channel neural network scheduling policy model based on the fused feature vectors, and realizing hierarchical optimization by combining deep reinforcement learning to generate a candidate set of scheduling actions and evaluate their performance. This method effectively improves system load balancing, access latency, resource utilization, and data locality, and has significant technical advantages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to heterogeneous data technology, and in particular to an intelligent scheduling optimization method and system for heterogeneous metadata based on a reinforcement learning algorithm. Background Art

[0002] In a modern computing environment, the scheduling optimization of heterogeneous metadata systems has become an important research direction for improving system performance and resource utilization. With the explosion of data volume and the diversification of application scenarios, traditional scheduling methods are facing more and more challenges. Existing technologies usually rely on static rules or simple heuristic algorithms and are difficult to adapt to dynamic load and resource conditions.

[0003] First of all, traditional scheduling methods lack the ability to dynamically analyze real-time metadata and cannot respond to changes in system status in a timely manner, resulting in lagging scheduling decisions. Secondly, existing technologies often rely on a fixed feature set in feature extraction and fail to fully explore the potential information of multi-source heterogeneous metadata, affecting the accuracy of scheduling strategies. Finally, existing optimization mechanisms usually only focus on a single goal and fail to comprehensively consider multi-dimensional factors such as system load balancing, access latency, and resource utilization, resulting in unsatisfactory scheduling effects. Summary of the Invention

[0004] Embodiments of the present invention provide an intelligent scheduling optimization method and system for heterogeneous metadata based on a reinforcement learning algorithm, which can solve the problems in the prior art.

[0005] In the first aspect of the embodiments of the present invention,

[0006] An intelligent scheduling optimization method for heterogeneous metadata based on a reinforcement learning algorithm is provided, including:

[0007] Construct a multi-level dynamic feature extraction mechanism, obtain real-time metadata information in a multi-source heterogeneous metadata system through a distributed data collection engine, perform feature decomposition on the real-time metadata information using a self-attention mechanism to obtain a set of feature matrices of metadata, the set of feature matrices including a type feature matrix, a time series feature matrix, a relationship feature matrix, and a load feature matrix; input the set of feature matrices into a transformer encoder for feature fusion to generate a fused feature vector, and construct a feature vector space through an adaptive dimensionality reduction algorithm; train a scheduling policy model with a two-channel neural network structure based on the feature vector space, where the graph convolutional neural network channel extracts spatial features based on the type feature matrix and the relationship feature matrix, and the gated recurrent unit network channel extracts time series features based on the time series feature matrix and the load feature matrix, and outputs initial scheduling policy parameters;

[0008] Based on the initial scheduling policy parameters, a hierarchical optimization mechanism based on deep reinforcement learning is established, taking the fused feature vector as the state input to generate a candidate set of scheduling actions; a multi-objective evaluation function is designed to evaluate the candidate actions, and the system load balance score, access delay score, resource utilization score, and data locality score are calculated respectively through an adaptive weight system. The system load balance, the access delay score, the resource utilization score, and the data locality score are weighted and combined to obtain a comprehensive multi-dimensional system performance score, and a comprehensive reward function is constructed based on the comprehensive multi-dimensional system performance score; a dual-objective Q-network structure is used for policy optimization, where the online network updates the policy parameters in real time based on the comprehensive reward function, and the target network maintains training stability through a soft update mechanism, and the optimized policy parameters are output as a scheduling decision sequence;

[0009] The scheduling decision sequence is input into the distributed scheduling execution engine to perform metadata scheduling operations through a two-phase commit protocol; an adaptive monitoring threshold is set based on the fused feature vector, and a distributed monitoring network is deployed to collect system operation status data in real time; when the monitoring data deviates from the adaptive monitoring threshold, the monitoring data and the original fused feature vector are subjected to differential analysis, and a hierarchical scheduling optimization strategy is initiated according to the degree of difference. An incremental adjustment method is used for local anomalies, and a strategy re-optimization is triggered for global anomalies; the optimized scheduling effect data is used as a new training sample, and the dual-channel neural network and the dual-objective Q-network are distributed and incrementally trained through a federated learning framework to achieve continuous optimization of the scheduling policy model.

[0010] Based on the initial scheduling policy parameters, a hierarchical optimization mechanism based on deep reinforcement learning is established, taking the fused feature vector as the state input to generate a candidate set of scheduling actions; a multi-objective evaluation function is designed to evaluate the candidate actions, and the system load balance score, access delay score, resource utilization score, and data locality score are calculated respectively through an adaptive weight system, including:

[0011] Based on the initial scheduling policy parameters, a three-level linked optimization architecture is constructed. The three-level linked optimization architecture includes a policy layer and an execution layer; the policy layer incorporates a deep deterministic policy gradient network with a dual exploration mechanism, and the execution layer incorporates a self-correcting four-dimensional performance evaluation model; the deep deterministic policy gradient network adopts a four-layer fully connected structure, where the number of nodes in the input layer is the same as the dimension of the fused feature vector, and the number of nodes in the output layer is the same as the dimension of the action space;

[0012] Input the fused feature vector into the deep deterministic policy gradient network, perform hierarchical Bayesian modeling on the parameter distribution of the deep deterministic policy gradient network to obtain a hierarchical network weight distribution, and use the Thompson sampling method to obtain multiple groups of network parameter samples from each layer of the hierarchical network weight distribution to form a parameter sample set. At the same time, introduce Gaussian noise perturbation into the parameter sample set; load the parameter sample set added with the Gaussian noise perturbation into the deep deterministic policy gradient network respectively to generate multiple policy network instances, and the multiple policy network instances independently generate a scheduling action candidate set based on the initial scheduling policy parameters;

[0013] Input the scheduling action candidate set into the four-dimensional performance evaluation model of the execution layer. The four-dimensional performance evaluation model sets a self-correction threshold for the evaluation result of each dimension, and triggers a re-evaluation mechanism when the evaluation result exceeds the self-correction threshold; calculate the load balance degree score, access delay score, resource utilization score, and data locality score respectively through the four-dimensional performance evaluation model.

[0014] Construct a comprehensive reward function based on the multi-dimensional system performance comprehensive score; adopt a dual-target Q-network structure for policy optimization, where the online network updates the policy parameters in real time based on the comprehensive reward function, and the target network maintains training stability through a soft update mechanism, including:

[0015] Construct a comprehensive reward function based on the multi-dimensional system performance comprehensive score, calculate the difference between the multi-dimensional system performance comprehensive score at the current moment and the previous moment to obtain a performance change value; set the multi-dimensional system performance comprehensive score as the basic reward item, and set the performance change value as the priority reward item; adaptively adjust the combined weights of the basic reward item and the priority reward item based on the cumulative reward and average reward during the training process to obtain the comprehensive reward function;

[0016] Adopt a dual-target Q-network structure for policy optimization, construct an online network and a target network with a shared network structure; input the current system state into the online network to obtain the current action value, and input the next moment system state into the target network to obtain the target state value; substitute the current action value, the target state value, and the comprehensive reward function into the Bellman equation to calculate the temporal difference error;

[0017] Calculate the sample priority based on the absolute value of the temporal difference error; store the historical interaction samples and the corresponding sample priorities in the experience replay pool; perform importance sampling according to the sample priorities to obtain a training sample batch; input the training sample batch into the online network, and update the policy parameters of the online network in real time through the backpropagation algorithm based on the temporal difference error;

[0018] The strategy parameters of the target network are updated using a soft update mechanism, and the updated strategy parameters of the online network are weighted-averaged with the strategy parameters of the target network to obtain the updated parameters of the target network; where the weight of the weighted average is determined by a preset soft update coefficient, and the soft update coefficient is used to maintain the smoothness of the update of the target network parameters.

[0019] Calculating the sample priority based on the absolute value of the temporal difference error; storing the historical interaction samples and the corresponding sample priorities in an experience replay pool; performing importance sampling according to the sample priorities to obtain a training sample batch; inputting the training sample batch into the online network, and updating the strategy parameters of the online network in real time based on the temporal difference error through the backpropagation algorithm includes:

[0020] Performing an exponential operation on the sum of the absolute value of the temporal difference error and a priority reference value to obtain the sample priority, where the priority reference value is used to ensure that all samples have a non-zero sampling probability;

[0021] Combining the historical interaction samples with the sample priorities to form sample items and storing them in an experience replay pool with a priority queue structure, where the priority queue structure includes leaf nodes and non-leaf nodes, the leaf nodes store the sample items, and the non-leaf nodes store the total priorities of the subtrees;

[0022] Calculating the sampling probability based on the sample priorities, where the sampling probability increases monotonically with the sample priorities, and performing stratified sampling according to the sampling probability to obtain a training sample batch;

[0023] Calculating the importance weights of the samples in the training sample batch, where the importance weights are used to compensate for the sample distribution bias caused by priority sampling, and the importance weights are calculated through an adjustable importance sampling exponent;

[0024] Inputting the training sample batch into the online network, calculating a weighted temporal difference loss function based on the importance weights, and calculating the gradient of the loss function with respect to the parameters of the online network through the backpropagation algorithm;

[0025] The gradient calculation formula of the network parameters is as follows:

[0026] ;

[0027] where, ∇ θ L is the gradient of the loss function with respect to the parameter θ, N is the size of the experience replay pool, w i is the importance weight of the i-th sample, y i is the target Q value of the i-th sample, Q(s i ,a i ;θ) is the online network's prediction of the state-action pair (si , a i ) predicted Q value, ∇ θ Q is the gradient of the Q value with respect to the parameter θ.

[0028] Set an adaptive monitoring threshold based on the fused feature vector, deploy a distributed monitoring network to collect system operation status data in real time; when the monitoring data deviates from the adaptive monitoring threshold, perform a difference analysis on the monitoring data and the original fused feature vector, and start a hierarchical scheduling optimization strategy according to the degree of difference. For local anomalies, an incremental adjustment method is adopted, and for global anomalies, triggering strategy re-optimization includes:

[0029] Establish a dynamic baseline based on the fused feature vector, calculate an adaptive adjustment factor in combination with system operation environment parameters, and multiply the dynamic baseline by the adaptive adjustment factor to obtain an adaptive monitoring threshold;

[0030] Deploy a three-layer distributed monitoring network to collect real-time operation status data of the system. The three-layer distributed monitoring network consists of a local monitoring agent, a regional monitoring node, and a global monitoring center. The device-level status data collected by the local monitoring agent is summarized by the regional monitoring node and then transmitted to the global monitoring center;

[0031] The global monitoring center performs feature extraction and fusion processing on the received real-time operation status data of the system, generates a real-time fused feature vector, and compares the real-time fused feature vector with the adaptive monitoring threshold;

[0032] When the real-time fused feature vector deviates from the adaptive monitoring threshold, calculate the normalized difference value of the real-time fused feature vector relative to the initial fused feature vector, and calculate the comprehensive difference degree based on the normalized difference value;

[0033] Classify the comprehensive difference degree according to a preset difference degree threshold. When the comprehensive difference degree is less than the difference degree threshold, it is determined as a local anomaly and an incremental adjustment strategy is triggered. When the comprehensive difference degree is greater than or equal to the difference degree threshold, it is determined as a global anomaly and a re-optimization strategy is triggered.

[0034] Use the optimized scheduling effect data as a new training sample, and perform distributed incremental training on the dual-channel neural network and the dual-objective Q network through a federated learning framework to continuously optimize the scheduling strategy model, including:

[0035] Collect scheduling effect evaluation data as training samples. The scheduling effect evaluation data includes system state vectors before and after scheduling execution, actual executed scheduling action sequences, resource utilization data, and energy consumption data. Perform normalization preprocessing on the scheduling effect evaluation data to obtain standardized training samples;

[0036] Distribute the standardized training samples to local training nodes through the federated learning framework, and the local training nodes construct local model copies including a dual-channel neural network and a dual-objective Q-network;

[0037] At the local training node, input the state feature vector in the standardized training samples into the state representation network of the dual-channel neural network to obtain a state encoding, and generate a new scheduling action from the action generation network of the dual-channel neural network based on the state encoding;

[0038] Compare the new scheduling action with the actual execution action in the standardized training samples, calculate the policy gradient, and update the parameters of the action generation network;

[0039] Input the state encoding and the new scheduling action into the dual-objective Q-network, calculate the resource utilization Q-value and the energy efficiency Q-value, weight the resource utilization Q-value and the energy efficiency Q-value according to the performance metric value in the standardized training samples to obtain a comprehensive Q-value, and update the parameters of the dual-objective Q-network based on the comprehensive Q-value;

[0040] Regularly trigger the federated averaging algorithm, collect the model parameters of the local training nodes for global aggregation to obtain global model parameters, and distribute the global model parameters to the local training nodes to update the local model copies;

[0041] Collect new scheduling effect evaluation data as incremental training samples, and continuously optimize the dual-channel neural network and the dual-objective Q-network through the federated learning framework.

[0042] Input the state encoding and the new scheduling action into the dual-objective Q-network, calculate the resource utilization Q-value and the energy efficiency Q-value, weight the resource utilization Q-value and the energy efficiency Q-value according to the performance metric value in the standardized training samples to obtain a comprehensive Q-value, and updating the parameters of the dual-objective Q-network based on the comprehensive Q-value includes:

[0043] Concatenate the state encoding and the new scheduling action to obtain a state-action vector, and use the state-action vector as the common input of the resource utilization evaluation network and the energy efficiency evaluation network;

[0044] Input the state-action vector into the resource utilization evaluation network, perform non-linear transformation through the first hidden layer and the second hidden layer of the resource utilization evaluation network to obtain a resource utilization feature, and calculate the resource utilization Q-value based on the resource utilization feature; d

[0045] The formula for calculating the resource utilization Q-value is as follows:

[0046] ;

[0047] Among them, Q r (s t , a t ) is the Q value of resource utilization rate, s t is the state encoding at time t, a t is the action at time t, W 3 is the weight matrix of the output layer, h 2 is the resource utilization rate feature, b 3 is the bias scalar of the output layer;

[0048] Input the state-action vector into the energy efficiency evaluation network, perform non-linear transformation through the first hidden layer and the second hidden layer of the energy efficiency evaluation network to obtain the energy efficiency feature, and calculate the energy efficiency Q value based on the energy efficiency feature;

[0049] The calculation formula of the energy efficiency Q value is as follows:

[0050] ;

[0051] Among them, Q e (s t , a t ) is the energy efficiency Q value, W 3 e is the weight matrix of the output layer of the energy efficiency evaluation network, h 2 e is the energy efficiency feature, b 3 e is the bias scalar of the output layer of the energy efficiency evaluation network;

[0052] Calculate the relative change in resource utilization rate based on the performance metrics in the current batch of training samples, input the relative change in resource utilization rate into the adaptive weight calculation unit, and calculate the resource utilization rate weight and the energy efficiency weight by the adaptive weight calculation unit;

[0053] Multiply the resource utilization rate Q value by the resource utilization rate weight to obtain the weighted resource utilization rate Q value, multiply the energy efficiency Q value by the energy efficiency weight to obtain the weighted energy efficiency Q value, and add the weighted resource utilization rate Q value and the weighted energy efficiency Q value to obtain the comprehensive Q value;

[0054] Construct a temporal difference loss function based on the comprehensive Q value, and update the network parameters of the resource utilization rate evaluation network and the energy efficiency evaluation network by minimizing the temporal difference loss function.

[0055] In the second aspect of the embodiments of the present invention,

[0056] Provide a heterogeneous metadata intelligent scheduling optimization system based on a reinforcement learning algorithm, including:

[0057] The first unit is used to construct a multi-level dynamic feature extraction mechanism. It obtains real-time metadata information in a multi-source heterogeneous metadata system through a distributed data acquisition engine, performs feature decomposition on the real-time metadata information using a self-attention mechanism to obtain a set of feature matrices of the metadata, where the set of feature matrices includes a type feature matrix, a time-series feature matrix, a relationship feature matrix, and a load feature matrix; inputs the set of feature matrices into a transformer encoder for feature fusion to generate a fused feature vector, and constructs a feature vector space through an adaptive dimensionality reduction algorithm; trains a scheduling policy model with a dual-channel neural network structure based on the feature vector space, where the graph convolutional neural network channel extracts spatial features based on the type feature matrix and the relationship feature matrix, and the gated recurrent unit network channel extracts time-series features based on the time-series feature matrix and the load feature matrix, and outputs initial scheduling policy parameters;

[0058] The second unit is used to establish a hierarchical optimization mechanism based on deep reinforcement learning according to the initial scheduling policy parameters, use the fused feature vector as the state input to generate a set of candidate scheduling actions; design a multi-objective evaluation function to evaluate the candidate actions, calculate the system load balance score, access delay score, resource utilization score, and data locality score respectively through an adaptive weight system, perform weighted combination on the system load balance, the access delay score, the resource utilization score, and the data locality score to obtain a comprehensive multi-dimensional system performance score, construct a comprehensive reward function based on the comprehensive multi-dimensional system performance score; use a dual-objective Q-network structure for policy optimization, where the online network updates the policy parameters in real time based on the comprehensive reward function, and the target network maintains training stability through a soft update mechanism, and outputs the optimized policy parameters as a scheduling decision sequence;

[0059] The third unit is used to input the scheduling decision sequence into a distributed scheduling execution engine to perform metadata scheduling operations through a two-phase commit protocol; set an adaptive monitoring threshold based on the fused feature vector, deploy a distributed monitoring network to collect system operation status data in real time; when the monitoring data deviates from the adaptive monitoring threshold, perform differential analysis on the monitoring data and the original fused feature vector, and start a hierarchical scheduling optimization strategy according to the degree of difference, adopt an incremental adjustment method for local anomalies, and trigger policy re-optimization for global anomalies; use the optimized scheduling effect data as a new training sample to perform distributed incremental training on the dual-channel neural network and the dual-objective Q network through a federated learning framework to achieve continuous optimization of the scheduling policy model.

[0060] In the third aspect of the embodiments of the present invention,

[0061] A kind of electronic device is provided, including:

[0062] A processor;

[0063] A memory for storing instructions executable by the processor;

[0064] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0065] In the fourth aspect of the embodiments of the present invention,

[0066] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0067] The beneficial effects of this application are as follows:

[0068] 1. Improve scheduling efficiency: By constructing a multi-level dynamic feature extraction mechanism and an adaptive dimensionality reduction algorithm, it is possible to obtain and fuse multi-source heterogeneous metadata features in real time, thereby optimizing the scheduling strategy and significantly improving the scheduling efficiency of the system.

[0069] 2. Enhance system performance: Using a multi-objective evaluation function to evaluate scheduling actions, comprehensively considering multiple dimensions such as system load balancing, access latency, resource utilization, and data locality, can effectively improve the overall performance of the system and ensure the reasonable utilization of resources.

[0070] 3. Achieve continuous optimization: Through the federated learning framework for distributed incremental training of the scheduling strategy model, it is possible to continuously optimize the scheduling strategy according to real-time monitoring data and scheduling effect data, realizing the adaptive adjustment and continuous improvement of the system. Description of the Drawings

[0071] Figure 1 It is a schematic flowchart of the heterogeneous metadata intelligent scheduling optimization method based on the reinforcement learning algorithm in the embodiments of the present invention;

[0072] Figure 2 It is a schematic structural diagram of the heterogeneous metadata intelligent scheduling optimization system based on the reinforcement learning algorithm in the embodiments of the present invention. Detailed Embodiments

[0073] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0074] The technical solution of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0075] Figure 1 It is a schematic flowchart of the heterogeneous metadata intelligent scheduling optimization method based on the reinforcement learning algorithm in the embodiments of the present invention. As Figure 1 shown, the method includes:

[0076] S11. Construct a multi-level dynamic feature extraction mechanism, obtain real-time metadata information in the multi-source heterogeneous metadata system through a distributed data collection engine, perform feature decomposition on the real-time metadata information using a self-attention mechanism to obtain a set of feature matrices of the metadata, and the set of feature matrices includes a type feature matrix, a time series feature matrix, a relationship feature matrix, and a load feature matrix; input the set of feature matrices into a transformer encoder for feature fusion to generate a fused feature vector, and construct a feature vector space through an adaptive dimensionality reduction algorithm; train a scheduling policy model with a two-channel neural network structure based on the feature vector space, where the graph convolutional neural network channel extracts spatial features based on the type feature matrix and the relationship feature matrix, and the gated recurrent unit network channel extracts time series features based on the time series feature matrix and the load feature matrix, and outputs initial scheduling policy parameters;

[0077] S12. According to the initial scheduling policy parameters, establish a hierarchical optimization mechanism based on deep reinforcement learning, use the fused feature vector as the state input to generate a set of candidate scheduling actions; design a multi-objective evaluation function to evaluate the candidate actions, calculate the system load balance score, access delay score, resource utilization score, and data locality score respectively through an adaptive weight system, perform a weighted combination of the system load balance, the access delay score, the resource utilization score, and the data locality score to obtain a multi-dimensional system performance comprehensive score, construct a comprehensive reward function based on the multi-dimensional system performance comprehensive score; use a dual-objective Q-network structure for policy optimization, where the online network updates the policy parameters in real time based on the comprehensive reward function, and the target network maintains training stability through a soft update mechanism, and outputs the optimized policy parameters as a scheduling decision sequence;

[0078] S13. Input the scheduling decision sequence into the distributed scheduling execution engine, and execute the metadata scheduling operation through the two-phase commit protocol; set an adaptive monitoring threshold based on the fusion feature vector, deploy a distributed monitoring network to collect system operation status data in real time; when the monitoring data deviates from the adaptive monitoring threshold, perform a difference analysis on the monitoring data and the original fusion feature vector, and start a hierarchical scheduling optimization strategy according to the degree of difference. For local anomalies, adopt an incremental adjustment method, and trigger a strategy re-optimization for global anomalies; use the optimized scheduling effect data as a new training sample, and perform distributed incremental training on the dual-channel neural network and the dual-objective Q network through the federated learning framework to continuously optimize the scheduling strategy model.

[0079] In an alternative implementation, according to the initial scheduling policy parameters, establish a hierarchical optimization mechanism based on deep reinforcement learning, use the fusion feature vector as the state input, and generate a candidate set of scheduling actions; design a multi-objective evaluation function to evaluate the candidate actions, and calculate the system load balancing score, access delay score, resource utilization score, and data locality score respectively through an adaptive weight system, including:

[0080] Construct a three-level linkage optimization architecture based on the initial scheduling policy parameters. The three-level linkage optimization architecture includes a policy layer and an execution layer; the policy layer incorporates a deep deterministic policy gradient network with a dual exploration mechanism, and the execution layer incorporates a self-correcting four-dimensional performance evaluation model; the deep deterministic policy gradient network adopts a four-layer fully connected structure, the number of input layer nodes is the same as the dimension of the fusion feature vector, and the number of output layer nodes is the same as the dimension of the action space;

[0081] Input the fusion feature vector into the deep deterministic policy gradient network, perform hierarchical Bayesian modeling on the parameter distribution of the deep deterministic policy gradient network to obtain a hierarchical network weight distribution, use the Thompson sampling method to obtain multiple groups of network parameter samples from each layer of the hierarchical network weight distribution to form a parameter sample set, and introduce Gaussian noise perturbation into the parameter sample set at the same time; load the parameter sample set added with the Gaussian noise perturbation into the deep deterministic policy gradient network respectively to generate multiple policy network instances, and the multiple policy network instances independently generate a candidate set of scheduling actions based on the initial scheduling policy parameters;

[0082] Input the candidate set of scheduling actions into the four-dimensional performance evaluation model of the execution layer. The four-dimensional performance evaluation model sets a self-correcting threshold for the evaluation result of each dimension, and triggers a re-evaluation mechanism when the evaluation result exceeds the self-correcting threshold; calculate the load balancing score, access delay score, resource utilization score, and data locality score respectively through the four-dimensional performance evaluation model.

[0083] First, construct a comprehensive reward function based on the comprehensive score of the multi-dimensional system performance. Calculate the performance change value by computing the difference between the comprehensive score of the system performance at the current moment and the previous moment. For example, if the comprehensive score of the system performance at the current moment is 0.85 and that at the previous moment is 0.78, then the performance change value is 0.07. Take the comprehensive score of the system performance as the basic reward item and the performance change value as the priority reward item. In practical applications, set the initial weight of the basic reward item to 0.6 and the weight of the priority reward item to 0.4. During the training process, when the cumulative reward exceeds 1000 and the average reward is greater than 0.8, dynamically adjust the weight ratio, adjust the weight of the basic reward item to 0.5, and the weight of the priority reward item to 0.5, to obtain the comprehensive reward function.

[0084] Next, adopt a dual-objective Q-network structure for policy optimization. Construct an online network and a target network with the same network structure, including an input layer, three hidden layers, and an output layer. The input layer receives the current system state information, including features such as resource utilization rate and load distribution. Input the current state into the online network to obtain the action value evaluation result. For example, the evaluation value of a certain scheduling action is 0.92. At the same time, input the state at the next moment into the target network to obtain the target state value of 0.88. Combine the above action value, state value, and comprehensive reward to calculate the temporal difference error.

[0085] Then, calculate the sample priority based on the absolute value of the temporal difference error. For example, if the absolute value of the temporal difference error is 0.15, the corresponding sample priority is 0.85. Store the historical interaction samples and priority information in the experience replay pool, and set the capacity of the replay pool to 10,000. Perform importance sampling according to the sample priority, and the higher the priority, the greater the probability of being sampled. Sample a training sample batch of size 256 and input it into the online network for training. Update the parameters of the online network through the backpropagation algorithm to gradually approximate the predicted value to the target value.

[0086] Finally, adopt a soft update mechanism to update the parameters of the target network. Set the soft update coefficient to 0.01, and perform weighted averaging on the updated parameters of the online network and the parameters of the target network. In specific implementation, the new value of the target network parameters is equal to 0.99 times the original target network parameters plus 0.01 times the parameters of the online network. Ensure the smooth update of the target network parameters through a small soft update coefficient and improve the training stability.

[0087] The solution of this application can:

[0088] In terms of algorithm performance: By constructing a comprehensive reward function and dynamically adjusting the weights, the adaptability of the model to environmental changes is improved; the use of a dual-target Q-network structure and a soft update mechanism significantly enhances the training stability and avoids policy oscillation. In terms of system efficiency: The priority-based experience replay mechanism improves the sample utilization efficiency; importance sampling ensures that high-value samples can be fully trained; the soft update mechanism reduces the volatility of parameter updates. In terms of practicality: The algorithm design fully considers the actual application requirements and has strong scalability; the parameter configuration is flexible and adjustable, facilitating optimization in different scenarios; the training process is stable and controllable, suitable for engineering practice deployment.

[0089] In an alternative embodiment, a comprehensive reward function is constructed based on the multi-dimensional system performance comprehensive score; a dual-target Q-network structure is used for policy optimization, where the online network updates the policy parameters in real time based on the comprehensive reward function, and the target network maintains the training stability through a soft update mechanism, including:

[0090] Construct a comprehensive reward function based on the multi-dimensional system performance comprehensive score, calculate the difference between the multi-dimensional system performance comprehensive score at the current moment and the previous moment to obtain the performance change value; set the multi-dimensional system performance comprehensive score as the basic reward item, and set the performance change value as the priority reward item; adaptively adjust the combined weights of the basic reward item and the priority reward item based on the cumulative reward and average reward during the training process to obtain the comprehensive reward function;

[0091] Use a dual-target Q-network structure for policy optimization, construct an online network and a target network with a shared network structure; input the current system state into the online network to obtain the current action value, and input the system state at the next moment into the target network to obtain the target state value; substitute the current action value, the target state value, and the comprehensive reward function into the Bellman equation to calculate the temporal difference error;

[0092] Calculate the sample priority based on the absolute value of the temporal difference error; store the historical interaction samples and the corresponding sample priorities in the experience replay pool; perform importance sampling according to the sample priorities to obtain a training sample batch; input the training sample batch into the online network, and update the policy parameters of the online network in real time through the backpropagation algorithm based on the temporal difference error;

[0093] Use a soft update mechanism to update the policy parameters of the target network, and perform weighted averaging on the updated policy parameters of the online network and the policy parameters of the target network to obtain the updated parameters of the target network; where the weight of the weighted average is determined by a preset soft update coefficient, and the soft update coefficient is used to maintain the smoothness of the target network parameter update.

[0094] First, the basis for constructing the comprehensive reward function is the comprehensive evaluation score of the multi-dimensional system performance. This score is obtained by evaluating the system's performance in different dimensions. The score for each dimension can be quantified according to the set criteria, such as response time, resource utilization, and system stability. By comparing the comprehensive evaluation score at the current moment with that at the previous moment, the performance change value is calculated. This change value reflects the improvement or decline of the system performance.

[0095] Next, the comprehensive evaluation score of the multi-dimensional system performance is used as the basic reward item, and the performance change value is used as the priority reward item. The basic reward item provides the basic feedback of the system, while the priority reward item emphasizes the dynamic change of the performance. By analyzing the cumulative reward and the average reward during the training process, the combined weights of these two reward items are dynamically adjusted to ensure that the comprehensive reward function can adapt to different training stages and environmental changes.

[0096] In terms of policy optimization, a dual-objective Q-network structure is adopted. The online network and the target network share the same network structure. The online network is responsible for processing the current system state in real time and outputting the value of the current action, while the target network is used to evaluate the value of the system state at the next moment. By combining the current action value, the target state value, and the comprehensive reward function, the temporal difference error is calculated using the Bellman equation. This error provides an important basis for subsequent policy updates.

[0097] In the calculation of sample priorities, the importance of samples is determined based on the absolute value of the temporal difference error. The historical interaction samples and their corresponding priorities are stored in the experience replay pool for subsequent training use. Importance sampling is performed according to the sample priorities to obtain a batch of training samples. These samples are input into the online network, and the policy parameters are updated in real time through the backpropagation algorithm to improve the learning efficiency of the network.

[0098] Finally, a soft update mechanism is adopted to update the policy parameters of the target network. The updated parameters of the online network and the parameters of the target network are weighted averaged to ensure a smooth parameter update process for the target network. The weight of the weighted average is determined by the preset soft update coefficient. This mechanism effectively avoids drastic fluctuations during the parameter update process, thereby improving the stability of the training.

[0099] The solution of this application can:

[0100] Through the dynamic adjustment of the comprehensive reward function, the system can better adapt to different environmental changes, improving the learning efficiency and the effect of policy optimization. The application of the dual-objective Q-network structure enhances the stability of policy optimization, reduces the fluctuations during the training process, and makes the system perform more reliably in complex environments. The introduction of the soft update mechanism ensures the smooth update of the target network parameters, further improving the overall performance and response ability of the system.

[0101] In an alternative embodiment, sample priorities are calculated based on the absolute value of the temporal difference error; historical interaction samples and their corresponding sample priorities are stored in an experience replay pool; training sample batches are obtained through importance sampling according to the sample priorities; and the training sample batches are input into the online network, and the policy parameters of the online network are updated in real time based on the temporal difference error through backpropagation algorithm, including:

[0102] The sample priority is obtained by performing an exponential operation on the sum of the absolute value of the temporal difference error and a priority reference value, and the priority reference value is used to ensure that all samples have a non-zero sampling probability;

[0103] The historical interaction samples and the sample priorities are combined into sample items and stored in an experience replay pool with a priority queue structure. The priority queue structure includes leaf nodes and non-leaf nodes. The leaf nodes store the sample items, and the non-leaf nodes store the total priorities of the subtrees;

[0104] The sampling probability is calculated based on the sample priority, and the sampling probability increases monotonically with the sample priority. Training sample batches are obtained through stratified sampling according to the sampling probability;

[0105] The importance weights of the samples in the training sample batches are calculated. The importance weights are used to compensate for the sample distribution bias caused by priority sampling, and the importance weights are calculated through an adjustable importance sampling exponent;

[0106] The training sample batches are input into the online network, and a weighted temporal difference loss function is calculated based on the importance weights. The gradient of the loss function with respect to the online network parameters is calculated through backpropagation algorithm;

[0107] The gradient calculation formula of the network parameters is as follows:

[0108] ;

[0109] where, ∇ θ L is the gradient of the loss function with respect to the parameter θ, N is the size of the experience replay pool, w i is the importance weight of the i-th sample, y i is the target Q value of the i-th sample, Q(s i ,a i ;θ) is the predicted Q value of the online network for the state-action pair (s i ,a i ), and ∇ θ Q is the gradient of the Q value with respect to the parameter θ.

[0110] First, collect historical interaction samples and calculate the temporal difference error for each sample. The temporal difference error refers to the difference between the current state and the next state, reflecting the performance of the model under the current decision. By processing the absolute values of these errors, the priority of each sample can be obtained. The priority is calculated by adding the absolute value to a preset priority reference value and then performing an exponential operation to ensure that all samples have a non-zero sampling probability.

[0111] Next, combine the historical interaction samples with the calculated sample priorities into sample items and store them in the experience replay pool. The experience replay pool adopts a priority queue structure, where the leaf nodes are used to store sample items, and the non-leaf nodes store the sum of the priorities of the subtrees. This structure can effectively manage the priorities of samples, enabling the prioritized selection of more important samples during sampling.

[0112] After the samples are stored, calculate the sampling probability based on the sample priorities. The sampling probability refers to the likelihood of selecting a sample from the experience replay pool, and it monotonically increases as the sample priority increases. In this way, hierarchical sampling can be achieved to ensure that important samples are used more frequently during training.

[0113] Then, calculate the importance weight for each sample in the training sample batch. The importance weight is used to compensate for the sample distribution bias caused by priority sampling. Through an adjustable importance sampling exponent, the calculation method of the weight can be flexibly adjusted to adapt to different training requirements.

[0114] After inputting the training sample batch into the online network, calculate the weighted temporal difference loss function based on the importance weights. The loss function reflects the difference between the model prediction and the actual target. Through the backpropagation algorithm, the gradient of the loss function with respect to the online network parameters can be calculated. This process can effectively update the model parameters, thereby improving the learning effect of the model.

[0115] The solution of this application can:

[0116] Improve the sample utilization efficiency: Through priority sampling, important samples are used more frequently, thus accelerating the learning process of the model. Enhance the stability of the model: The introduction of importance weights can effectively compensate for the sample distribution bias and reduce the fluctuations during training. Optimize the training effect: The weighted loss function enables the model to pay more attention to important samples when updating parameters, thereby improving the overall prediction accuracy.

[0117] In an alternative embodiment, an adaptive monitoring threshold is set based on the fused feature vector, and a distributed monitoring network is deployed to collect system operation status data in real time; when the monitoring data deviates from the adaptive monitoring threshold, the monitoring data is subjected to differential analysis with the original fused feature vector, and according to the degree of difference, a hierarchical scheduling optimization strategy is initiated, adopting an incremental adjustment method for local anomalies and triggering a policy re-optimization for global anomalies, including:

[0118] A dynamic baseline is established based on the fused feature vector, an adaptive adjustment factor is calculated in combination with system operation environment parameters, and the dynamic baseline is multiplied by the adaptive adjustment factor to obtain an adaptive monitoring threshold;

[0119] A three-layer distributed monitoring network is deployed to collect real-time operation status data of the system. The three-layer distributed monitoring network consists of local monitoring agents, regional monitoring nodes, and a global monitoring center. The device-level status data collected by the local monitoring agents is aggregated by the regional monitoring nodes and then transmitted to the global monitoring center;

[0120] The global monitoring center performs feature extraction and fusion processing on the received real-time operation status data of the system, generates a real-time fused feature vector, and compares the real-time fused feature vector with the adaptive monitoring threshold;

[0121] When the real-time fused feature vector deviates from the adaptive monitoring threshold, calculate the standardized difference value of the real-time fused feature vector relative to the initial fused feature vector, and calculate the comprehensive difference degree based on the standardized difference value;

[0122] The comprehensive difference degree is classified according to a preset difference degree threshold. When the comprehensive difference degree is less than the difference degree threshold, it is determined as a local anomaly and an incremental adjustment strategy is triggered. When the comprehensive difference degree is greater than or equal to the difference degree threshold, it is determined as a global anomaly and a re-optimization strategy is triggered.

[0123] First, establish the basis of the fused feature vector. By performing multi-dimensional feature extraction on the system operation status data, a comprehensive feature vector is formed. This feature vector contains information such as the operation parameters of the device, environmental factors, and historical data, ensuring its comprehensiveness and accuracy.

[0124] Next, set the adaptive monitoring threshold. According to the fused feature vector and in combination with the operation environment parameters of the system, a dynamic baseline is calculated. The dynamic baseline is set according to the normal operation state of the device and can reflect the performance changes of the device under different environments. The adaptive adjustment factor is adjusted according to the fluctuations of the real-time monitoring data to ensure that the monitoring threshold can reflect the actual operation state of the device in real time.

[0125] Then, deploy a three - layer distributed monitoring network. This network consists of local monitoring agents, regional monitoring nodes, and a global monitoring center. The local monitoring agents are responsible for collecting device - level status data and transmitting the data to the regional monitoring nodes. The regional monitoring nodes aggregate and perform preliminary analysis on the data from multiple local monitoring agents, and finally transmit the processed data to the global monitoring center.

[0126] At the global monitoring center, the received real - time operating status data of the system will undergo feature extraction and fusion processing to generate a real - time fusion feature vector. This vector will be compared with the adaptive monitoring threshold to determine whether the device is operating normally.

[0127] When the real - time fusion feature vector deviates from the adaptive monitoring threshold, it is necessary to calculate the normalized difference value between it and the initial fusion feature vector. This difference value reflects the deviation degree between the current state and the normal state. Based on this difference value, the comprehensive difference degree is further calculated to evaluate the abnormal situation of the device.

[0128] According to the preset difference degree threshold, the comprehensive difference degree is classified. When the comprehensive difference degree is less than the difference degree threshold, it is determined as a local anomaly, and an incremental adjustment strategy is triggered to perform a small - scale parameter adjustment to restore the normal state of the device; when the comprehensive difference degree is greater than or equal to the difference degree threshold, it is determined as a global anomaly, and a re - optimization strategy is triggered to perform a comprehensive system optimization and adjustment.

[0129] The solution of this application can:

[0130] Improve the real - time monitoring ability of the system, be able to detect and handle device anomalies in a timely manner, and reduce the failure rate. By setting the adaptive monitoring threshold, the flexibility and adaptability of the monitoring system are ensured, and it can adapt to the operating states in different environments. The implementation of the hierarchical scheduling optimization strategy makes the handling of local and global anomalies more efficient, and improves the overall operating efficiency of the system.

[0131] In an optional implementation manner, taking the optimized scheduling effect data as new training samples, performing distributed incremental training on the dual - channel neural network and the dual - objective Q - network through the federated learning framework, and realizing the continuous optimization of the scheduling policy model includes:

[0132] Collecting scheduling effect evaluation data as training samples, where the scheduling effect evaluation data includes system state vectors before and after scheduling execution, the actual executed scheduling action sequence, resource utilization data, and energy consumption data, and performing normalization pre - processing on the scheduling effect evaluation data to obtain standardized training samples;

[0133] Distribute the standardized training samples to local training nodes through the federated learning framework, and each local training node constructs a local model copy containing a dual-channel neural network and a dual-objective Q-network;

[0134] At the local training node, input the state feature vector in the standardized training samples into the state representation network of the dual-channel neural network to obtain a state encoding, and generate a new scheduling action by the action generation network of the dual-channel neural network based on the state encoding;

[0135] Compare the new scheduling action with the actual execution action in the standardized training samples, calculate the policy gradient and update the parameters of the action generation network;

[0136] Input the state encoding and the new scheduling action into the dual-objective Q-network, calculate the resource utilization Q-value and the energy efficiency Q-value, weight the resource utilization Q-value and the energy efficiency Q-value according to the performance metric value in the standardized training samples to obtain a comprehensive Q-value, and update the parameters of the dual-objective Q-network based on the comprehensive Q-value;

[0137] Regularly trigger the federated averaging algorithm, collect the model parameters of the local training nodes for global aggregation to obtain global model parameters, and distribute the global model parameters to the local training nodes to update the local model copies;

[0138] Collect new scheduling effect evaluation data as incremental training samples, and continuously optimize the dual-channel neural network and the dual-objective Q-network through the federated learning framework.

[0139] First, collect scheduling effect evaluation data as training samples. This data includes the system state vectors before and after scheduling execution, the actual executed scheduling action sequence, resource utilization data, and energy consumption data. Through normalizing and preprocessing these data, standardized training samples are obtained. The purpose of normalization is to eliminate the influence of different dimensions and ranges on model training, enabling the data to be compared and learned under the same standard.

[0140] Next, distribute the standardized training samples to local training nodes through the federated learning framework. Each local training node constructs a local model copy containing a dual-channel neural network and a dual-objective Q-network. This step ensures that each node can independently perform model training while protecting data privacy.

[0141] At the local training node, input the state feature vector in the standardized training samples into the state representation network of the dual-channel neural network to obtain a state encoding. Based on this state encoding, the action generation network of the dual-channel neural network generates a new scheduling action. This process optimizes the current scheduling decision by learning historical scheduling effects.

[0142] Subsequently, the generated new scheduling actions are compared with the actual execution actions in the standardized training samples, the policy gradient is calculated, and the parameters of the action generation network are updated. In this way, the model can continuously adjust and optimize its decision-making strategy to improve the scheduling effect.

[0143] Next, the state encoding and the new scheduling actions are input into the dual-objective Q-network to calculate the resource utilization Q-value and the energy efficiency Q-value. According to the performance metric values in the standardized training samples, the resource utilization Q-value and the energy efficiency Q-value are weighted to obtain a comprehensive Q-value. The parameters of the dual-objective Q-network are updated based on the comprehensive Q-value to ensure that the model achieves the best balance between resource utilization and energy efficiency.

[0144] The federated averaging algorithm is triggered periodically to collect the model parameters of local training nodes for global aggregation to obtain global model parameters. The global model parameters are distributed to local training nodes to update the local model copies. This process ensures that the models of each node can share the learning results and improve the performance of the overall model.

[0145] Finally, new scheduling effect evaluation data is collected as incremental training samples, and the dual-channel neural network and the dual-objective Q-network are continuously optimized through the federated learning framework. Through continuous iteration and update, the model can adapt to new scheduling environments and requirements and maintain an efficient scheduling strategy.

[0146] The solution of this application can:

[0147] Improve the accuracy and efficiency of the scheduling strategy, optimize resource utilization and energy efficiency, and reduce energy consumption. Protect data privacy through federated learning and avoid security risks brought by centralized data storage. Achieve continuous optimization of the model, adapt to dynamic scheduling environments, and enhance the flexibility and responsiveness of the system.

[0148] In an alternative embodiment, inputting the state encoding and the new scheduling actions into the dual-objective Q-network, calculating the resource utilization Q-value and the energy efficiency Q-value, weighting the resource utilization Q-value and the energy efficiency Q-value according to the performance metric values in the standardized training samples to obtain a comprehensive Q-value, and updating the dual-objective Q-network parameters based on the comprehensive Q-value includes:

[0149] Concatenate the state encoding and the new scheduling actions to obtain a state-action vector, and use the state-action vector as the common input of the resource utilization evaluation network and the energy efficiency evaluation network;

[0150] Input the state-action vector into the resource utilization evaluation network, perform non-linear transformation through the first hidden layer and the second hidden layer of the resource utilization evaluation network to obtain resource utilization features, and calculate the resource utilization Q-value based on the resource utilization features; d

[0151] The calculation formula for the resource utilization Q value is as follows:

[0152] ;

[0153] where Q r (s t , a t ) is the resource utilization Q value, s t is the state encoding at time t, a t is the action at time t, W 3 is the weight matrix of the output layer, h 2 is the resource utilization feature, b 3 is the bias scalar of the output layer;

[0154] Input the state-action vector into the energy efficiency evaluation network, perform non-linear transformation through the first hidden layer and the second hidden layer of the energy efficiency evaluation network to obtain the energy efficiency feature, and calculate the energy efficiency Q value based on the energy efficiency feature;

[0155] The calculation formula for the energy efficiency Q value is as follows:

[0156] ;

[0157] where Q e (s t , a t ) is the energy efficiency Q value, W 3 e is the weight matrix of the output layer of the energy efficiency evaluation network, h 2 e is the energy efficiency feature, b 3 e is the bias scalar of the output layer of the energy efficiency evaluation network;

[0158] Calculate the relative change in resource utilization based on the performance metrics in the current batch of training samples, input the relative change in resource utilization into the adaptive weight calculation unit, and calculate the resource utilization weight and the energy efficiency weight by the adaptive weight calculation unit;

[0159] Multiply the resource utilization Q value by the resource utilization weight to obtain the weighted resource utilization Q value, multiply the energy efficiency Q value by the energy efficiency weight to obtain the weighted energy efficiency Q value, and add the weighted resource utilization Q value and the weighted energy efficiency Q value to obtain the comprehensive Q value;

[0160] Construct a temporal difference loss function based on the comprehensive Q value, and update the network parameters of the resource utilization evaluation network and the energy efficiency evaluation network by minimizing the temporal difference loss function.

[0161] First, process the state encoding and the new scheduling action. Concatenate the state encoding vector and the scheduling action vector to generate a state-action vector. For example, if the state encoding dimension is 128 and the scheduling action dimension is 64, the concatenated state-action vector has a dimension of 192. This vector serves as the common input feature for both the resource utilization evaluation network and the energy efficiency evaluation network.

[0162] For the resource utilization evaluation network, a three-layer neural network structure is adopted. The first hidden layer contains 256 neurons and uses the ReLU activation function for non-linear transformation; the second hidden layer contains 128 neurons and also uses the ReLU activation function to obtain the resource utilization feature vector. Input this feature vector into the output layer with 1 neuron to get the evaluation result of the resource utilization Q value. In practical applications, assume that the resource utilization Q value of a certain state-action combination is 0.85.

[0163] The energy efficiency evaluation network adopts the same network architecture. Through the non-linear transformation of two hidden layers, an energy efficiency feature vector is obtained and input into the output layer to get the energy efficiency Q value. For example, the energy efficiency Q value of the same state-action combination is 0.72. The two evaluation networks share the input layer parameters but have independent hidden layer and output layer parameters.

[0164] Calculate the relative change in resource utilization based on the current training sample batch. Assume that the average resource utilization of the current batch of samples is 75% and that of the previous batch is 70%, then the relative change is 7.14%. Input this change into the adaptive weight calculation unit and calculate the weight coefficients of the two networks according to the preset weight adjustment rule. When the resource utilization improves significantly, increase the weight of the resource utilization evaluation. For example, the calculated resource utilization weight is 0.6 and the energy efficiency weight is 0.4.

[0165] Multiply the resource utilization Q value by the corresponding weight to get a weighted result of 0.51, and multiply the energy efficiency Q value by the weight to get a weighted result of 0.29. Add the two weighted Q values to get a comprehensive Q value of 0.80. This comprehensive Q value reflects both the resource utilization efficiency and the energy consumption level.

[0166] Finally, construct a temporal difference loss function, with the gap between the current comprehensive Q value and the target Q value as the optimization objective. Calculate the gradient through the backpropagation algorithm and update the parameters of the two evaluation networks simultaneously. Use the Adam optimizer for parameter optimization, with the learning rate set to 0.001.

[0167] The solution of this application can:

[0168] In terms of optimization effect: The joint optimization of resource utilization rate and energy efficiency is achieved through the dual-objective Q-network structure; the adaptive weight mechanism ensures the dynamic balance of the two optimization objectives; the shared input layer design reduces the redundant calculation of feature extraction. In terms of algorithm performance: The use of a deep neural network improves the ability to extract state-action features; the multi-level non-linear transformation enhances the expressive ability of the model; the independent evaluation network ensures the professionalism of different optimization objectives. In terms of engineering practicability: The network structure is reasonably designed and convenient for actual deployment; the weight adaptive mechanism can be flexibly adjusted according to actual needs; the selection of evaluation indicators meets the requirements of actual application scenarios.

[0169] Figure 2 FIG. is a schematic structural diagram of a heterogeneous metadata intelligent scheduling optimization system based on a reinforcement learning algorithm according to an embodiment of the present invention, as Figure 2 shown, the system includes:

[0170] The first unit is used to construct a multi-level dynamic feature extraction mechanism, obtain real-time metadata information in a multi-source heterogeneous metadata system through a distributed data collection engine, perform feature decomposition on the real-time metadata information by using a self-attention mechanism to obtain a set of feature matrices of metadata, where the set of feature matrices includes a type feature matrix, a time-series feature matrix, a relationship feature matrix, and a load feature matrix; input the set of feature matrices into a transformer encoder for feature fusion to generate a fused feature vector, and construct a feature vector space through an adaptive dimensionality reduction algorithm; train a scheduling policy model with a two-channel neural network structure based on the feature vector space, where the graph convolutional neural network channel extracts spatial features based on the type feature matrix and the relationship feature matrix, and the gated recurrent unit network channel extracts time-series features based on the time-series feature matrix and the load feature matrix, and outputs initial scheduling policy parameters;

[0171] The second unit is used to establish a hierarchical optimization mechanism based on deep reinforcement learning according to the initial scheduling policy parameters, use the fused feature vector as a state input to generate a set of candidate scheduling actions; design a multi-objective evaluation function to evaluate the candidate actions, calculate the system load balance score, access delay score, resource utilization rate score, and data locality score respectively through an adaptive weight system, perform weighted combination on the system load balance, the access delay score, the resource utilization rate score, and the data locality score to obtain a multi-dimensional comprehensive system performance score, construct a comprehensive reward function based on the multi-dimensional comprehensive system performance score; use a dual-objective Q-network structure for policy optimization, where the online network updates the policy parameters in real time based on the comprehensive reward function, and the target network maintains training stability through a soft update mechanism, and outputs the optimized policy parameters as a scheduling decision sequence;

[0172] A third unit, configured to input the scheduling decision sequence into a distributed scheduling execution engine, and perform metadata scheduling operations through a two-phase commit protocol; set an adaptive monitoring threshold based on the fusion feature vector, deploy a distributed monitoring network to collect system operation status data in real time; when the monitoring data deviates from the adaptive monitoring threshold, perform differential analysis on the monitoring data and the original fusion feature vector, and start a hierarchical scheduling optimization strategy according to the degree of difference, adopt an incremental adjustment method for local anomalies, and trigger a strategy re-optimization for global anomalies; use the optimized scheduling effect data as a new training sample, and perform distributed incremental training on the dual-channel neural network and the dual-objective Q network through a federated learning framework to continuously optimize the scheduling policy model.

[0173] In a third aspect of the embodiments of the present invention,

[0174] a provided electronic device includes:

[0175] a processor;

[0176] a memory for storing instructions executable by the processor;

[0177] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0178] In a fourth aspect of the embodiments of the present invention,

[0179] a provided computer-readable storage medium has computer program instructions stored thereon, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0180] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions for performing various aspects of the present invention loaded thereon.

[0181] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A heterogeneous metadata intelligent scheduling optimization method based on reinforcement learning algorithm, characterized in that: include: Construct a multi-level dynamic feature extraction mechanism, obtain real-time metadata information in a multi-source heterogeneous metadata system through a distributed data acquisition engine, perform feature decomposition on the real-time metadata information using a self-attention mechanism, and obtain a feature matrix set of metadata, wherein the feature matrix set includes a type feature matrix, a time series feature matrix, a relationship feature matrix, and a load feature matrix; input the feature matrix set into a transformer encoder for feature fusion, generate a fused feature vector, and construct a feature vector space through an adaptive dimensionality reduction algorithm; A scheduling strategy model of a dual-channel neural network structure is trained based on the feature vector space, wherein a graph convolutional neural network channel extracts spatial features based on the type feature matrix and the relationship feature matrix, and a gated recurrent unit network channel extracts timing features based on the timing feature matrix and the load feature matrix, and outputs initial scheduling strategy parameters; According to the initial scheduling strategy parameters, a hierarchical optimization mechanism based on deep reinforcement learning is established, and the fused feature vector is used as a state input to generate a candidate set of scheduling actions; a multi-objective evaluation function is designed to evaluate the candidate actions, and the system load balance score, access delay score, resource utilization score and data locality score are calculated respectively through an adaptive weight system, and the system load balance, the access delay score, the resource utilization score and the data locality score are weightedly combined to obtain a multi-dimensional system performance comprehensive score, and a comprehensive reward function is constructed based on the multi-dimensional system performance comprehensive score; A dual-objective Q network structure is used for policy optimization, wherein the online network updates the policy parameters in real time based on the comprehensive reward function, the target network maintains training stability through a soft update mechanism, and outputs the optimized policy parameters as a scheduling decision sequence; Inputting the scheduling decision sequence into a distributed scheduling execution engine, and executing metadata scheduling operations through a two-phase commit protocol; Setting an adaptive monitoring threshold based on the fused feature vector and deploying a distributed monitoring network to collect system operation status data in real time; When the monitoring data deviates from the adaptive monitoring threshold, a difference analysis is performed on the monitoring data and the original fused feature vector, and a hierarchical scheduling optimization strategy is started according to the degree of difference. An incremental adjustment method is used for local anomalies, and the global anomaly triggering strategy is re-optimized. The optimized scheduling effect data is used as a new training sample, and distributed incremental training is performed on the dual-channel neural network and the dual-objective Q network through a federated learning framework to achieve continuous optimization of the scheduling strategy model.

2. The method according to claim 1, characterized in that According to the initial scheduling strategy parameters, a hierarchical optimization mechanism based on deep reinforcement learning is established, and the fused feature vector is used as the state input to generate a candidate set of scheduling actions; a multi-objective evaluation function is designed to evaluate the candidate actions, and the system load balance score, access delay score, resource utilization score and data locality score are calculated respectively through an adaptive weight system, including: A three-level linkage optimization architecture is constructed based on the initial scheduling strategy parameters, and the three-level linkage optimization architecture includes a strategy layer and an execution layer; the strategy layer has a built-in deep deterministic policy gradient network with a dual exploration mechanism, and the execution layer has a built-in self-correcting four-dimensional performance evaluation model; the deep deterministic policy gradient network adopts a four-layer fully connected structure, the number of nodes in the input layer is the same as the dimension of the fusion feature vector, and the number of nodes in the output layer is the same as the dimension of the action space; The fused feature vector is input into the deep deterministic policy gradient network, and the parameter distribution of the deep deterministic policy gradient network is modeled by hierarchical Bayesian modeling to obtain a hierarchical network weight distribution. A Thompson sampling method is used to obtain multiple groups of network parameter samples from each layer of the hierarchical network weight distribution to form a parameter sample set, and Gaussian noise perturbation is introduced into the parameter sample set; the parameter sample sets after adding the Gaussian noise perturbation are respectively loaded into the deep deterministic policy gradient network to generate multiple policy network instances, and the multiple policy network instances independently generate scheduling action candidate sets based on the initial scheduling policy parameters; The scheduling action candidate set is input into the four-dimensional performance evaluation model of the execution layer. The four-dimensional performance evaluation model sets a self-correction threshold for the evaluation result of each dimension, and triggers a re-evaluation mechanism when the evaluation result exceeds the self-correction threshold. The load balancing score, access delay score, resource utilization score and data locality score are calculated respectively through the four-dimensional performance evaluation model.

3. The method according to claim 1, characterized in that A comprehensive reward function is constructed based on the comprehensive score of the multi-dimensional system performance; a dual-objective Q network structure is used for policy optimization, wherein the online network updates the policy parameters in real time based on the comprehensive reward function, and the target network maintains the training stability through a soft update mechanism, including: A comprehensive reward function is constructed based on the multi-dimensional system performance comprehensive score, and the difference between the multi-dimensional system performance comprehensive score at the current moment and the previous moment is calculated to obtain a performance change value; the multi-dimensional system performance comprehensive score is set as a basic reward item, and the performance change value is set as a priority reward item; based on the accumulated reward and average reward during the training process, the combined weights of the basic reward item and the priority reward item are adaptively adjusted to obtain the comprehensive reward function; A dual-objective Q network structure is used for strategy optimization to construct an online network and a target network with a shared network structure; the current system state is input into the online network to obtain the current action value, and the next moment system state is input into the target network to obtain the target state value; the current action value, the target state value and the comprehensive reward function are substituted into the Bellman equation to calculate the temporal difference error; Calculating sample priority based on the absolute value of the temporal difference error; storing historical interaction samples and corresponding sample priorities in an experience replay pool; performing importance sampling according to the sample priorities to obtain training sample batches; inputting the training sample batches into the online network, and updating the strategy parameters of the online network in real time based on the temporal difference error through a back propagation algorithm; The soft update mechanism is adopted to update the policy parameters of the target network, and the updated policy parameters of the online network and the policy parameters of the target network are weighted averaged to obtain the updated parameters of the target network; wherein the weight of the weighted average is determined by a preset soft update coefficient, and the stability of the target network parameter update is maintained by a smaller soft update coefficient.

4. The method according to claim 3, characterized in that Calculating the sample priority based on the absolute value of the temporal difference error; storing the historical interaction samples and the corresponding sample priorities in the experience replay pool; Performing importance sampling according to the sample priorities to obtain a training sample batch; Inputting the training sample batches into the online network, and updating the strategy parameters of the online network in real time based on the temporal difference error by a back propagation algorithm includes: The absolute value of the timing difference error is added to the priority reference value and then an exponential operation is performed to obtain the sample priority, wherein the priority reference value is used to ensure that all samples have a non-zero sampling probability; The historical interaction samples and the sample priorities form a sample item and store it in an experience replay pool of a priority queue structure, wherein the priority queue structure includes leaf nodes and non-leaf nodes, the leaf nodes store the sample items, and the non-leaf nodes store the priority sum of the subtrees; Calculate the sampling probability based on the sample priority, the sampling probability monotonically increases with the sample priority, and perform stratified sampling according to the sampling probability to obtain a training sample batch; Calculating the importance weight of each sample in the training sample batch, wherein the importance weight is used to compensate for the sample distribution deviation caused by priority sampling, and the importance weight is calculated by an adjustable importance sampling index; Inputting the training sample batches into the online network, calculating a weighted temporal difference loss function based on the importance weights, and calculating the gradient of the loss function with respect to the online network parameters by a back propagation algorithm; The gradient calculation formula of the network parameters is as follows: ; Among them, ∇ θ L is the gradient of the loss function with respect to the parameter θ, N is the size of the experience replay pool, and w i is the importance weight of the i-th sample, y i is the target Q value of the i-th sample, Q(s i ,a i ;θ) is the state-action pair of the online network (s i ,a i )’s predicted Q value, ∇ θ Q is the gradient of the Q value with respect to the parameter θ.

5. The method according to claim 1, characterized in that: Based on the fused feature vector, an adaptive monitoring threshold is set, and a distributed monitoring network is deployed to collect system operation status data in real time; when the monitoring data deviates from the adaptive monitoring threshold, the monitoring data and the original fused feature vector are analyzed for differences, and a hierarchical scheduling optimization strategy is started according to the degree of difference, and an incremental adjustment method is adopted for local anomalies. The global anomaly triggering strategy is re-optimized, including: Establishing a dynamic baseline based on the fused feature vector, calculating an adaptive adjustment factor in combination with system operating environment parameters, and multiplying the dynamic baseline by the adaptive adjustment factor to obtain an adaptive monitoring threshold; Deploy a three-layer distributed monitoring network to collect real-time operating status data of the system. The three-layer distributed monitoring network consists of a local monitoring agent, a regional monitoring node, and a global monitoring center. The device-level status data collected by the local monitoring agent is aggregated by the regional monitoring node and transmitted to the global monitoring center; The global monitoring center performs feature extraction and fusion processing on the received real-time operating status data of the system to generate a real-time fusion feature vector, and compares the real-time fusion feature vector with the adaptive monitoring threshold; When the real-time fused feature vector deviates from the adaptive monitoring threshold, calculating a standardized difference value of the real-time fused feature vector relative to the initial fused feature vector, and calculating a comprehensive difference degree based on the standardized difference value; The comprehensive difference degree is graded according to a preset difference degree threshold. When the comprehensive difference degree is less than the difference degree threshold, it is determined as a local abnormality and triggers an incremental adjustment strategy. When the comprehensive difference degree is greater than the difference degree threshold or equal to the difference degree threshold, it is determined as a global abnormality and triggers a re-optimization strategy.

6. The method according to claim 1, characterized in that The optimized scheduling effect data is used as a new training sample, and the dual-channel neural network and the dual-objective Q network are subjected to distributed incremental training through a federated learning framework to achieve continuous optimization of the scheduling strategy model, including: Collecting scheduling effect evaluation data as training samples, the scheduling effect evaluation data includes system state vectors before and after scheduling execution, actually executed scheduling action sequences, resource utilization data and energy consumption data, and performing normalization preprocessing on the scheduling effect evaluation data to obtain standardized training samples; Distributing the standardized training samples to local training nodes through a federated learning framework, and the local training nodes construct local model copies including a dual-channel neural network and a dual-objective Q network; At the local training node, inputting the state feature vector in the standardized training sample into the state representation network of the dual-channel neural network to obtain a state code, and generating a new scheduling action based on the state code by the action generation network of the dual-channel neural network; Comparing the new scheduled action with the actual executed action in the standardized training sample, calculating the policy gradient and updating the parameters of the action generation network; Inputting the state code and the new scheduling action into the dual-objective Q network, calculating a resource utilization Q value and an energy efficiency Q value, weighting the resource utilization Q value and the energy efficiency Q value according to the performance indicator value in the standardized training sample to obtain a comprehensive Q value, and updating the parameters of the dual-objective Q network based on the comprehensive Q value; Periodically triggering a federated averaging algorithm to collect model parameters of the local training nodes for global aggregation to obtain global model parameters, and distributing the global model parameters to the local training nodes to update the local model copies; New scheduling effect evaluation data is collected as incremental training samples, and the dual-channel neural network and the dual-objective Q network are continuously optimized through the federated learning framework.

7. The method according to claim 6, characterized in that Inputting the state code and the new scheduling action into the dual-objective Q network, calculating the resource utilization Q value and the energy efficiency Q value, weighting the resource utilization Q value and the energy efficiency Q value according to the performance indicator value in the standardized training sample to obtain a comprehensive Q value, and updating the dual-objective Q network parameters based on the comprehensive Q value includes: The state code and the new scheduling action are concatenated to obtain a state action vector, and the state action vector is used as a common input of a resource utilization evaluation network and an energy efficiency evaluation network; Inputting the state action vector into the resource utilization evaluation network, performing nonlinear transformation through the first hidden layer and the second hidden layer of the resource utilization evaluation network to obtain resource utilization characteristics, and calculating the resource utilization Q value based on the resource utilization characteristics; The resource utilization Q value calculation formula is as follows: ; Among them, Q r (s t ,a t ) is the resource utilization Q value, s t is the state code at time t, a t is the action at time t, W3 is the weight matrix of the output layer, h2 is the resource utilization feature, and b3 is the bias scalar of the output layer; Inputting the state action vector into the energy efficiency evaluation network, performing nonlinear transformation through the first hidden layer and the second hidden layer of the energy efficiency evaluation network to obtain energy efficiency characteristics, and calculating the energy efficiency Q value based on the energy efficiency characteristics; The energy efficiency Q value calculation formula is as follows: ; Among them, Q e (s t ,a t ) is the energy efficiency Q value, W3 e is the weight matrix of the output layer of the energy efficiency evaluation network, h2 e is the energy efficiency characteristic, b3 e The bias scalar of the output layer of the network for energy efficiency evaluation; Calculate the relative change of resource utilization based on the performance indicators in the current batch of training samples, input the relative change of resource utilization into the adaptive weight calculation unit, and calculate the resource utilization weight and energy efficiency weight by the adaptive weight calculation unit; Multiplying the resource utilization rate Q value by the resource utilization rate weight to obtain a weighted resource utilization rate Q value, multiplying the energy efficiency Q value by the energy efficiency weight to obtain a weighted energy efficiency Q value, and adding the weighted resource utilization rate Q value and the weighted energy efficiency Q value to obtain a comprehensive Q value; A temporal difference loss function is constructed based on the comprehensive Q value, and network parameters of the resource utilization evaluation network and the energy efficiency evaluation network are simultaneously updated by minimizing the temporal difference loss function.

8. A heterogeneous metadata intelligent scheduling optimization system based on a reinforcement learning algorithm, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to construct a multi-level dynamic feature extraction mechanism, obtain real-time metadata information in a multi-source heterogeneous metadata system through a distributed data acquisition engine, perform feature decomposition on the real-time metadata information using a self-attention mechanism, and obtain a feature matrix set of metadata, wherein the feature matrix set includes a type feature matrix, a time series feature matrix, a relationship feature matrix, and a load feature matrix; the feature matrix set is input into a transformer encoder for feature fusion to generate a fused feature vector, and a feature vector space is constructed through an adaptive dimensionality reduction algorithm; A scheduling strategy model of a dual-channel neural network structure is trained based on the feature vector space, wherein a graph convolutional neural network channel extracts spatial features based on the type feature matrix and the relationship feature matrix, and a gated recurrent unit network channel extracts timing features based on the timing feature matrix and the load feature matrix, and outputs initial scheduling strategy parameters; The second unit is used to establish a hierarchical optimization mechanism based on deep reinforcement learning according to the initial scheduling strategy parameters, use the fused feature vector as a state input, and generate a candidate set of scheduling actions; design a multi-objective evaluation function to evaluate the candidate actions, calculate the system load balance score, access delay score, resource utilization score and data locality score respectively through an adaptive weight system, perform a weighted combination of the system load balance, the access delay score, the resource utilization score and the data locality score to obtain a multi-dimensional system performance comprehensive score, and construct a comprehensive reward function based on the multi-dimensional system performance comprehensive score; A dual-objective Q network structure is used for policy optimization, wherein the online network updates the policy parameters in real time based on the comprehensive reward function, the target network maintains training stability through a soft update mechanism, and outputs the optimized policy parameters as a scheduling decision sequence; A third unit is used to input the scheduling decision sequence into a distributed scheduling execution engine and perform metadata scheduling operations through a two-phase commit protocol; Setting an adaptive monitoring threshold based on the fused feature vector and deploying a distributed monitoring network to collect system operation status data in real time; When the monitoring data deviates from the adaptive monitoring threshold, a difference analysis is performed on the monitoring data and the original fused feature vector, and a hierarchical scheduling optimization strategy is started according to the degree of difference. An incremental adjustment method is used for local anomalies, and the global anomaly triggering strategy is re-optimized. The optimized scheduling effect data is used as a new training sample, and distributed incremental training is performed on the dual-channel neural network and the dual-objective Q network through a federated learning framework to achieve continuous optimization of the scheduling strategy model.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Power grid topology optimization method and system based on search sorting

    CN118539441A

  • Multi-source heterogeneous remote sensing data organization and management method and system, medium and computer program product

    CN118885544A