Intelligent evolution method and system of distributed computing resources based on digital twins
By deploying intelligent perception probes in distributed computing clusters and building a digital twin model for computing resources, combining hybrid entropy optimization algorithms and hierarchical reinforcement learning frameworks, the problems of insufficient resource state perception and poor scheduling strategies in traditional resource management methods are solved, and more efficient resource utilization and system performance improvement are achieved.
Patent Information
- Application Number
- CN202510153895.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-12
AI Technical Summary
Traditional distributed computing resource management methods are difficult to adapt to dynamically changing loads and complex application needs, and lack the integration of refined perception of resource state, modeling of dynamic correlation between computing nodes, and multi-time scale resource utilization prediction.
The intelligent evolution method of distributed computing resources based on digital twins is adopted. By deploying intelligent perception probes on each computing node, using multi-task parallel acquisition algorithm and a two-layer attention mechanism to extract the characteristics of computing resource operation data, a digital twin model of computing resources is constructed, a dynamic correlation diagram of nodes is established, and a hybrid entropy optimization algorithm and a hierarchical reinforcement learning framework are used to optimize resource allocation and scheduling strategies.
It improves resource perception and modeling accuracy, optimizes resource allocation and scheduling efficiency, enhances resource prediction and configuration capabilities, and improves the overall performance of the system.
Smart Images

Figure CN119597493B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to resource evolution technology, and in particular to a distributed computing resource intelligent evolution method and system based on digital twins. Background Art
[0002] Distributed computing systems have become an important infrastructure for processing large-scale data and complex applications. With the rapid development of cloud computing and edge computing, the scale and complexity of distributed computing clusters are increasing. How to efficiently manage and schedule computing resources has become a key challenge. Traditional resource management methods usually rely on static configuration and rule-based scheduling strategies, which are difficult to adapt to dynamically changing loads and complex application requirements. In addition, traditional monitoring methods mainly focus on the performance indicators of a single computing node and lack global perception and collaborative optimization of the resource status of the entire cluster.
[0003] Lack of refined perception of resource status: Traditional monitoring methods have difficulty capturing subtle changes and mutation points in resource status, resulting in the inability to identify potential performance bottlenecks and failure risks in a timely manner. This limits the effectiveness and flexibility of resource management and scheduling strategies.
[0004] Lack of modeling of dynamic relationships between computing nodes: Traditional resource scheduling methods usually treat computing nodes as isolated individuals, ignoring the mutual influence and dependencies between nodes. This leads to the inability to fully utilize the overall computing power of the cluster, which easily leads to resource waste and performance bottlenecks.
[0005] Lack of integration of multi-time scale resource utilization forecasts: Traditional resource forecasting methods usually only focus on a single time scale, making it difficult to balance the accuracy of short-term forecasts with the stability of long-term forecasts. This results in the inability to provide a reliable decision-making basis for resource management and scheduling. Summary of the invention
[0006] The embodiments of the present invention provide a method and system for intelligent evolution of distributed computing resources based on digital twins, which can solve the problems in the prior art.
[0007] According to a first aspect of the embodiments of the present invention,
[0008] Provides a distributed computing resource intelligent evolution method based on digital twins, including:
[0009] Intelligent perception probes are deployed on each computing node in the distributed computing cluster. The intelligent perception probes use a multi-task parallel acquisition algorithm to obtain computing resource operation data. A two-layer attention mechanism is used to extract features from the computing resource operation data, wherein the first-layer attention mechanism identifies the mutation point of the resource state to generate a mutation feature sequence, and the second-layer attention mechanism extracts the resource fluctuation feature based on the mutation feature sequence to obtain a feature tensor. The feature tensor is input into a graph neural network to construct a computing resource digital twin model. The graph neural network integrates the temporal convolution layer and the spatial relationship layer to establish a node dynamic association graph between computing nodes. Based on the node dynamic association graph, a hybrid entropy optimization algorithm is used to calculate the feature importance to obtain a feature weight matrix, which is used to characterize the multi-dimensional correlation of resource states.
[0010] Input the feature weight matrix and the node dynamic association graph into a hierarchical reinforcement learning framework, adopt an improved policy gradient algorithm at the macro level to learn a global resource allocation strategy based on the feature weight matrix, and adopt an improved Q learning algorithm at the micro level to optimize local scheduling decisions based on the node dynamic association graph;
[0011] The global resource allocation strategy and the local scheduling decision are input into the attention-enhanced sequence prediction network to generate a resource utilization prediction sequence for a multi-scale time window; based on the resource utilization prediction sequence, a dynamic fusion mechanism is designed to combine the long-term and short-term prediction results, and an improved Bayesian optimization algorithm is used to calculate the prediction weights of different time scales to obtain a fused prediction result; a multi-objective optimization model is constructed based on the fused prediction result, the system performance indicator is set as the optimization target, and the particle swarm algorithm is used to solve the optimal resource allocation plan and scheduling strategy.
[0012] The feature tensor is input into the graph neural network to construct a digital twin model of computing resources. The graph neural network integrates the temporal convolution layer and the spatial relationship layer to establish a node dynamic association graph between computing nodes. Based on the node dynamic association graph, the hybrid entropy optimization algorithm is used to calculate the feature importance to obtain a feature weight matrix. The feature weight matrix is used to characterize the multi-dimensional correlation of resource status, including:
[0013] The feature tensor is input into a graph neural network that integrates a temporal convolution layer and a spatial relationship layer, wherein the temporal convolution layer adopts a multi-scale parallel convolution structure, and obtains multi-scale temporal features by keeping the sequence length unchanged through causal filling; the spatial relationship layer adopts a dynamic attention mechanism, calculates the attention weights between nodes based on a learnable parameter matrix, and adaptively aggregates the node features according to the attention weights between nodes;
[0014] Based on the output of the graph neural network, a node dynamic association graph between computing nodes is constructed, a Pearson correlation coefficient matrix and a resource call dependency matrix between nodes are calculated through a sliding time window mechanism, the Pearson correlation coefficient matrix and the resource call dependency matrix are fused with an adaptive weight coefficient to obtain a comprehensive association strength matrix, and the comprehensive association strength matrix is dynamically updated based on an exponential decay method to obtain the node dynamic association graph;
[0015] Based on the node dynamic association graph, a hybrid entropy optimization algorithm is used to calculate feature importance. The hybrid entropy optimization algorithm calculates the information entropy of a single feature through kernel density estimation, and then calculates the cross entropy between the feature and the target variable. Conditional mutual information is calculated based on the information entropy and the cross entropy, and the conditional mutual information is normalized to obtain a feature weight matrix;
[0016] Dynamically optimize the feature weight matrix, introduce a time decay factor to adjust the weight according to the timeliness of the feature, calculate the weight gradient based on the prediction error loss, and update the weight using the gradient descent method of the momentum term, where the momentum term is used to smooth the weight update process;
[0017] The updated feature weight matrix is fed back to the graph neural network to guide the dynamic weighted aggregation of node features in the spatial relationship layer to achieve iterative optimization of the digital twin model of computing resources, wherein the feature weight matrix represents the multi-dimensional correlation of the computing resource status.
[0018] Inputting the feature weight matrix and the node dynamic association graph into a hierarchical reinforcement learning framework, using an improved policy gradient algorithm at a macro level to learn a global resource allocation strategy based on the feature weight matrix, and using an improved Q learning algorithm at a micro level to optimize local scheduling decisions based on the node dynamic association graph include:
[0019] Inputting the feature weight matrix and the node dynamic association graph into a hierarchical reinforcement learning framework, wherein the hierarchical reinforcement learning framework includes a macro layer and a micro layer, and the hierarchical reinforcement learning framework realizes the strategy coordination between the macro layer and the micro layer based on a hierarchical reward mechanism, wherein the hierarchical reward mechanism constructs a strategy learning target by a weighted combination of a global reward value and a local reward value;
[0020] At the macro level, an improved policy gradient algorithm is used to learn a global resource allocation strategy based on the feature weight matrix. The improved policy gradient algorithm maps feature weight information to resource allocation action probabilities through a policy network, introduces trust domain constraints to limit the policy update step, and uses a clipping method of the ratio of new and old strategies to construct an optimization objective function, and combines a time-decayed adaptive learning rate to update policy network parameters;
[0021] At the micro level, an improved Q learning algorithm is used to optimize local scheduling decisions based on the node dynamic association graph. The improved Q learning algorithm constructs a dual Q network structure, including a main value network and an auxiliary value network, extracts state features of target nodes and neighbor nodes based on the node dynamic association graph, combines the outputs of the two value networks through a learnable fusion coefficient, and uses a priority experience replay mechanism based on temporal difference error for sample learning;
[0022] A two-way information flow mechanism is used to achieve collaborative decision-making between the macro layer and the micro layer, and the global resource allocation strategy learned by the macro layer is converted into local decision constraints to guide the action selection of the micro layer. At the same time, the local state information of the micro layer is aggregated through pooling operations and fed back to the macro layer;
[0023] A cross-layer value evaluation function is constructed, the global value evaluation of the macro layer and the local value evaluation of the micro layer are weightedly combined, and the global resource allocation strategy and the local scheduling decision are jointly optimized based on the cross-layer value evaluation function.
[0024] Based on the resource utilization prediction sequence, a dynamic fusion mechanism is designed to combine the long-term and short-term prediction results, and the prediction weights of different time scales are calculated using an improved Bayesian optimization algorithm to obtain a fusion prediction result; a multi-objective optimization model is constructed based on the fusion prediction result, and the system performance index is set as the optimization target. The particle swarm algorithm is used to solve the optimal resource allocation plan and scheduling strategy, including:
[0025] Dividing the resource utilization prediction sequence into a short-term prediction sequence and a long-term prediction sequence according to different time spans, wherein the short-term prediction sequence includes prediction values of continuous time points, and the long-term prediction sequence includes prediction values of interval sampling;
[0026] Inputting the short-term prediction sequence and the long-term prediction sequence into an improved Bayesian optimization algorithm, the improved Bayesian optimization algorithm constructs a dynamic Gaussian process kernel function including a time-varying noise term, adaptively adjusts the exploration parameters based on the prediction error, and updates the Gaussian process model parameters online through a sliding time window to obtain the fusion weights of the short-term prediction sequence and the long-term prediction sequence;
[0027] The fusion weight is weightedly combined with the short-term prediction sequence and the long-term prediction sequence to obtain a fusion prediction result, and a multi-objective optimization model is constructed based on the fusion prediction result. The multi-objective optimization model includes a system throughput target, a resource utilization target, and an energy efficiency target, and sets a resource capacity upper limit constraint, a task completion deadline constraint, and a load balancing constraint;
[0028] A particle swarm algorithm is used to solve the multi-objective optimization model. The particle swarm algorithm adaptively adjusts the inertia weight through a nonlinear adjustment factor, introduces the neighborhood optimal solution into the speed update formula to enhance the local search ability, calculates the individual crowding degree based on the target space distance and constructs a selection mechanism, saves non-dominated solutions through an external archive set, and generates the optimal resource allocation plan and scheduling strategy based on the solution results of the particle swarm algorithm.
[0029] The multi-objective optimization model is solved by using a particle swarm algorithm. The particle swarm algorithm adaptively adjusts the inertia weight through a nonlinear adjustment factor, introduces the neighborhood optimal solution into the speed update formula to enhance the local search capability, calculates the individual crowding degree based on the target space distance and constructs a selection mechanism, and saves non-dominated solutions through an external archive set. The optimal resource allocation scheme and scheduling strategy based on the solution results of the particle swarm algorithm include:
[0030] Initialize the position vector and velocity vector of the particle swarm, design a nonlinear adjustment factor, the nonlinear adjustment factor adopts a composite form of a cosine function and a power function, and calculates the current inertia weight according to the ratio of the current number of iterations to the maximum number of iterations;
[0031] A speed update mechanism is constructed based on the current inertia weight, the Euclidean distance of the particle in the target space is calculated to obtain a distance matrix, the neighborhood range of each particle is determined according to the distance matrix, an optimal solution is selected within the neighborhood range, the difference between the optimal solution and the current position is calculated to obtain a neighborhood guide vector, and the neighborhood guide vector is combined with an individual optimal position term and a global optimal position term to form a complete speed update expression;
[0032] Using the speed update expression to update the particle position to obtain a candidate solution set, calculating the difference in function values of adjacent solutions in the target space in the candidate solution set to obtain a crowding index, performing non-dominated sorting on the candidate solution set according to the crowding index, and selecting non-dominated solutions to construct an external archive set;
[0033] Add the newly generated non-dominated solutions to the external archive set and re-sort them. When the external archive set exceeds the preset capacity, sort the archive members in descending order based on the crowding index and retain the top 5% solutions. The external archive set is used to save high-quality solutions with convergence and diversity.
[0034] A comprehensive evaluation value is calculated for the solutions in the external archive set, and the solution with the best comprehensive evaluation value is selected as the final solution. A resource allocation matrix and a task scheduling sequence are generated according to the decision variables corresponding to the final solution. The optimal resource configuration plan and scheduling strategy are determined based on the resource allocation matrix and the task scheduling sequence.
[0035] The method further comprises:
[0036] Based on the optimal resource configuration scheme and the scheduling strategy, dynamic resource adjustment is performed through a distributed feedback controller, and the distributed feedback controller adopts a sliding mode control algorithm to achieve precise control; a multi-level anomaly detection framework is constructed to monitor the execution effect in real time, and the feature weight matrix is used as a benchmark, combined with an improved isolation forest algorithm to identify state deviations, and the deviation propagation path is analyzed based on the node dynamic association graph;
[0037] Constructing a system control target based on the optimal resource configuration scheme and the scheduling strategy, converting the system control target into a state space expression, designing a sliding mode switching function of a distributed feedback controller according to the state space expression, constructing a composite control law of an equivalent control term and a switching control term based on the sliding mode switching function, and generating a resource dynamic adjustment instruction through boundary layer smoothing;
[0038] Execute the resource dynamic adjustment instruction and collect real-time status data, build a multi-level anomaly detection framework including status features, weight matrix and detection threshold, compare the real-time status data with the reference state represented by the feature weight matrix, and calculate the status deviation vector;
[0039] The state deviation vector is input into the improved isolation forest algorithm, a dynamic split probability and an adaptive depth limit are set for the improved isolation forest algorithm according to the importance of resource features, the detection results of multiple isolated trees are integrated by weighted voting, and the location information and degree information of the abnormal state are output;
[0040] The abnormal source node is located in the node dynamic association graph based on the position information, the product of the feature weight and the association strength between adjacent nodes is calculated according to the degree information to obtain the propagation intensity of the state deviation, the node dynamic association graph is used to analyze the diffusion path of the propagation intensity, and the key influencing link is determined.
[0041] The method further comprises:
[0042] An improved meta-learning algorithm is used to optimize the digital twin model based on the updated feature weight matrix and dynamic association graph, and rapid model iteration is achieved through task decomposition and gradient accumulation. Finally, the root cause of performance degradation is analyzed based on the improved causal reasoning model, and the analysis results are fed back to the hierarchical reinforcement learning framework to achieve dynamic optimization of decision-making strategies.
[0043] An improved meta-learning algorithm is used to initialize a feature weight matrix, where each element in the feature weight matrix represents the weight of the input feature on the output indicator. A task-specific weight matrix is obtained by calculating the loss function of each task. A meta-gradient is calculated based on the task-specific weight matrix and a global weight matrix is updated.
[0044] A dynamic association graph is constructed based on the global weight matrix, wherein the nodes of the dynamic association graph represent system components, and the edges represent the association strength between components. Significant association relationships are screened by setting dynamic thresholds, and the association strength is continuously updated by using a sliding time window;
[0045] Decomposing the optimization target into multiple subtasks and calculating the gradients in parallel, assigning a weight coefficient to each subtask according to the degree of association of the components in the dynamic association graph, and updating the parameters of the digital twin model by using a weighted gradient accumulation method;
[0046] The causal graph structure is initialized based on the dynamic association graph, and a bidirectional search algorithm is used to optimize causal relationships and eliminate false associations. The contribution of influencing factors to performance degradation is quantified through counterfactual reasoning, and the contribution is input into a hierarchical reinforcement learning framework as a reward signal. The hierarchical reinforcement learning framework includes a macro layer for selecting macro actions and a micro layer for performing specific operations. The coordinated optimization of the macro layer and the micro layer is achieved through a hierarchical reward mechanism.
[0047] According to a second aspect of the embodiments of the present invention,
[0048] Provides a distributed computing resource intelligent evolution system based on digital twins, including:
[0049] The first unit is used to deploy an intelligent perception probe on each computing node of a distributed computing cluster, and the intelligent perception probe adopts a multi-task parallel acquisition algorithm to obtain computing resource operation data; a two-layer attention mechanism is used to extract features of the computing resource operation data, wherein the first-layer attention mechanism identifies the mutation point of the resource state to generate a mutation feature sequence, and the second-layer attention mechanism extracts the resource fluctuation feature based on the mutation feature sequence to obtain a feature tensor; the feature tensor is input into a graph neural network to construct a computing resource digital twin model, and the graph neural network integrates the temporal convolution layer and the spatial relationship layer to establish a node dynamic association graph between computing nodes; based on the node dynamic association graph, a hybrid entropy optimization algorithm is used to calculate the feature importance to obtain a feature weight matrix, and the feature weight matrix is used to characterize the multi-dimensional correlation of resource states;
[0050] The second unit is used to input the feature weight matrix and the node dynamic association graph into a hierarchical reinforcement learning framework, adopt an improved policy gradient algorithm at a macro level to learn a global resource allocation strategy based on the feature weight matrix, and adopt an improved Q learning algorithm at a micro level to optimize local scheduling decisions based on the node dynamic association graph;
[0051] The third unit is used to input the global resource allocation strategy and the local scheduling decision into the attention-enhanced sequence prediction network to generate a resource utilization prediction sequence for a multi-scale time window; based on the resource utilization prediction sequence, a dynamic fusion mechanism is designed to combine the long-term and short-term prediction results, and an improved Bayesian optimization algorithm is used to calculate the prediction weights of different time scales to obtain a fusion prediction result; a multi-objective optimization model is constructed according to the fusion prediction result, the system performance indicator is set as the optimization target, and the particle swarm algorithm is used to solve and obtain the optimal resource allocation plan and scheduling strategy.
[0052] According to a third aspect of the embodiments of the present invention,
[0053] An electronic device is provided, comprising:
[0054] processor;
[0055] a memory for storing processor-executable instructions;
[0056] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0057] A fourth aspect of the embodiments of the present invention is:
[0058] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0059] The beneficial effects of this application are as follows:
[0060] 1. Improve resource perception and modeling accuracy: The multi-task parallel acquisition algorithm and double-layer attention mechanism can more accurately perceive and extract key features in computing resource operation data. Combined with graph neural networks and time series convolution, a more accurate digital twin model of computing resources can be built to more accurately reflect the resource status and dynamic relationship between nodes.
[0061] 2. Optimize resource allocation and scheduling efficiency: Based on the hierarchical reinforcement learning framework, combined with the global resource allocation strategy and local scheduling decision, it can more effectively utilize computing resources and improve resource utilization and scheduling efficiency. At the same time, the application of the hybrid entropy optimization algorithm can better capture the multi-dimensional correlation of resource status and further improve the level of intelligent decision-making.
[0062] 3. Enhanced resource prediction and configuration capabilities: The attention-enhanced sequence prediction network and dynamic fusion mechanism can generate more accurate resource utilization prediction sequences for multi-scale time windows, providing a reliable basis for resource configuration. Combined with the multi-objective optimization model and particle swarm algorithm, better resource configuration solutions and scheduling strategies can be obtained, thereby improving the overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a flow chart of a method for intelligent evolution of distributed computing resources based on digital twins according to an embodiment of the present invention;
[0064] Figure 2 This is a structural diagram of a distributed computing resource intelligent evolution system based on digital twins according to an embodiment of the present invention. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0066] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0067] Figure 1 Schematic diagram of the process of the distributed computing resource intelligent evolution method based on digital twins according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0068] S11. Deploy an intelligent perception probe on each computing node in the distributed computing cluster, and the intelligent perception probe adopts a multi-task parallel acquisition algorithm to obtain computing resource operation data; use a two-layer attention mechanism to extract features from the computing resource operation data, wherein the first-layer attention mechanism identifies the mutation point of the resource state to generate a mutation feature sequence, and the second-layer attention mechanism extracts the resource fluctuation feature based on the mutation feature sequence to obtain a feature tensor; input the feature tensor into the graph neural network to construct a computing resource digital twin model, and the graph neural network integrates the temporal convolution layer and the spatial relationship layer to establish a node dynamic association graph between computing nodes; based on the node dynamic association graph, a hybrid entropy optimization algorithm is used to calculate the feature importance to obtain a feature weight matrix, and the feature weight matrix is used to characterize the multi-dimensional correlation of resource states;
[0069] S12. Input the feature weight matrix and the node dynamic association graph into a hierarchical reinforcement learning framework, adopt an improved policy gradient algorithm at the macro level to learn the global resource allocation strategy based on the feature weight matrix, and adopt an improved Q learning algorithm at the micro level to optimize the local scheduling decision based on the node dynamic association graph;
[0070] S13. Input the global resource allocation strategy and the local scheduling decision into the attention-enhanced sequence prediction network to generate a resource utilization prediction sequence for a multi-scale time window; based on the resource utilization prediction sequence, design a dynamic fusion mechanism to combine the long-term and short-term prediction results, and use an improved Bayesian optimization algorithm to calculate the prediction weights of different time scales to obtain a fusion prediction result; construct a multi-objective optimization model based on the fusion prediction result, set the system performance indicator as the optimization target, and use the particle swarm algorithm to solve and obtain the optimal resource allocation plan and scheduling strategy.
[0071] In an optional implementation, the feature tensor is input into a graph neural network to construct a computing resource digital twin model, and the graph neural network integrates the temporal convolution layer and the spatial relationship layer to establish a node dynamic association graph between computing nodes; based on the node dynamic association graph, a hybrid entropy optimization algorithm is used to calculate the feature importance to obtain a feature weight matrix, and the feature weight matrix is used to characterize the multi-dimensional correlation of resource status, including:
[0072] The feature tensor is input into a graph neural network that integrates a temporal convolution layer and a spatial relationship layer, wherein the temporal convolution layer adopts a multi-scale parallel convolution structure, and obtains multi-scale temporal features by keeping the sequence length unchanged through causal filling; the spatial relationship layer adopts a dynamic attention mechanism, calculates the attention weights between nodes based on a learnable parameter matrix, and adaptively aggregates the node features according to the attention weights between nodes;
[0073] Based on the output of the graph neural network, a node dynamic association graph between computing nodes is constructed, a Pearson correlation coefficient matrix and a resource call dependency matrix between nodes are calculated through a sliding time window mechanism, the Pearson correlation coefficient matrix and the resource call dependency matrix are fused with an adaptive weight coefficient to obtain a comprehensive association strength matrix, and the comprehensive association strength matrix is dynamically updated based on an exponential decay method to obtain the node dynamic association graph;
[0074] Based on the node dynamic association graph, a hybrid entropy optimization algorithm is used to calculate feature importance. The hybrid entropy optimization algorithm calculates the information entropy of a single feature through kernel density estimation, and then calculates the cross entropy between the feature and the target variable. Conditional mutual information is calculated based on the information entropy and the cross entropy, and the conditional mutual information is normalized to obtain a feature weight matrix;
[0075] Dynamically optimize the feature weight matrix, introduce a time decay factor to adjust the weight according to the timeliness of the feature, calculate the weight gradient based on the prediction error loss, and update the weight using the gradient descent method of the momentum term, where the momentum term is used to smooth the weight update process;
[0076] The updated feature weight matrix is fed back to the graph neural network to guide the dynamic weighted aggregation of node features in the spatial relationship layer to achieve iterative optimization of the digital twin model of computing resources, wherein the feature weight matrix represents the multi-dimensional correlation of the computing resource status.
[0077] The method for constructing a digital twin model of computing resources integrates the temporal and spatial characteristics of computing resources through a graph neural network, and uses a hybrid entropy optimization algorithm to dynamically adjust the feature weights to achieve accurate characterization of the state of computing resources.
[0078] First, collect various performance indicator data of computing resources, such as CPU utilization, memory usage, network traffic, disk IO, etc., and construct these data into feature tensors. The resource indicators at each time point constitute a feature vector, and the feature vectors of multiple time points are arranged in chronological order to form a feature tensor. For example, if the four indicators of CPU utilization, memory usage, network traffic, and disk IO are collected every 5 seconds, and data at 100 time points are collected continuously, a 100x4 feature tensor can be constructed.
[0079] Next, the feature tensor is input into a graph neural network that combines a temporal convolution layer and a spatial relationship layer. The temporal convolution layer uses a multi-scale parallel convolution structure to extract temporal features. For example, three parallel convolution kernels with kernel sizes of 3, 5, and 7 are used to perform convolution operations on the input feature tensor. In order to keep the sequence length unchanged, causal padding is used, that is, 0 is filled at the beginning of the sequence. Through multi-scale convolution, temporal features of different time granularities can be captured. The spatial relationship layer uses a dynamic attention mechanism to adaptively aggregate node features. Assuming there are three computing nodes, the spatial relationship layer dynamically calculates the attention weights based on the correlation between the nodes. For example, if node 1 and node 2 are highly correlated, the attention weight between them is high, and when aggregating features, the features of node 1 and node 2 will be given greater weights.
[0080] The output of the graph neural network is used to construct a node dynamic association graph between computing nodes. Through the sliding time window mechanism, for example, by selecting resource data from the past 10 time points, the Pearson correlation coefficient matrix and the resource call dependency matrix between the nodes are calculated. The Pearson correlation coefficient matrix reflects the linear correlation of changes in resource indicators between nodes. The resource call dependency matrix represents the call relationship between nodes, such as the number of times node 1 calls node 2. Then, the Pearson correlation coefficient matrix and the resource call dependency matrix are fused using adaptive weight coefficients to obtain a comprehensive association strength matrix. Finally, the comprehensive association strength matrix is dynamically updated based on the exponential decay method to obtain a node dynamic association graph. The exponential decay method means that recent data has a greater impact on the association strength.
[0081] Based on the node dynamic association graph, a hybrid entropy optimization algorithm is used to calculate the importance of features. First, the information entropy of a single feature is calculated by kernel density estimation. The higher the information entropy, the more information the feature contains. Then, the cross entropy between the feature and the target variable (for example, the overall performance of the system) is calculated. The lower the cross entropy, the higher the correlation between the feature and the target variable. Next, the conditional mutual information is calculated based on the information entropy and cross entropy. The conditional mutual information indicates the degree to which the uncertainty of the target variable is reduced when the feature is known. Finally, the conditional mutual information is normalized to obtain the feature weight matrix. For example, the conditional mutual information of CPU utilization is the highest, and its corresponding weight value in the feature weight matrix is the largest.
[0082] Dynamically optimize the feature weight matrix. Introduce a time decay factor to adjust the weight according to the timeliness of the feature. For example, if the CPU utilization has a greater impact on system performance in the recent period, its weight will be higher. Calculate the weight gradient based on the prediction error loss, and use the gradient descent method with momentum term to update the weight. The momentum term can smooth the weight update process and avoid oscillation. Feed the updated feature weight matrix back to the graph neural network to guide the dynamic weighted aggregation of node features in the spatial relationship layer, and realize the iterative optimization of the digital twin model of computing resources.
[0083] The solution of this application can:
[0084] Improve model accuracy: By fusing temporal and spatial features and dynamically adjusting feature weights, the multi-dimensional correlation of computing resource status can be more accurately reflected, thereby improving the accuracy of the digital twin model. Enhance model adaptability: The dynamic attention mechanism and time decay factor are used to enable the model to adapt to the dynamic changes in the computing resource status, improving the robustness and generalization ability of the model. Optimize resource management: By building an accurate digital twin model, the operating status of computing resources can be better understood, thereby optimizing resource scheduling and management strategies and improving resource utilization efficiency.
[0085] In an optional implementation, the feature weight matrix and the node dynamic association graph are input into a hierarchical reinforcement learning framework, an improved policy gradient algorithm is used at a macro level to learn a global resource allocation strategy based on the feature weight matrix, and an improved Q learning algorithm is used at a micro level to optimize local scheduling decisions based on the node dynamic association graph, including:
[0086] Inputting the feature weight matrix and the node dynamic association graph into a hierarchical reinforcement learning framework, wherein the hierarchical reinforcement learning framework includes a macro layer and a micro layer, and the hierarchical reinforcement learning framework realizes the strategy coordination between the macro layer and the micro layer based on a hierarchical reward mechanism, wherein the hierarchical reward mechanism constructs a strategy learning target by a weighted combination of a global reward value and a local reward value;
[0087] At the macro level, an improved policy gradient algorithm is used to learn a global resource allocation strategy based on the feature weight matrix. The improved policy gradient algorithm maps feature weight information to resource allocation action probabilities through a policy network, introduces trust domain constraints to limit the policy update step, and uses a clipping method of the ratio of new and old strategies to construct an optimization objective function, and combines a time-decayed adaptive learning rate to update policy network parameters;
[0088] At the micro level, an improved Q learning algorithm is used to optimize local scheduling decisions based on the node dynamic association graph. The improved Q learning algorithm constructs a dual Q network structure, including a main value network and an auxiliary value network, extracts state features of target nodes and neighbor nodes based on the node dynamic association graph, combines the outputs of the two value networks through a learnable fusion coefficient, and uses a priority experience replay mechanism based on temporal difference error for sample learning;
[0089] A two-way information flow mechanism is used to achieve collaborative decision-making between the macro layer and the micro layer, and the global resource allocation strategy learned by the macro layer is converted into local decision constraints to guide the action selection of the micro layer. At the same time, the local state information of the micro layer is aggregated through pooling operations and fed back to the macro layer;
[0090] A cross-layer value evaluation function is constructed, the global value evaluation of the macro layer and the local value evaluation of the micro layer are weightedly combined, and the global resource allocation strategy and the local scheduling decision are jointly optimized based on the cross-layer value evaluation function.
[0091] A resource scheduling method based on hierarchical reinforcement learning is used to optimize resource allocation and task scheduling in complex systems. This method decouples global resource allocation and local task scheduling, and uses improved reinforcement learning algorithms at the macro and micro levels for collaborative optimization.
[0092] First, a hierarchical reinforcement learning framework is constructed, which consists of a macro layer and a micro layer. The macro layer is responsible for global resource allocation, and the micro layer is responsible for local task scheduling. The core of the framework is a hierarchical reward mechanism, which weightedly combines global rewards and local rewards to construct the goal of policy learning. Global rewards measure the overall system performance, such as throughput or latency; local rewards measure the performance of a single node, such as task completion time or energy consumption. For example, the global reward can be set to the total time to complete all tasks, and the local reward can be set to the time to complete the task on each node. The weighting coefficient can be adjusted according to the system requirements. For example, if the overall performance is more concerned, the weight of the global reward can be set higher.
[0093] At the macro level, an improved policy gradient algorithm is used to learn the global resource allocation strategy. The algorithm uses a policy network to map feature weight information to resource allocation action probabilities. The feature weight matrix describes the importance of different resource types to different tasks. For example, a row in the matrix represents a resource type, a column represents a task type, and each element represents the importance of the resource type to the task type. The input of the policy network is the feature weight matrix, and the output is the probability of each resource allocation scheme. In order to stabilize the training process, a trust region constraint is introduced to limit the policy update step size, and the optimization objective function is constructed by clipping the ratio of the new and old strategies. In addition, a time-decayed adaptive learning rate is used to update the policy network parameters to improve learning efficiency. For example, the initial learning rate is set to 0.01 and gradually decreases with the increase in the number of training steps.
[0094] At the micro level, an improved Q-learning algorithm is used to optimize local scheduling decisions. The algorithm constructs a dual Q-network structure, including a main value network and an auxiliary value network. Both networks extract state features of target nodes and neighbor nodes based on the node dynamic association graph. The node dynamic association graph describes the connection relationship and state information between nodes. For example, each node in the graph represents a computing node, the edge represents the communication link between nodes, and the characteristics of the nodes can include CPU utilization, memory occupancy, etc. The outputs of the two value networks are combined through a learnable fusion coefficient to obtain a more accurate value estimate. For example, the fusion coefficient can be initialized to 0.5 and adjusted as the training process progresses. In addition, a priority experience replay mechanism based on temporal difference error is used for sample learning to improve learning efficiency. For example, samples with larger temporal difference errors are given higher priority for faster learning.
[0095] In order to achieve collaborative decision-making at the macro and micro levels, a two-way information flow mechanism is adopted. The global resource allocation strategy learned by the macro level is converted into local decision constraints to guide the action selection of the micro level. For example, the total amount of resources allocated to a node by the macro level can be used as the upper limit of the number of tasks scheduled for the node. At the same time, the local state information of the micro level is aggregated through pooling operations and fed back to the macro level so that the macro level can better perceive the system state. For example, the average CPU utilization of each node can be used as the input feature of the macro level.
[0096] Finally, a cross-layer value evaluation function is constructed to perform a weighted combination of the global value evaluation at the macro level and the local value evaluation at the micro level. Based on this cross-layer value evaluation function, the global resource allocation strategy and local scheduling decision are jointly optimized to achieve the optimal overall system performance.
[0097] The solution of this application can:
[0098] Improve resource utilization: Through the coordinated optimization of global resource allocation and local task scheduling, system resources can be used more effectively to avoid resource waste and bottlenecks. Improve system performance: The hierarchical reinforcement learning framework can dynamically adjust resource allocation and task scheduling strategies according to the system state, thereby improving overall system performance, such as throughput and latency. Enhance system adaptability: This method can adapt to dynamically changing environments and task loads, and automatically learn the optimal resource scheduling strategy to improve the robustness and adaptability of the system.
[0099] In an optional implementation, based on the resource utilization prediction sequence, a dynamic fusion mechanism is designed to combine the long-term and short-term prediction results, and an improved Bayesian optimization algorithm is used to calculate the prediction weights of different time scales to obtain a fusion prediction result; a multi-objective optimization model is constructed based on the fusion prediction result, and the system performance index is set as the optimization target. The particle swarm algorithm is used to solve the optimal resource allocation plan and scheduling strategy, including:
[0100] Dividing the resource utilization prediction sequence into a short-term prediction sequence and a long-term prediction sequence according to different time spans, wherein the short-term prediction sequence includes prediction values of continuous time points, and the long-term prediction sequence includes prediction values of interval sampling;
[0101] Inputting the short-term prediction sequence and the long-term prediction sequence into an improved Bayesian optimization algorithm, the improved Bayesian optimization algorithm constructs a dynamic Gaussian process kernel function including a time-varying noise term, adaptively adjusts the exploration parameters based on the prediction error, and updates the Gaussian process model parameters online through a sliding time window to obtain the fusion weights of the short-term prediction sequence and the long-term prediction sequence;
[0102] The fusion weight is weightedly combined with the short-term prediction sequence and the long-term prediction sequence to obtain a fusion prediction result, and a multi-objective optimization model is constructed based on the fusion prediction result. The multi-objective optimization model includes a system throughput target, a resource utilization target, and an energy efficiency target, and sets a resource capacity upper limit constraint, a task completion deadline constraint, and a load balancing constraint;
[0103] A particle swarm algorithm is used to solve the multi-objective optimization model. The particle swarm algorithm adaptively adjusts the inertia weight through a nonlinear adjustment factor, introduces the neighborhood optimal solution into the speed update formula to enhance the local search ability, calculates the individual crowding degree based on the target space distance and constructs a selection mechanism, saves non-dominated solutions through an external archive set, and generates the optimal resource allocation plan and scheduling strategy based on the solution results of the particle swarm algorithm.
[0104] First, the resource utilization prediction series is divided into time scales. The short-term prediction series selects the last 24 hours of continuous
[0105] In an optional implementation, a particle swarm algorithm is used to solve the multi-objective optimization model. The particle swarm algorithm adaptively adjusts the inertia weight through a nonlinear adjustment factor, introduces the neighborhood optimal solution into the speed update formula to enhance the local search capability, calculates the individual crowding degree based on the target space distance and constructs a selection mechanism, and saves non-dominated solutions through an external archive set. The optimal resource allocation scheme and scheduling strategy based on the solution result of the particle swarm algorithm include:
[0106] Initialize the position vector and velocity vector of the particle swarm, design a nonlinear adjustment factor, the nonlinear adjustment factor adopts a composite form of a cosine function and a power function, and calculates the current inertia weight according to the ratio of the current number of iterations to the maximum number of iterations;
[0107] A speed update mechanism is constructed based on the current inertia weight, the Euclidean distance of the particle in the target space is calculated to obtain a distance matrix, the neighborhood range of each particle is determined according to the distance matrix, an optimal solution is selected within the neighborhood range, the difference between the optimal solution and the current position is calculated to obtain a neighborhood guide vector, and the neighborhood guide vector is combined with an individual optimal position term and a global optimal position term to form a complete speed update expression;
[0108] Using the speed update expression to update the particle position to obtain a candidate solution set, calculating the difference in function values of adjacent solutions in the target space in the candidate solution set to obtain a crowding index, performing non-dominated sorting on the candidate solution set according to the crowding index, and selecting non-dominated solutions to construct an external archive set;
[0109] Add the newly generated non-dominated solutions to the external archive set and re-sort them. When the external archive set exceeds the preset capacity, sort the archive members in descending order based on the crowding index and retain the top 5% solutions. The external archive set is used to save high-quality solutions with convergence and diversity.
[0110] A comprehensive evaluation value is calculated for the solutions in the external archive set, and the solution with the best comprehensive evaluation value is selected as the final solution. A resource allocation matrix and a task scheduling sequence are generated according to the decision variables corresponding to the final solution. The optimal resource configuration plan and scheduling strategy are determined based on the resource allocation matrix and the task scheduling sequence.
[0111] A multi-objective resource allocation and scheduling method based on an improved particle swarm algorithm aims to improve resource utilization and optimize scheduling efficiency. This method enhances the algorithm's search ability and convergence by nonlinearly adjusting the inertia weight, introducing the neighborhood optimal solution, building a selection mechanism based on congestion, and maintaining an external archive set, and ultimately generates the optimal resource allocation solution and scheduling strategy.
[0112] First, the particle swarm is initialized. Each particle represents a resource configuration and scheduling scheme, and its position vector represents the decision variables in the scheme, such as resource allocation ratio, task execution order, etc. The velocity vector represents the movement direction and speed of the particle in the search space. The initial position and velocity are randomly generated to ensure the diversity of the population. For example, assuming there are 10 resources and 5 tasks, the position vector of each particle contains 15 elements, which represent which task each resource is assigned to and the execution order of the tasks. The initial velocity vector also contains 15 elements, and the value range is randomly generated, for example, between -1 and 1.
[0113] Next, a nonlinear adjustment factor is designed to dynamically adjust the inertia weight. The inertia weight affects the degree to which a particle inherits its previous velocity. The current inertia weight is calculated based on the ratio of the current number of iterations to the maximum number of iterations using a composite form of a cosine function and a power function. For example, assuming the maximum number of iterations is 100 and the current number of iterations is 50, the current inertia weight can be calculated to be 0.5 according to the preset formula. This nonlinear adjustment method can balance the capabilities of global and local search. A larger inertia weight in the early stage is conducive to global exploration, while a smaller inertia weight in the later stage promotes local development.
[0114] Then, a speed update mechanism is constructed. First, the Euclidean distance of the particle in the target space is calculated to obtain a distance matrix. The target space consists of multiple objective function values, such as completion time, cost consumption, etc. The distance matrix reflects the similarity between particles. The neighborhood range of each particle is determined according to the distance matrix, for example, the K particles closest to each other are selected as neighbors. The optimal solution is selected within the neighborhood range, and the difference between the optimal solution and the current position is calculated to obtain the neighborhood guidance vector. The neighborhood guidance vector guides the particle to move toward the neighborhood optimal solution. The neighborhood guidance vector is combined with the individual optimal position term and the global optimal position term to form a complete speed update expression. The individual optimal position is the optimal position of the particle itself in history, and the global optimal position is the optimal position in the history of the entire population. The speed update expression comprehensively considers neighborhood guidance, individual experience, and global experience, so that particles can search for the optimal solution more effectively.
[0115] The particle positions are updated with the updated velocity to obtain the candidate solution set. The crowding index is obtained by calculating the difference in function values of adjacent solutions in the target space. The crowding index reflects the distribution density of the solution. Solutions with low crowding are more conducive to maintaining the diversity of the population. Non-dominated sorting is performed on the candidate solution set according to the crowding index. Non-dominated sorting divides solutions into different levels, and the lower the level, the better the solution. Non-dominated solutions are selected to construct an external archive set. The external archive set is used to save non-dominated solutions and represents the optimal solution set currently found.
[0116] The newly generated non-dominated solutions are added to the external archive set and re-sorted. When the external archive set exceeds the preset capacity, the archive members are sorted in descending order based on the crowding index and the top 5% solutions are retained. This selection mechanism can effectively control the size of the archive set and retain representative high-quality solutions.
[0117] Finally, the comprehensive evaluation value is calculated for the solutions in the external archive set. The comprehensive evaluation value can be calculated based on the weights of different objective functions. For example, if completion time is more important than cost consumption, a higher weight can be given to completion time. The solution with the best comprehensive evaluation value is selected as the final solution. The resource allocation matrix and task scheduling sequence are generated according to the decision variables corresponding to the final solution, and the optimal resource allocation scheme and scheduling strategy are finally determined. For example, the position vector of the final solution contains resource allocation and task sorting information, which can be converted into a resource allocation matrix and a task scheduling Gantt chart.
[0118] The solution of this application can:
[0119] Improved solution efficiency: By nonlinearly adjusting the inertia weight and introducing the optimal solution in the neighborhood, the algorithm's local search capability is enhanced, and the optimal solution can be found more quickly. Enhanced solution quality: Based on the target space distance, the individual congestion degree is calculated and the selection mechanism is constructed, as well as the maintenance strategy of the external archive set, which ensures the convergence and diversity of the algorithm and can find better solutions. Good adaptability: This method is applicable to various multi-objective resource allocation and scheduling problems, and the objective function and parameters can be adjusted according to actual needs, which has strong practical value.
[0120] In an optional implementation, the method further includes:
[0121] Based on the optimal resource configuration scheme and the scheduling strategy, dynamic resource adjustment is performed through a distributed feedback controller, and the distributed feedback controller adopts a sliding mode control algorithm to achieve precise control; a multi-level anomaly detection framework is constructed to monitor the execution effect in real time, and the feature weight matrix is used as a benchmark, combined with an improved isolation forest algorithm to identify state deviations, and the deviation propagation path is analyzed based on the node dynamic association graph;
[0122] Constructing a system control target based on the optimal resource configuration scheme and the scheduling strategy, converting the system control target into a state space expression, designing a sliding mode switching function of a distributed feedback controller according to the state space expression, constructing a composite control law of an equivalent control term and a switching control term based on the sliding mode switching function, and generating a resource dynamic adjustment instruction through boundary layer smoothing;
[0123] Execute the resource dynamic adjustment instruction and collect real-time status data, build a multi-level anomaly detection framework including status features, weight matrix and detection threshold, compare the real-time status data with the reference state represented by the feature weight matrix, and calculate the status deviation vector;
[0124] The state deviation vector is input into the improved isolation forest algorithm, a dynamic split probability and an adaptive depth limit are set for the improved isolation forest algorithm according to the importance of resource features, the detection results of multiple isolated trees are integrated by weighted voting, and the location information and degree information of the abnormal state are output;
[0125] The abnormal source node is located in the node dynamic association graph based on the position information, the product of the feature weight and the association strength between adjacent nodes is calculated according to the degree information to obtain the propagation intensity of the state deviation, the node dynamic association graph is used to analyze the diffusion path of the propagation intensity, and the key influencing link is determined.
[0126] First, obtain the current resource status information of the system, including the CPU usage, memory usage, network bandwidth usage, etc. of each node. For example, the CPU usage of node 1 is 70%, the memory usage is 80%, and the network bandwidth usage is 50%; the CPU usage of node 2 is 60%, the memory usage is 70%, and the network bandwidth usage is 60%.
[0127] Then, according to the system load and service quality requirements, a resource allocation plan and scheduling strategy are formulated. For example, high-priority tasks are scheduled to nodes with sufficient resources, and low-priority tasks are scheduled to nodes with relatively idle resources.
[0128] Next, based on the optimal resource allocation scheme and scheduling strategy, the system control objectives are constructed. For example, the CPU usage of node 1 is controlled below 75%, the memory usage is controlled below 85%, and the network bandwidth usage is controlled below 60%.
[0129] The system control objective is converted into a state space expression, and the sliding mode switching function of the distributed feedback controller is designed. The sliding mode switching function is used to describe the deviation between the system state and the expected state.
[0130] Based on the sliding mode switching function, a composite control law of equivalent control items and switching control items is constructed. The equivalent control item is used to maintain the system near the desired state, and the switching control item is used to overcome the uncertainty and interference of the system. Through boundary layer smoothing, dynamic resource adjustment instructions are generated. For example, the CPU allocation quota of node 1 is reduced, and the memory allocation quota of node 2 is increased.
[0131] Execute resource dynamic adjustment instructions and collect real-time status data. For example, the CPU usage of node 1 is reduced to 72%, the memory usage is reduced to 78%, and the network bandwidth usage is reduced to 55%.
[0132] A multi-level anomaly detection framework is constructed that includes state features, weight matrix, and detection threshold. State features include CPU usage, memory usage, network bandwidth usage, etc. The weight matrix indicates the importance of different state features. The detection threshold is used to determine whether the system state is abnormal.
[0133] The real-time status data is compared with the baseline status represented by the feature weight matrix to calculate the status deviation vector. For example, the CPU usage deviation of node 1 is -3%, the memory usage deviation is -2%, and the network bandwidth usage deviation is -5%.
[0134] The state deviation vector is input into the improved isolation forest algorithm. According to the importance of resource features, dynamic split probability and adaptive depth limit are set for the improved isolation forest algorithm. For example, if the importance of CPU usage is higher than that of memory usage, the split probability corresponding to CPU usage is higher and the depth limit is smaller.
[0135] The detection results of multiple isolated trees are fused by weighted voting to output the location and degree information of abnormal status. For example, the CPU usage of node 1 is abnormal and the degree is slight.
[0136] The abnormal source node is located in the node dynamic association graph based on the location information. For example, node 1 is the abnormal source node.
[0137] The product of the feature weight and the association strength between adjacent nodes is calculated based on the degree information to obtain the propagation strength of the state deviation. For example, the propagation strength between node 1 and node 2 is 0.2.
[0138] The node dynamic association graph is used to analyze the diffusion path of the propagation intensity and determine the key influencing links. For example, the abnormal state of node 1 may spread to nodes 2 and 3.
[0139] The solution of this application can:
[0140] Improve resource utilization: By dynamically adjusting resource allocation, resource waste can be effectively avoided and overall resource utilization can be improved. Ensure service quality: Through precise control algorithms and multi-level anomaly detection frameworks, abnormal situations can be discovered and handled in a timely manner to ensure service stability and reliability. Enhance system robustness: By analyzing the deviation propagation path and determining the key influencing links, the spread of abnormal conditions can be effectively suppressed, and the robustness and fault tolerance of the system can be enhanced.
[0141] In an optional implementation, the method further includes:
[0142] An improved meta-learning algorithm is used to optimize the digital twin model based on the updated feature weight matrix and dynamic association graph, and rapid model iteration is achieved through task decomposition and gradient accumulation. Finally, the root cause of performance degradation is analyzed based on the improved causal reasoning model, and the analysis results are fed back to the hierarchical reinforcement learning framework to achieve dynamic optimization of decision-making strategies.
[0143] An improved meta-learning algorithm is used to initialize a feature weight matrix, where each element in the feature weight matrix represents the weight of the input feature on the output indicator. A task-specific weight matrix is obtained by calculating the loss function of each task. A meta-gradient is calculated based on the task-specific weight matrix and a global weight matrix is updated.
[0144] A dynamic association graph is constructed based on the global weight matrix, wherein the nodes of the dynamic association graph represent system components, and the edges represent the association strength between components. Significant association relationships are screened by setting dynamic thresholds, and the association strength is continuously updated by using a sliding time window;
[0145] Decomposing the optimization target into multiple subtasks and calculating the gradients in parallel, assigning a weight coefficient to each subtask according to the degree of association of the components in the dynamic association graph, and updating the parameters of the digital twin model by using a weighted gradient accumulation method;
[0146] The causal graph structure is initialized based on the dynamic association graph, and a bidirectional search algorithm is used to optimize causal relationships and eliminate false associations. The contribution of influencing factors to performance degradation is quantified through counterfactual reasoning, and the contribution is input into a hierarchical reinforcement learning framework as a reward signal. The hierarchical reinforcement learning framework includes a macro layer for selecting macro actions and a micro layer for performing specific operations. The coordinated optimization of the macro layer and the micro layer is achieved through a hierarchical reward mechanism.
[0147] First, collect system operation data. Taking a manufacturing execution system as an example, the collected data includes equipment status data (e.g., operating speed, temperature, pressure, etc.), production process data (e.g., output, qualified rate, energy consumption, etc.), and environmental data (e.g., temperature, humidity, etc.). Assume that 10,000 data samples are collected, each containing 10 features, such as the rotation speed of equipment A, the temperature of equipment B, the qualified rate of the product, etc.
[0148] Then, the improved meta-learning algorithm is used to initialize the feature weight matrix. Each row of the matrix represents a feature, and each column represents an output indicator. Each element in the initial matrix can be set to the same value, such as 0.1, indicating that each input feature has the same weight on each output indicator. Collect data for multiple tasks, such as data from different production batches or data under different operating conditions. Suppose that data for 5 tasks are collected, each task contains 2000 samples. For each task, a loss function, such as mean squared error, is calculated to evaluate the difference between the model's predicted value and the true value. According to the loss function of each task, the corresponding feature weight matrix is adjusted. For example, if a feature contributes more to the loss function of a task, the corresponding element value of the feature in the weight matrix of the task is increased. Based on the weight matrix specific to each task, the meta-gradient is calculated and the global weight matrix is updated. The meta-gradient represents the update direction and size of the global weight matrix, with the goal of making the model have better performance on all tasks.
[0149] Next, a dynamic association graph is constructed based on the global weight matrix. The nodes of the dynamic association graph represent system components, such as devices, sensors, controllers, etc. The edges represent the strength of association between components. The strength of association can be determined by calculating the correlation coefficient between the feature weight vectors corresponding to the two components. A dynamic threshold is set, such as 0.8, to filter significant associations and remove edges with association strengths below the threshold. A sliding time window is used, such as a time window containing the past 100 data samples, to continuously update the association strength to reflect the dynamic changes in the relationship between system components. For example, if the rotation speed of device A and the temperature of device B are highly correlated in the past 100 samples, the strength of association between them is high.
[0150] After that, the optimization goal is decomposed into multiple subtasks and the gradients are calculated in parallel. For example, the goal of improving production efficiency is decomposed into subtasks such as improving equipment utilization, reducing production time, and improving product quality. A weight coefficient is assigned to each subtask according to the degree of association of the components in the dynamic association graph. For example, if the association strength between the components involved in a subtask is high, a larger weight coefficient is assigned to the subtask. The parameters of the digital twin model are updated using a weighted gradient accumulation method. For example, the gradient calculated for each subtask is multiplied by the corresponding weight coefficient, and then the weighted gradients of all subtasks are accumulated to update the parameters of the digital twin model.
[0151] Then, the causal graph structure is initialized based on the dynamic association graph. The nodes and edges of the causal graph are the same as those of the dynamic association graph, but the edges represent causal relationships. A bidirectional search algorithm is used to optimize causal relationships and eliminate false associations. For example, if there is a correlation between the rotation speed of device A and the temperature of device B, but it is not a causal relationship, then the edge between them is removed. The contribution of influencing factors to performance degradation is quantified through counterfactual reasoning. For example, the impact of the failure of device A on production efficiency is simulated to determine the contribution of the failure of device A to performance degradation. The contribution degree is input into the hierarchical reinforcement learning framework as a reward signal. The hierarchical reinforcement learning framework includes a macro layer for selecting macro actions and a micro layer for performing specific operations. For example, the macro layer selects a macro action to increase production, and the micro layer selects a specific operation to increase the rotation speed of device A. The coordinated optimization of the macro layer and the micro layer is achieved through a hierarchical reward mechanism. For example, if the macro action selected by the macro layer and the specific operation selected by the micro layer can improve production efficiency, a positive reward is given; otherwise, a negative reward is given.
[0152] Finally, the digital twin model is optimized based on the updated feature weight matrix and dynamic association graph, and the model is quickly iterated through task decomposition and gradient accumulation. The above steps are repeated until the model performance reaches the predetermined target.
[0153] The solution of this application can:
[0154] Improve model accuracy: Optimize the feature weight matrix through meta-learning algorithms, capture the complex relationship between features and output indicators, and improve the prediction accuracy of digital twin models. The construction and update of dynamic association graphs can better reflect the dynamic interactions between system components and further improve the accuracy of the model. The application of causal reasoning models can identify the key factors of performance degradation and provide more accurate guidance for optimizing decisions. Accelerate model iteration: The strategies of task decomposition and gradient accumulation can parallelize the model training process, significantly speed up the model iteration speed, and shorten the model optimization cycle. The application of dynamic association graphs can effectively guide the weight allocation of subtasks and improve the efficiency of model training. Enhance decision optimization: The hierarchical reinforcement learning framework can dynamically adjust the decision strategy according to the analysis results of the causal reasoning model to achieve coordinated optimization of the macro and micro levels. Through continuous learning and optimization, the performance of complex systems can be improved.
[0155] Figure 2 Schematic diagram of the structure of the distributed computing resource intelligent evolution system based on digital twins according to an embodiment of the present invention. Figure 2 As shown, the system comprises:
[0156] The first unit is used to deploy an intelligent perception probe on each computing node of a distributed computing cluster, and the intelligent perception probe adopts a multi-task parallel acquisition algorithm to obtain computing resource operation data; a two-layer attention mechanism is used to extract features of the computing resource operation data, wherein the first-layer attention mechanism identifies the mutation point of the resource state to generate a mutation feature sequence, and the second-layer attention mechanism extracts the resource fluctuation feature based on the mutation feature sequence to obtain a feature tensor; the feature tensor is input into a graph neural network to construct a computing resource digital twin model, and the graph neural network integrates the temporal convolution layer and the spatial relationship layer to establish a node dynamic association graph between computing nodes; based on the node dynamic association graph, a hybrid entropy optimization algorithm is used to calculate the feature importance to obtain a feature weight matrix, and the feature weight matrix is used to characterize the multi-dimensional correlation of resource states;
[0157] The second unit is used to input the feature weight matrix and the node dynamic association graph into a hierarchical reinforcement learning framework, adopt an improved policy gradient algorithm at a macro level to learn a global resource allocation strategy based on the feature weight matrix, and adopt an improved Q learning algorithm at a micro level to optimize local scheduling decisions based on the node dynamic association graph;
[0158] The third unit is used to input the global resource allocation strategy and the local scheduling decision into the attention-enhanced sequence prediction network to generate a resource utilization prediction sequence for a multi-scale time window; based on the resource utilization prediction sequence, a dynamic fusion mechanism is designed to combine the long-term and short-term prediction results, and an improved Bayesian optimization algorithm is used to calculate the prediction weights of different time scales to obtain a fusion prediction result; a multi-objective optimization model is constructed according to the fusion prediction result, the system performance indicator is set as the optimization target, and the particle swarm algorithm is used to solve and obtain the optimal resource allocation plan and scheduling strategy.
[0159] According to a third aspect of the embodiments of the present invention,
[0160] An electronic device is provided, comprising:
[0161] processor;
[0162] a memory for storing processor-executable instructions;
[0163] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0164] A fourth aspect of the embodiments of the present invention is:
[0165] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0166] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0167] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A distributed computing resource intelligent evolution method based on digital twins, characterized in that: include: Deploy an intelligent sensing probe on each computing node of the distributed computing cluster, and the intelligent sensing probe uses a multi-task parallel acquisition algorithm to obtain computing resource operation data; A double-layer attention mechanism is used to extract features from the computing resource operation data, wherein the first-layer attention mechanism identifies the mutation points of resource status to generate a mutation feature sequence, and the second-layer attention mechanism extracts resource fluctuation features based on the mutation feature sequence to obtain a feature tensor; the feature tensor is input into a graph neural network to construct a computing resource digital twin model, and the graph neural network integrates the temporal convolution layer and the spatial relationship layer to establish a node dynamic association graph between computing nodes; based on the node dynamic association graph, a hybrid entropy optimization algorithm is used to calculate the feature importance to obtain a feature weight matrix, and the feature weight matrix is used to characterize the multi-dimensional correlation of resource status; Input the feature weight matrix and the node dynamic association graph into a hierarchical reinforcement learning framework, adopt an improved policy gradient algorithm at the macro level to learn a global resource allocation strategy based on the feature weight matrix, and adopt an improved Q learning algorithm at the micro level to optimize local scheduling decisions based on the node dynamic association graph; The global resource allocation strategy and the local scheduling decision are input into the attention-enhanced sequence prediction network to generate a resource utilization prediction sequence for a multi-scale time window; based on the resource utilization prediction sequence, a dynamic fusion mechanism is designed to combine the long-term and short-term prediction results, and an improved Bayesian optimization algorithm is used to calculate the prediction weights of different time scales to obtain a fused prediction result; a multi-objective optimization model is constructed based on the fused prediction result, the system performance indicator is set as the optimization target, and the particle swarm algorithm is used to solve the optimal resource allocation plan and scheduling strategy.
2. The method according to claim 1, characterized in that The feature tensor is input into the graph neural network to construct a digital twin model of computing resources. The graph neural network integrates the temporal convolution layer and the spatial relationship layer to establish a node dynamic association graph between computing nodes. Based on the node dynamic association graph, the hybrid entropy optimization algorithm is used to calculate the feature importance to obtain a feature weight matrix. The feature weight matrix is used to characterize the multi-dimensional correlation of resource status, including: The feature tensor is input into a graph neural network that integrates a temporal convolution layer and a spatial relationship layer, wherein the temporal convolution layer adopts a multi-scale parallel convolution structure, and obtains multi-scale temporal features by keeping the sequence length unchanged through causal filling; the spatial relationship layer adopts a dynamic attention mechanism, calculates the attention weights between nodes based on a learnable parameter matrix, and adaptively aggregates the node features according to the attention weights between nodes; Based on the output of the graph neural network, a node dynamic association graph between computing nodes is constructed, a Pearson correlation coefficient matrix and a resource call dependency matrix between nodes are calculated through a sliding time window mechanism, the Pearson correlation coefficient matrix and the resource call dependency matrix are fused with an adaptive weight coefficient to obtain a comprehensive association strength matrix, and the comprehensive association strength matrix is dynamically updated based on an exponential decay method to obtain the node dynamic association graph; Based on the node dynamic association graph, a hybrid entropy optimization algorithm is used to calculate feature importance. The hybrid entropy optimization algorithm calculates the information entropy of a single feature through kernel density estimation, and then calculates the cross entropy between the feature and the target variable. Conditional mutual information is calculated based on the information entropy and the cross entropy, and the conditional mutual information is normalized to obtain a feature weight matrix; Dynamically optimize the feature weight matrix, introduce a time decay factor to adjust the weight according to the timeliness of the feature, calculate the weight gradient based on the prediction error loss, and update the weight using the gradient descent method of the momentum term, where the momentum term is used to smooth the weight update process; The updated feature weight matrix is fed back to the graph neural network to guide the dynamic weighted aggregation of node features in the spatial relationship layer to achieve iterative optimization of the digital twin model of computing resources, wherein the feature weight matrix represents the multi-dimensional correlation of the computing resource status.
3. The method according to claim 1, characterized in that Inputting the feature weight matrix and the node dynamic association graph into a hierarchical reinforcement learning framework, using an improved policy gradient algorithm at a macro level to learn a global resource allocation strategy based on the feature weight matrix, and using an improved Q learning algorithm at a micro level to optimize local scheduling decisions based on the node dynamic association graph include: Inputting the feature weight matrix and the node dynamic association graph into a hierarchical reinforcement learning framework, wherein the hierarchical reinforcement learning framework includes a macro layer and a micro layer, and the hierarchical reinforcement learning framework realizes the strategy coordination between the macro layer and the micro layer based on a hierarchical reward mechanism, wherein the hierarchical reward mechanism constructs a strategy learning target by a weighted combination of a global reward value and a local reward value; At the macro level, an improved policy gradient algorithm is used to learn a global resource allocation strategy based on the feature weight matrix. The improved policy gradient algorithm maps feature weight information to resource allocation action probabilities through a policy network, introduces trust domain constraints to limit the policy update step, and uses a clipping method of the ratio of new and old strategies to construct an optimization objective function, and combines a time-decayed adaptive learning rate to update policy network parameters; At the micro level, an improved Q learning algorithm is used to optimize local scheduling decisions based on the node dynamic association graph. The improved Q learning algorithm constructs a dual Q network structure, including a main value network and an auxiliary value network, extracts state features of target nodes and neighbor nodes based on the node dynamic association graph, combines the outputs of the two value networks through a learnable fusion coefficient, and uses a priority experience replay mechanism based on temporal difference error for sample learning; A two-way information flow mechanism is used to achieve collaborative decision-making between the macro layer and the micro layer, and the global resource allocation strategy learned by the macro layer is converted into local decision constraints to guide the action selection of the micro layer. At the same time, the local state information of the micro layer is aggregated through pooling operations and fed back to the macro layer; A cross-layer value evaluation function is constructed, the global value evaluation of the macro layer and the local value evaluation of the micro layer are weightedly combined, and the global resource allocation strategy and the local scheduling decision are jointly optimized based on the cross-layer value evaluation function.
4. The method according to claim 1, characterized in that: Based on the resource utilization prediction sequence, a dynamic fusion mechanism is designed to combine the long-term and short-term prediction results, and the prediction weights of different time scales are calculated using an improved Bayesian optimization algorithm to obtain a fusion prediction result; a multi-objective optimization model is constructed based on the fusion prediction result, and the system performance index is set as the optimization target. The particle swarm algorithm is used to solve the optimal resource allocation plan and scheduling strategy, including: Dividing the resource utilization prediction sequence into a short-term prediction sequence and a long-term prediction sequence according to different time spans, wherein the short-term prediction sequence includes prediction values of continuous time points, and the long-term prediction sequence includes prediction values of interval sampling; Inputting the short-term prediction sequence and the long-term prediction sequence into an improved Bayesian optimization algorithm, the improved Bayesian optimization algorithm constructs a dynamic Gaussian process kernel function including a time-varying noise term, adaptively adjusts the exploration parameters based on the prediction error, and updates the Gaussian process model parameters online through a sliding time window to obtain the fusion weights of the short-term prediction sequence and the long-term prediction sequence; The fusion weight is weightedly combined with the short-term prediction sequence and the long-term prediction sequence to obtain a fusion prediction result, and a multi-objective optimization model is constructed based on the fusion prediction result. The multi-objective optimization model includes a system throughput target, a resource utilization target, and an energy efficiency target, and sets a resource capacity upper limit constraint, a task completion deadline constraint, and a load balancing constraint; A particle swarm algorithm is used to solve the multi-objective optimization model. The particle swarm algorithm adaptively adjusts the inertia weight through a nonlinear adjustment factor, introduces the neighborhood optimal solution into the speed update formula to enhance the local search ability, calculates the individual crowding degree based on the target space distance and constructs a selection mechanism, saves non-dominated solutions through an external archive set, and generates the optimal resource allocation plan and scheduling strategy based on the solution results of the particle swarm algorithm.
5. The method according to claim 4, characterized in that The multi-objective optimization model is solved by using a particle swarm algorithm. The particle swarm algorithm adaptively adjusts the inertia weight through a nonlinear adjustment factor, introduces the neighborhood optimal solution into the speed update formula to enhance the local search capability, calculates the individual crowding degree based on the target space distance and constructs a selection mechanism, and saves non-dominated solutions through an external archive set. The optimal resource allocation scheme and scheduling strategy based on the solution results of the particle swarm algorithm include: Initialize the position vector and velocity vector of the particle swarm, design a nonlinear adjustment factor, the nonlinear adjustment factor adopts a composite form of a cosine function and a power function, and calculates the current inertia weight according to the ratio of the current number of iterations to the maximum number of iterations; A speed update mechanism is constructed based on the current inertia weight, the Euclidean distance of the particle in the target space is calculated to obtain a distance matrix, the neighborhood range of each particle is determined according to the distance matrix, an optimal solution is selected within the neighborhood range, the difference between the optimal solution and the current position is calculated to obtain a neighborhood guide vector, and the neighborhood guide vector is combined with an individual optimal position term and a global optimal position term to form a complete speed update expression; Using the speed update expression to update the particle position to obtain a candidate solution set, calculating the difference in function values of adjacent solutions in the target space in the candidate solution set to obtain a crowding index, performing non-dominated sorting on the candidate solution set according to the crowding index, and selecting non-dominated solutions to construct an external archive set; Add the newly generated non-dominated solutions to the external archive set and re-sort them. When the external archive set exceeds the preset capacity, sort the archive members in descending order based on the crowding index and retain the top 5% solutions. The external archive set is used to save high-quality solutions with convergence and diversity. A comprehensive evaluation value is calculated for the solutions in the external archive set, and the solution with the best comprehensive evaluation value is selected as the final solution. A resource allocation matrix and a task scheduling sequence are generated according to the decision variables corresponding to the final solution. The optimal resource configuration plan and scheduling strategy are determined based on the resource allocation matrix and the task scheduling sequence.
6. The method according to claim 1, characterized in that The method further comprises: Based on the optimal resource configuration scheme and the scheduling strategy, dynamic resource adjustment is performed through a distributed feedback controller, and the distributed feedback controller adopts a sliding mode control algorithm to achieve precise control; a multi-level anomaly detection framework is constructed to monitor the execution effect in real time, and the feature weight matrix is used as a benchmark, combined with an improved isolation forest algorithm to identify state deviations, and the deviation propagation path is analyzed based on the node dynamic association graph; Constructing a system control target based on the optimal resource configuration scheme and the scheduling strategy, converting the system control target into a state space expression, designing a sliding mode switching function of a distributed feedback controller according to the state space expression, constructing a composite control law of an equivalent control term and a switching control term based on the sliding mode switching function, and generating a resource dynamic adjustment instruction through boundary layer smoothing; Execute the resource dynamic adjustment instruction and collect real-time status data, build a multi-level anomaly detection framework including status features, weight matrix and detection threshold, compare the real-time status data with the reference state represented by the feature weight matrix, and calculate the status deviation vector; The state deviation vector is input into the improved isolation forest algorithm, a dynamic split probability and an adaptive depth limit are set for the improved isolation forest algorithm according to the importance of resource features, the detection results of multiple isolated trees are integrated by weighted voting, and the location information and degree information of the abnormal state are output; The abnormal source node is located in the node dynamic association graph based on the position information, the product of the feature weight and the association strength between adjacent nodes is calculated according to the degree information to obtain the propagation intensity of the state deviation, the node dynamic association graph is used to analyze the diffusion path of the propagation intensity, and the key influencing link is determined.
7. The method according to claim 1, characterized in that The method further comprises: An improved meta-learning algorithm is used to optimize the digital twin model based on the updated feature weight matrix and dynamic association graph, and rapid model iteration is achieved through task decomposition and gradient accumulation. Finally, the root cause of performance degradation is analyzed based on the improved causal reasoning model, and the analysis results are fed back to the hierarchical reinforcement learning framework to achieve dynamic optimization of decision-making strategies. An improved meta-learning algorithm is used to initialize a feature weight matrix, where each element in the feature weight matrix represents the weight of the input feature on the output indicator. A task-specific weight matrix is obtained by calculating the loss function of each task. A meta-gradient is calculated based on the task-specific weight matrix and a global weight matrix is updated. A dynamic association graph is constructed based on the global weight matrix, wherein the nodes of the dynamic association graph represent system components, and the edges represent the association strength between components. Significant association relationships are screened by setting dynamic thresholds, and the association strength is continuously updated by using a sliding time window; Decomposing the optimization target into multiple subtasks and calculating the gradients in parallel, assigning a weight coefficient to each subtask according to the degree of association of the components in the dynamic association graph, and updating the parameters of the digital twin model by using a weighted gradient accumulation method; The causal graph structure is initialized based on the dynamic association graph, and a bidirectional search algorithm is used to optimize causal relationships and eliminate false associations. The contribution of influencing factors to performance degradation is quantified through counterfactual reasoning, and the contribution is input into a hierarchical reinforcement learning framework as a reward signal. The hierarchical reinforcement learning framework includes a macro layer for selecting macro actions and a micro layer for performing specific operations. The coordinated optimization of the macro layer and the micro layer is achieved through a hierarchical reward mechanism.
8. A distributed computing resource intelligent evolution system based on digital twins, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to deploy an intelligent sensing probe on each computing node of the distributed computing cluster, and the intelligent sensing probe adopts a multi-task parallel acquisition algorithm to obtain computing resource operation data; A double-layer attention mechanism is used to extract features from the computing resource operation data, wherein the first-layer attention mechanism identifies the mutation points of resource status to generate a mutation feature sequence, and the second-layer attention mechanism extracts resource fluctuation features based on the mutation feature sequence to obtain a feature tensor; the feature tensor is input into a graph neural network to construct a computing resource digital twin model, and the graph neural network integrates the temporal convolution layer and the spatial relationship layer to establish a node dynamic association graph between computing nodes; based on the node dynamic association graph, a hybrid entropy optimization algorithm is used to calculate the feature importance to obtain a feature weight matrix, and the feature weight matrix is used to characterize the multi-dimensional correlation of resource status; The second unit is used to input the feature weight matrix and the node dynamic association graph into a hierarchical reinforcement learning framework, adopt an improved policy gradient algorithm at a macro level to learn a global resource allocation strategy based on the feature weight matrix, and adopt an improved Q learning algorithm at a micro level to optimize local scheduling decisions based on the node dynamic association graph; The third unit is used to input the global resource allocation strategy and the local scheduling decision into the attention-enhanced sequence prediction network to generate a resource utilization prediction sequence for a multi-scale time window; based on the resource utilization prediction sequence, a dynamic fusion mechanism is designed to combine the long-term and short-term prediction results, and an improved Bayesian optimization algorithm is used to calculate the prediction weights of different time scales to obtain a fusion prediction result; a multi-objective optimization model is constructed according to the fusion prediction result, the system performance indicator is set as the optimization target, and the particle swarm algorithm is used to solve and obtain the optimal resource allocation plan and scheduling strategy.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Fusion networking method and system based on satellite communication and short-wave communication
CN118233936A
Automatic operation and maintenance method and system for power distribution network based on artificial intelligence
CN119090490A