A load balancing method for SDN controller based on graph neural network and evolutionary algorithm

By building a graph model and using graph neural networks and evolutionary algorithms, the performance bottleneck of a single controller in the SDN network is solved, load balancing and rapid adaptation to dynamic network environments are achieved, and network performance is improved.

CN120434126BActive Publication Date: 2025-09-26QINGCHUANG WANGYU (HEFEI) TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510949074.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-09-26
Estimated Expiration
2045-07-10

AI Technical Summary

Technical Problem

In an SDN network, a single controller can easily become a performance bottleneck, leading to increased network latency. Therefore, how to reasonably deploy multiple SDN controllers and achieve load balancing becomes the core issue for achieving efficient operation.

Method used

A method based on graph neural networks and evolutionary algorithms is adopted. By building a graph model, a multi-layer graph convolutional network is used for feature aggregation and updating. Combined with meta-reinforcement learning and covariance matrix adaptation evolutionary algorithm, the scheduling strategy is adjusted in real time to achieve load balancing.

Benefits of technology

It realizes the rapid adaptation and load balancing of SDN networks in the face of dynamic network environments, finds the global optimal solution through multi-agent collaboration mechanism and real-time feedback updates, and improves network performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120434126B_ABST
    Figure CN120434126B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computer network optimization and controller load balancing, and discloses an SDN controller load balancing method based on graph neural network and evolutionary algorithm. According to the state vector containing various operating indicators of the controller and the adjacency matrix containing the connection relationship between each controller, a graph model is constructed and spliced ​​to obtain a feature matrix; the adjacency matrix is ​​normalized, and the feature matrix is ​​aggregated and updated using a multi-layer graph convolutional network to output a high-dimensional embedding vector, and the final state representation and reward function of the controller are set. A scheduling strategy that adapts to the dynamic network environment is obtained through meta-reinforcement learning and actor-critic network training; a global search is performed on the parameters of the meta-reinforcement learning in a simulation environment to find the optimal solution; the latest state of each controller is collected in real time, the embedding is updated through a pre-trained GCN, and the trained meta-reinforcement learning model is used to output the scheduling strategy according to the global state, and Q-learning is used to perform online parameter fine-tuning and update of the network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of computer network optimization and controller load balancing, and in particular to an SDN controller load balancing method based on graph neural networks and evolutionary algorithms. Background Art

[0002] With the continuous growth of network traffic, traditional static network architecture has deficiencies in flexibility and adaptability. As a network architecture, SDN achieves centralized control and flexible programming by separating the control plane from the data plane. However, the controller under the SDN architecture is the core of the entire network and must process a large number of events and instructions. A single controller can easily become a performance bottleneck, resulting in increased network latency. How to reasonably deploy multiple SDN controllers and achieve load balancing has become a core issue for achieving efficient operation of SDN networks. Summary of the Invention

[0003] In order to reasonably deploy multiple SDN controllers and achieve load balancing, this application provides an SDN controller load balancing method based on graph neural network and evolutionary algorithm, which adopts the following technical solutions:

[0004] A SDN controller load balancing method based on graph neural network and evolutionary algorithm includes the following steps:

[0005] S1. Based on the state vector containing various operating indicators of the controller and the adjacency matrix containing the connection relationship between each controller, a graphical model is constructed and spliced ​​to obtain a feature matrix;

[0006] S2. Normalize the adjacency matrix, use a multi-layer graph convolutional network to aggregate and update the feature matrix, output a high-dimensional embedding vector, and perform contrastive loss training;

[0007] S3. Set the controller’s final state representation and reward function, and obtain a scheduling strategy that adapts to the dynamic network environment through meta-reinforcement learning and actor-critic network training;

[0008] S4. Use the covariance matrix adaptive evolutionary algorithm to perform a global search for the parameters of meta-reinforcement learning in a simulation environment, and find the optimal solution through candidate solution sampling, fitness evaluation, and parameter update;

[0009] S5. Collect the latest status of each controller in real time, update the embedding through the pre-trained GCN, use the trained meta-reinforcement learning model to output the scheduling strategy based on the global state, and use Q-learning to fine-tune and update the network parameters online.

[0010] Optionally, step S1 includes:

[0011] S1-1. For each controller i, regularly collect CPU utilization , memory usage , message queue length , current traffic rate , the number of connections with other controllers , forming the state vector ;

[0012] S1-2, the connection relationship between each controller is represented by the adjacency matrix A;

[0013] S1-3. Build the entire SDN network as a graph model , each node represents a controller, the node set is represented by V, the direct connection relationship between nodes is determined according to the non-zero elements in the adjacency matrix A, and the edge set is represented by E;

[0014] S1-4, the state vectors of all controllers Arrange rows and splice them into feature matrix .

[0015] Optionally, step S2 includes:

[0016] S2-1. Obtain the normalized adjacency matrix ;

[0017] S2-2. Use the graph convolutional network (GCN) to perform multi-layer feature aggregation on nodes. Through the update of each layer, the representation of the node contains its own information and incorporates the features of its neighbors. After L layers of GCN operation, the final embedding vector of each node is represented as ;

[0018] S2-3, using contrast loss function, the embedding generated by GNN reflects the similarity and difference between nodes, and the objective function uses express.

[0019] Optionally, step S3 includes:

[0020] S3-1. For each controller i, obtain the state vector and GNN embedding vector , concatenate the two to form the final state representation of the controller , then the state information set of the entire network is expressed as S={ }, ;

[0021] S3-2, design the action space, discretize the action into two types: new flow allocation and flow migration operations, where a=j means assigning new flow to controller j, a=(i j, ) represents the flow of part of controller i Migrate to controller j;

[0022] S3-3. Load of each controller , calculate the average load as , calculate the load variance , define the load balancing reward =- , set the Shapley value to measure the marginal contribution of each controller, and the Shapley value of each controller i is , calculate the average contribution of all controllers and define the fairness reward , define the final total reward R= + ;

[0023] S3-4. Construct multiple tasks , each task represents a network environment, the initial parameters of the model Is globally shared, for tasks ,from Initially, a small amount of task data is used for gradient descent update, the update results of all tasks are summarized, and the gradients of all tasks are used to update the initial parameters. Make adjustments.

[0024] Optionally, step S4 includes:

[0025] S4-1. Use the CMA-ES algorithm to find the optimal solution and sample candidate solutions from the multivariate normal distribution. ,and Represents the mean of the current search distribution, Indicates the search step size, which is adjusted according to the search degree. The search range is wider, the smaller The search scope is more refined. Where N represents normal distribution, C represents covariance matrix, and for each candidate solution Calculate the comprehensive fitness F( );

[0026] S4-2, according to the fitness evaluation results, the best performing Candidate solutions ( ) is retained and weights are assigned for subsequent updates , the updated mean is used Indicates that the update formula for setting the covariance matrix C is set, and the step size is Adaptively adjust based on the search success rate and the distribution of candidate solutions;

[0027] S4-3, repeat the sampling, fitness calculation and covariance matrix update process until the set number of iterations is reached or the fitness converges, and the optimal candidate solution is finally obtained As the optimal parameter configuration in the meta-reinforcement learning module, it is passed to the online decision-making module.

[0028] Optionally, step S5 includes:

[0029] S5-1. Each controller regularly collects the latest operating indicators and updates the state vector , according to the latest collected node feature matrix and the adjacency matrix , use the pre-trained GCN model to calculate the updated embedding vector of each node , combined with the updated original state and the updated embedding vector , construct the final state representation of controller i , the state collection of all controllers constitutes the global state ;

[0030] S5-2. Use the trained meta-reinforcement learning actor network and input the current global state Then output the action probability distribution and select the optimal action based on the distribution , the selected action Converted into actual scheduling instructions;

[0031] S5-3. Collect new status , calculate the immediate reward according to the reward function , use Q-learning method to update the actor and critic networks;

[0032] S5-4. Regularly aggregate the state, action, and feedback data collected online over a period of time to the overall platform, use the aggregated data to retrain the entire meta-reinforcement learning model, and combine it with the global optimization module (CMA-ES) to re-search the optimal parameters.

[0033] Optionally, the comprehensive fitness function calculation formula is as follows:

[0034]

[0035] in, represents the comprehensive fitness value, represents sampling candidate solutions from a multivariate normal distribution, Indicates the calculation of variance, Indicates the balance weight parameter, which is a positive number pre-set according to the actual situation. Indicates that it is in use As the solution, the Shapley value of the j-th controller is, Indicates that it is in use As the solution, the arithmetic mean of the Shapley values ​​of all N controllers;

[0036] The calculation formula of the covariance matrix C is as follows:

[0037]

[0038] in, Represents a preset learning rate parameter, which is used to control the step size of the covariance matrix update. represents the updated covariance matrix, represents the covariance matrix before updating, Indicates assignment to The weight of a solution, represents the mean calculated based on this generation, represents the total number of candidate solutions generated in each generation, It represents the total number of candidate solutions in each generation, which is ranked in the ith position after sorting by fitness from low to high. Indicates that the candidate with the smallest fitness is selected from all candidate solutions in this generation. solutions, T represents transpose.

[0039] Optionally, the update calculation formula of Q-learning is as follows:

[0040]

[0041] in, represents the online learning rate, represents the discount factor, Indicates the current state, Indicates the current action. Indicates the new state after execution.

[0042] In summary, this application includes at least one of the following beneficial technical effects:

[0043] In this application, a graph neural network is used to obtain a higher-dimensional node representation through information transmission and aggregation between nodes. Meta-reinforcement learning can enable the model to quickly adapt to new tasks. In SDN load balancing, due to the constant changes in network traffic and topology, meta-reinforcement learning is used to quickly update the scheduling strategy when facing sudden traffic or controller state changes. In the process of searching for candidate solutions, the covariance matrix adaptive evolutionary algorithm explores the parameter space through random sampling, uses the covariance matrix to capture the correlation between the parameters, and realizes adaptive search direction and step size adjustment. By continuously updating the mean, covariance and step size of the candidate solutions, it jumps out of the local optimum in the search space and finds the global optimum. The multi-agent collaboration mechanism shares the local state of each controller to the overall platform, constructs a global state diagram, and realizes collaborative decision-making. The real-time feedback update mechanism enables the system to continuously adjust according to the execution effect and continuously search for the optimal solution. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is a flowchart of the implementation of the SDN controller load balancing method based on graph neural network and evolutionary algorithm in this application;

[0045] Figure 2 This is the relevant flow chart of meta-reinforcement learning in this application;

[0046] Figure 3 This is a flowchart related to the CMA-ES optimization algorithm in this application. DETAILED DESCRIPTION

[0047] Embodiments of the present application are described in detail below, examples of which are illustrated in the accompanying drawings.

[0048] Throughout this specification, reference to the terms "certain embodiments," "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with the embodiment or example is included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0049] The present application embodiment discloses a SDN controller load balancing method based on graph neural network and evolutionary algorithm, referring to Figure 1 , including the following steps:

[0050] S1, based on the state vector containing various operating indicators of the controller And the adjacency matrix A containing the connection relationship between each controller, build a graph model and splice to obtain the feature matrix ;

[0051] Specifically, step S1 includes:

[0052] S1-1. Regularly collect various operating indicators of each controller i, including but not limited to CPU utilization , memory usage , message queue length , current traffic rate , the number of connections with other controllers , the state vector is formed by the above operating indicators ;

[0053]

[0054] S1-2, the connection relationship between each controller is represented by the adjacency matrix A;

[0055]

[0056] S1-3. Build the entire SDN network as a graph model , each node represents a controller, the node set is represented by V, the direct connection relationship between nodes is determined according to the non-zero elements in the adjacency matrix A, and the edge set is represented by E;

[0057] S1-4, the state vectors of all controllers Arrange rows and splice them into feature matrix .

[0058] S2. Normalize the adjacency matrix, use a multi-layer graph convolutional network to aggregate and update the feature matrix, output a high-dimensional embedding vector, and perform contrastive loss training;

[0059] Specifically, step S2 includes:

[0060] S2-1. To prevent numerical overflow during the aggregation process, the adjacency matrix is ​​normalized. The normalized adjacency matrix is ​​used express;

[0061] Normalized adjacency matrix The calculation formula is as follows:

[0062]

[0063] Where I represents the identity matrix and D represents the degree matrix;

[0064] The calculation formula of the diagonal elements of the degree matrix D is as follows:

[0065]

[0066] in, represents the diagonal elements of the degree matrix, represents the elements of the adjacency matrix, represents the elements of the identity matrix;

[0067] S2-2. Use the graph convolutional network (GCN) to perform multi-layer feature aggregation on nodes. Through the update of each layer, the representation of the node contains its own information and incorporates the features of its neighbors. After L layers of GCN operation, the final embedding vector of each node is represented as ;

[0068] The update formula for each layer of GCN is as follows:

[0069]

[0070] in, =X, represents the weight matrix of the lth layer, represents the activation function, represents the normalized adjacency matrix;

[0071] The final embedding vector for each node is as follows:

[0072]

[0073] in, represents the final embedding vector of each node, Represents the i-th row of the final feature matrix of GCN;

[0074] S2-3, using the contrast loss function, the embedding generated by GNN can reflect the similarities and differences between nodes, and the objective function uses express;

[0075] Objective function usage The calculation formula is as follows:

[0076]

[0077] Among them, P represents the positive sample pair, N represents the negative sample pair, and 、 、 Represent the final embedding vector of each node respectively;

[0078] S3. Set the controller’s final state representation and reward function, and obtain a scheduling strategy that adapts to the dynamic network environment through meta-reinforcement learning and actor-critic network training;

[0079] Specifically, step S3 includes:

[0080] S3-1. For each controller i, obtain the state vector and GNN embedding vector , concatenate the two to form the final state representation of the controller , then the state information set of the entire network is expressed as S={ }, ;

[0081] in, represents the state vector of controller i, Represents the GNN embedding vector;

[0082] S3-2, design the action space, discretize the action into two types: new flow allocation and flow migration operations, where a=j means assigning new flow to controller j, a=(i j, ) represents the flow of part of controller i Migrate to controller j;

[0083] S3-3. Load of each controller , calculate the average load as , calculate the load variance , define the load balancing reward =- , set the Shapley value to measure the marginal contribution of each controller, and the Shapley value of each controller i is , calculate the average contribution of all controllers and define the fairness reward , define the final total reward R= + ;

[0084] Load Average The calculation formula is as follows:

[0085]

[0086] in, ; Here 、 、 、 、 All use the Z-Score normalization method, using the original value minus the historical average value of the indicator divided by the historical standard deviation to obtain a normalized dimensionless value. 、 、 、 、 The sum of the five values ​​is 1, and these five values ​​are also put into the CMA-ES algorithm. CMA-ES searches for five values ​​in the unconstrained space and converts them into corresponding k values ​​through functions. CMA-ES continuously updates its mean and covariance matrix according to fitness and finds the best value;

[0087] The Shapley value of each controller i The calculation formula is as follows:

[0088]

[0089] Where S represents the controller set, f(S) represents the current subset scheduling performance function, and f(S) = −AvgLatency(S) represents the average response time under the controller subset;

[0090] The fairness reward calculation formula is as follows:

[0091]

[0092] in, represents the global average contribution; represents the Shapley value of controller i, The value is determined by the CMA-ES algorithm, which proposes different After multiple iterations, the optimal balance between network performance and load balancing is found. value;

[0093] S3-4. Construct multiple tasks , each task represents a network environment, the initial parameters of the model Is globally shared, for tasks ,from Initially, a small amount of task data is used for gradient descent update, the update results of all tasks are summarized, and the gradients of all tasks are used to update the initial parameters. Make adjustments;

[0094] S3-5. An actor-critic structure is used to reduce the variance of policy gradient updates. The actor network is responsible for outputting the scheduling strategy. Given a state s, it outputs the action probability distribution. The actor network generates a specific strategy based on the current state, and the critic network is responsible for evaluating the current state and estimating the action value.

[0095] S4. Use the covariance matrix adaptive evolutionary algorithm to perform a global search for the parameters of meta-reinforcement learning in a simulation environment, such as Figure 2 As shown, the optimal solution is found through candidate solution sampling, fitness evaluation and parameter updating;

[0096] Specifically, step S4 includes:

[0097] S4-1, such as Figure 3 As shown, the CMA-ES algorithm is used to find the optimal solution and sample candidate solutions from the multivariate normal distribution. ,and Represents the mean of the current search distribution, Indicates the search step size, which is adjusted according to the search degree. The search range is wider, the smaller The search scope is more refined. Where N represents normal distribution, C represents covariance matrix, and for each candidate solution Calculate the comprehensive fitness F( );

[0098] The calculation formula of the comprehensive fitness function is as follows:

[0099]

[0100] in, represents sampling candidate solutions from a multivariate normal distribution, Indicates the calculation of variance, Indicates the balance weight parameter, which is a positive number pre-set according to the actual situation. Indicates that it is in use As the solution, the Shapley value of the j-th controller is, Indicates that it is in use As the solution, the arithmetic mean of the Shapley values ​​of all N controllers;

[0101] S4-2, according to the fitness evaluation results, the best performing Candidate solutions ( ) is retained and weights are assigned for subsequent updates , the updated mean is used Indicates that the update formula for setting the covariance matrix C is set, and the step size is Adaptively adjust based on the search success rate and the distribution of candidate solutions;

[0102] mean The calculation formula is as follows:

[0103]

[0104] The calculation formula of the covariance matrix C is as follows:

[0105]

[0106] in, Represents a preset learning rate parameter, which is used to control the step size of the covariance matrix update. represents the updated covariance matrix, represents the covariance matrix before updating, Indicates assignment to The weight of a solution, represents the mean calculated based on this generation, represents the total number of candidate solutions generated in each generation, It represents the total number of candidate solutions in each generation, which is ranked in the ith position after sorting by fitness from low to high. Indicates that the candidate with the smallest fitness is selected from all candidate solutions in this generation. solutions, T represents transpose;

[0107] S4-3, repeat the sampling, fitness calculation and covariance matrix update process until the set number of iterations is reached or the fitness converges, and the optimal candidate solution is finally obtained As the optimal parameter configuration in the meta-reinforcement learning module, it is passed to the online decision-making module.

[0108] S5. Collect the latest status of each controller in real time, update the embedding through the pre-trained GCN, use the trained meta-reinforcement learning model to output the scheduling strategy based on the global state, and use Q-learning to fine-tune and update the network parameters online.

[0109] Specifically, step S5 includes:

[0110] S5-1. Each controller regularly collects the latest operating indicators and updates the state vector , according to the latest collected node feature matrix and the adjacency matrix , use the pre-trained GCN model to calculate the updated embedding vector of each node , combined with the updated original state and the updated embedding vector , construct the final state representation of controller i , the state collection of all controllers constitutes the global state ;

[0111] S5-2. Use the trained meta-reinforcement learning actor network and input the current global state Then output the action probability distribution and select the optimal action based on the distribution , the selected action Converted into actual scheduling instructions;

[0112] S5-3. Collect new status , calculate the immediate reward according to the reward function , use Q-learning method to update the actor and critic networks;

[0113] The update calculation formula of Q-learning is as follows:

[0114]

[0115] in, represents the online learning rate, represents the discount factor, Indicates the current state, Indicates the current action. Indicates the new state after execution.

[0116] S5-4. Regularly aggregate the state, action, and feedback data collected online over a period of time to the overall platform, use the aggregated data to retrain the entire meta-reinforcement learning model, and combine it with the global optimization module (CMA-ES) to re-search the optimal parameters.

[0117] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A SDN controller load balancing method based on graph neural network and evolutionary algorithm, characterized in that: The steps include: S1. Based on the state vector containing various operating indicators of the controller and the adjacency matrix containing the connection relationship between each controller, a graphical model is constructed and spliced ​​to obtain a feature matrix; S2. Normalize the adjacency matrix, use a multi-layer graph convolutional network to aggregate and update the feature matrix, output a high-dimensional embedding vector, and perform contrastive loss training; S3. Set the controller’s final state representation and reward function, and obtain a scheduling strategy that adapts to the dynamic network environment through meta-reinforcement learning and actor-critic network training; S4. Use the covariance matrix adaptive evolutionary algorithm to perform a global search for the parameters of meta-reinforcement learning in a simulation environment, and find the optimal solution through candidate solution sampling, fitness evaluation, and parameter update; S5: Collect the latest status of each controller in real time, update the embedding through the pre-trained GCN, use the trained meta-reinforcement learning model to output the scheduling strategy based on the global state, and use Q-learning to fine-tune and update the network parameters online; The step S3 comprises: S3-1. For each controller i, obtain the state vector and GNN embedding vector , concatenate the two to form the final state representation of the controller , then the state information set of the entire network is expressed as S={ }, ; S3-2, design the action space, discretize the action into two types: new flow allocation and flow migration operations, where a=j means assigning new flow to controller j, a=(i j, ) represents the flow of part of controller i Migrate to controller j; S3-3. Load of each controller , calculate the average load as , calculate the load variance , define the load balancing reward =- , set the Shapley value to measure the marginal contribution of each controller, and the Shapley value of each controller i is , calculate the average contribution of all controllers and define the fairness reward , define the final total reward R= + ; S3-4. Construct multiple tasks , each task represents a network environment, the initial parameters of the model Is globally shared, for tasks ,from Initially, a small amount of task data is used for gradient descent update, the update results of all tasks are summarized, and the gradients of all tasks are used to update the initial parameters. Make adjustments; S3-5. An actor-critic structure is used to reduce the variance of policy gradient updates. The actor network is responsible for outputting the scheduling strategy. Given a state s, it outputs the action probability distribution. The actor network generates a specific strategy based on the current state, and the critic network is responsible for evaluating the current state and estimating the action value.

2. The SDN controller load balancing method based on graph neural network and evolutionary algorithm according to claim 1 is characterized in that: The step S1 comprises: S1-1. For each controller i, regularly collect CPU utilization , memory usage , message queue length , current traffic rate , the number of connections with other controllers , forming the state vector ; S1-2, the connection relationship between each controller is represented by the adjacency matrix A; S1-3. Build the entire SDN network as a graph model , each node represents a controller, the node set is represented by V, the direct connection relationship between nodes is determined according to the non-zero elements in the adjacency matrix A, and the edge set is represented by E; S1-4, the state vectors of all controllers Arrange rows and splice them into feature matrix .

3. The SDN controller load balancing method based on graph neural network and evolutionary algorithm according to claim 1 is characterized in that: The step S2 comprises: S2-1. Obtain the normalized adjacency matrix ; S2-2. Use the graph convolutional network (GCN) to perform multi-layer feature aggregation on nodes. Through the update of each layer, the representation of the node contains its own information and incorporates the features of its neighbors. After L layers of GCN operation, the final embedding vector of each node is represented as ; S2-3, using contrast loss function, the embedding generated by GNN reflects the similarity and difference between nodes, and the objective function uses express.

4. The SDN controller load balancing method based on graph neural network and evolutionary algorithm according to claim 1 is characterized in that: The step S4 comprises: S4-1. Use the CMA-ES algorithm to find the optimal solution and sample candidate solutions from the multivariate normal distribution. ,and Represents the mean of the current search distribution, Indicates the search step size, which is adjusted according to the search degree. The search range is , Where N represents normal distribution, C represents covariance matrix, and for each candidate solution Calculate the comprehensive fitness F( ); S4-2, according to the fitness evaluation results, the best performing Candidate solutions ( ) is retained and weights are assigned for subsequent updates , the updated mean is used Indicates that the update formula for setting the covariance matrix C is set, and the step size is Adaptively adjust based on the search success rate and the distribution of candidate solutions; S4-3, repeat the sampling, fitness calculation and covariance matrix update process until the set number of iterations is reached or the fitness converges, and the optimal candidate solution is finally obtained As the optimal parameter configuration in the meta-reinforcement learning module, it is passed to the online decision-making module.

5. The SDN controller load balancing method based on graph neural network and evolutionary algorithm according to claim 1 is characterized in that: The step S5 comprises: S5-1. Each controller regularly collects the latest operating indicators and updates the state vector , according to the latest collected node feature matrix and the adjacency matrix , use the pre-trained GCN model to calculate the updated embedding vector of each node , combined with the updated original state and the updated embedding vector , construct the final state representation of controller i , the state collection of all controllers constitutes the global state ; S5-2. Use the trained meta-reinforcement learning actor network and input the current global state Then output the action probability distribution and select the optimal action based on the distribution , the selected action Converted into actual scheduling instructions; S5-3. Collect new status , calculate the immediate reward according to the reward function , use Q-learning method to update the actor and critic networks; S5-4. Regularly aggregate the state, action, and feedback data collected online over a period of time to the overall platform, use the aggregated data to retrain the entire meta-reinforcement learning model, and combine it with the global optimization module to re-search for the optimal parameters.

6. The SDN controller load balancing method based on graph neural network and evolutionary algorithm according to claim 4 is characterized in that: The calculation formula of the comprehensive fitness function is as follows: ; in, represents the comprehensive fitness value, represents sampling candidate solutions from a multivariate normal distribution, Indicates the calculation of variance, Indicates the balance weight parameter, which is a positive number pre-set according to the actual situation. Indicates that it is in use As the solution, the Shapley value of the j-th controller is, Indicates that it is in use As the solution, the arithmetic mean of the Shapley values ​​of all N controllers; The calculation formula of the covariance matrix C is as follows: ; in, Represents a preset learning rate parameter, which is used to control the step size of the covariance matrix update. represents the updated covariance matrix, represents the covariance matrix before updating, Indicates assignment to The weight of a solution, represents the mean calculated based on this generation, represents the total number of candidate solutions generated in each generation, It represents the total number of candidate solutions in each generation, which is ranked in the ith position after sorting by fitness from low to high. Indicates that the candidate with the smallest fitness is selected from all candidate solutions in this generation. solutions, T represents transpose.

7. The SDN controller load balancing method based on graph neural network and evolutionary algorithm according to claim 5 is characterized in that: The update calculation formula of Q-learning is as follows: ; in, represents the online learning rate, represents the discount factor, Indicates the current state, Indicates the current action. Indicates the new state after execution.

Citation Information

Patent Citations

  • Multi-target distributed deep learning container scheduling method and system and storage medium

    CN120029720A