Distributed test network topology optimization method based on deep reinforcement learning

By optimizing the distributed experimental network topology using the deep reinforcement learning algorithm A2C, the problem of automated configuration of node deployment in large-scale cross-domain simulation experiments was solved, improving network performance and the accuracy of experimental results, while reducing manual intervention and node load pressure.

CN119603161BActive Publication Date: 2025-11-21NO 15 INST OF CHINA ELECTRONICS TECH GRP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411612072.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-13
Publication Date
2025-11-21
Estimated Expiration
2044-11-13

AI Technical Summary

Technical Problem

In large-scale cross-domain distributed simulation experiments, the existing methods of deploying artificial nodes are difficult to optimize the network topology, leading to network latency and data loss, which affects the accuracy and efficiency of the experimental results.

Method used

Deep reinforcement learning (DRL) algorithms, especially the A2C algorithm, are used to automatically configure nodes and optimize the topology of the distributed experimental network. The optimal node deployment scheme is generated through graph neural network and buffer training.

Benefits of technology

It enables automated and rapid node deployment in large-scale cross-domain simulation experiments, reducing manual intervention, improving network performance and the accuracy of test results, and reducing node load pressure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119603161B_ABST
    Figure CN119603161B_ABST
Patent Text Reader

Abstract

The application provides a distributed test network topology optimization method based on deep reinforcement learning, comprising the following steps: constructing a graph data structure; initializing processing; configuring reward calculation selection parameters and initializing action and state selection parameters; sequentially performing a first algorithm and a second algorithm, and outputting an accuracy rate; if the accuracy rate is greater than a preset accuracy rate, the reward calculation selection parameters are represented as the first algorithm, otherwise, the reward calculation selection parameters are represented as the second algorithm; and repeating the above process. The application uses deep reinforcement learning to replace manual work, and optimizes a deployment scheme of a node network of large-scale tests. The method can effectively save manual work, can quickly recalculate according to a field condition, and can optimize a change of a few nodes in large-scale calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of distributed node network technology, and in particular to a method for optimizing the topology of a distributed experimental network based on deep reinforcement learning. Background Technology

[0002] In distributed node networks, network topology significantly impacts key performance metrics, including link utilization, throughput, and latency. In large-scale distributed simulation experiments across geographical regions, network transmission conditions vary considerably between different distributed nodes. Improving the overall performance of the experimental system requires a comprehensive node deployment scheme. In practice, manual optimization is often employed, which cannot guarantee finding the optimal solution. This application proposes a method using Deep Reinforcement Learning (DRL) algorithms to address the cross-domain joint experimental node deployment problem, employing the A2C algorithm to find the optimal solution for the node deployment scheme.

[0003] In industrial simulation technology, using single devices for simulation experiments is no longer sufficient. Currently, simulating a specific system has become an important research direction in computer simulation, applied in fields such as power and aerospace. To conduct complete and thorough simulation experiments on a system, distributed simulation systems are commonly used. These systems build simulation platforms, connecting the devices to be tested to achieve comprehensive and realistic simulation experiments on a system composed of multiple targets. In distributed simulation platform technology, heterogeneity exists due to the different communication methods of the various devices involved in the experiment. Experimenters typically use nodes as relay stations to re-encode and forward information sent by different types of devices to the appropriate receiving devices. In practical applications, the devices participating in the experiment often come from different regions with significant distances, resulting in cross-regional testing. This leads to significant transmission delays between nodes, and the node deployment scheme in the distributed network directly affects the performance of the entire cross-regional joint testing system. To minimize transmission delays and reduce the load on individual nodes, it is necessary to optimize the network topology and adopt a reasonable node deployment scheme based on data transmission requirements.

[0004] In current technological applications, the selection of node deployment in distributed experimental networks is done manually by experimenters using prior knowledge. For small-scale experiments in a single region, the low latency and low concurrency characteristics make the impact of the experimental network environment on the results negligible.

[0005] As the scale of experiments increases and the number of participating devices grows, the number of nodes required for the experiment and the data interaction relationships between nodes expand rapidly, making the distributed network highly complex. This leads to two problems. First, since the experimental facilities are located in different regions, unavoidable network latency will occur between links. If the accumulated latency is too high, it will affect the accuracy of the experimental results. Second, with the expansion of the number of nodes, individual nodes are prone to high data concurrency. Excessive data traffic exceeding the forwarding capacity can lead to data loss or even node crashes. Therefore, the network structure needs to be optimized in large-scale distributed simulation experiments.

[0006] In small-scale experiments, manually configuring nodes can meet the requirements, but in large-scale cross-domain simulation experiments, it is very difficult to optimize the complex and large-scale distributed network topology. Summary of the Invention

[0007] The technical problem to be solved by this invention is how to accurately achieve automated node configuration to meet the needs of large-scale cross-domain simulation experiments; in view of this, this invention provides a distributed experimental network topology optimization method based on deep reinforcement learning.

[0008] The technical solution adopted in this invention is a method for topology optimization of distributed experimental networks based on deep reinforcement learning, comprising:

[0009] Step S1: Construct a graph data structure, including the set of nodes in the distributed network and the set of communication relationships between the nodes;

[0010] Step S2: Initialize the graph neural network and the buffer, wherein the graph neural network is used for quantitative evaluation;

[0011] Step S3, initialize constants;

[0012] Step S4: Configure reward calculation selection parameters and initialize action and state selection parameters, wherein the value of the reward calculation selection parameter represents the source of the parameter acquisition, and the value of the action and state selection parameters represents whether the current state is a full space or a compressed space.

[0013] Step S5: Based on the reward calculation selection parameters and the action and state selection parameters, execute the first algorithm for a preset number of steps to determine the policy gradient and adjust the policy function parameters. If the reward calculation selection parameters indicate that the source of the parameters is the verifier of the first algorithm, then record the corresponding data in the buffer.

[0014] Step S6: Execute the second algorithm, input the data and threshold of the buffer into the preset graph neural network, and output the updated graph neural network;

[0015] Step S7: Use the configured policy function to generate action sequences and related network structures, execute Verifier to check the effectiveness of the current graph neural network, calculate the corresponding target value, and add the target value to the buffer;

[0016] Step S8: Execute the operation of the second algorithm, input the trained network of the second algorithm and the generated target value data, and output the accuracy. If the accuracy is greater than the preset value, the reward calculation selection parameter is represented by the first algorithm; otherwise, it is the second algorithm. Repeat steps S5 to S8.

[0017] In one implementation, the constants in step S3 include: preset number of steps, number of epochs, and threshold.

[0018] In one implementation, the first algorithm is A2C and the second algorithm is GNN.

[0019] In one implementation, the accuracy preset value is specifically configured to be at least 90%.

[0020] In one implementation, the training process in the network of the second algorithm, which is input to the network, includes:

[0021] Randomize the buffer, collect data pairs and represent them as icon data;

[0022] Add labels to the chart data; the labels are used to indicate the current network status.

[0023] The optimizer is used to update the input GNN layer parameters for each pair of data in the graph data;

[0024] Repeat the above steps for a total of the specified number of epochs.

[0025] Another aspect of the present invention provides an electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the distributed experimental network topology optimization method based on deep reinforcement learning as described in any of the preceding claims.

[0026] Another aspect of the present invention provides a computer storage medium storing a computer program that, when executed by a processor, implements the steps of the distributed experimental network topology optimization method based on deep reinforcement learning as described in any of the preceding claims.

[0027] Compared with the prior art, the present invention has at least the following advantages:

[0028] This invention provides a scheme for optimizing the deployment of node networks in large-scale experiments using deep reinforcement learning instead of manual labor. This method effectively saves manpower, can quickly recalculate based on field conditions, and can optimize the handling of changes in the situation of a few nodes in large-scale computations. Attached Figure Description

[0029] Figure 1 This is a schematic diagram of the topology optimization method for distributed experimental networks based on deep reinforcement learning according to an embodiment of the present invention.

[0030] Figure 2 A schematic diagram illustrating the implementation logic of the distributed experimental network topology optimization method based on deep reinforcement learning according to an embodiment of the present invention;

[0031] Figure 3 This is a schematic diagram of the electronic device according to an embodiment of the present invention. Detailed Implementation

[0032] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the present invention will be described in detail below with reference to the accompanying drawings and preferred embodiments.

[0033] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms (e.g., those defined in common dictionaries) shall be interpreted as having the meaning consistent with their meaning in the context of the relevant art and shall not be interpreted in an idealized or overly formal sense unless expressly so specified herein.

[0034] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0035] This invention provides a method for optimizing the topology of a distributed experimental network based on deep reinforcement learning, comprising:

[0036] Step S1: Construct a graph data structure, including the set of nodes in the distributed network and the set of communication relationships between the nodes;

[0037] Step S2: Initialize the graph neural network and the buffer, wherein the graph neural network is used for quantitative evaluation;

[0038] Step S3, initialize constants;

[0039] Step S4: Configure reward calculation selection parameters and initialize action and state selection parameters, wherein the value of the reward calculation selection parameter represents the source of the parameter acquisition, and the value of the action and state selection parameters represents whether the current state is a full space or a compressed space.

[0040] Step S5: Based on the reward calculation selection parameters and the action and state selection parameters, execute the first algorithm for a preset number of steps to determine the policy gradient and adjust the policy function parameters. If the reward calculation selection parameters indicate that the source of the parameters is the verifier of the first algorithm, then record the corresponding data in the buffer.

[0041] Step S6: Execute the second algorithm, input the data and threshold of the buffer into the preset graph neural network, and output the updated graph neural network;

[0042] Step S7: Use the configured policy function to generate action sequences and related network structures, execute Verifier to check the effectiveness of the current graph neural network, calculate the corresponding target value, and add the target value to the buffer;

[0043] Step S8: Execute the operation of the second algorithm, input the trained network of the second algorithm and the generated target value data, and output the accuracy. If the accuracy is greater than the preset value, the reward calculation selection parameter is represented by the first algorithm; otherwise, it is the second algorithm. Repeat steps S5 to S8.

[0044] In step S3, the constants include: preset number of steps, number of epochs, and threshold.

[0045] In this embodiment, the first algorithm is A2C, and the second algorithm is GNN.

[0046] In this embodiment, the accuracy preset value is specifically configured to be at least 90%.

[0047] In this embodiment, the training process, input into the trained network of the second algorithm, includes:

[0048] Randomize the buffer, collect data pairs and represent them as icon data;

[0049] Add labels to the chart data; the labels are used to indicate the current network status.

[0050] The optimizer is used to update the input GNN layer parameters for each pair of data in the graph data;

[0051] Repeat the above steps for a total of the specified number of epochs.

[0052] refer to Figure 1 as well as Figure 2The method provided in this implementation will be further explained below.

[0053] This application proposes a solution for optimizing distributed simulation experiment node deployment schemes using Deep Reinforcement Learning (DRL). Among numerous DRL algorithms, we chose the A2C algorithm to address this type of problem for the following reasons: First, the A2C algorithm has low variance. In network structure optimization problems, the optimization difficulty varies significantly depending on the structure. Compared to other DRL algorithms, A2C has a more stable training process and is better able to adapt to larger variances between samples. Second, the A2C algorithm is an online learning algorithm, which can update the policy at each time step without waiting for the entire round to end, thus exhibiting better real-time performance. In distributed simulation experiments, algorithms with better real-time performance are beneficial for quickly resolving various unexpected problems. Third, the A2C algorithm has fewer hyperparameters, making parameter tuning easier and allowing experimenters to quickly adjust it on-site according to actual conditions. A2C (Advantage Actor-Critic) is an algorithm for deep reinforcement learning.

[0054] It combines the Actor-Critic approach with the concept of Advantage learning. In A2C, there are two main components: the actor and the critic. The actor is responsible for selecting actions, while the critic evaluates the quality of those actions.

[0055] The A2C algorithm uses neural networks to approximate the actor and critic, and learns by optimizing these two networks. The advantage of A2C lies in its parallelizability, as the actor and critic can be updated independently during training. This makes A2C computationally more efficient, as it can collect data and perform updates in parallel across multiple environments. Furthermore, A2C leverages the concept of advantage learning; by calculating the advantage of each action, it can more accurately evaluate the value of actions, thus improving learning efficiency. In summary, the A2C algorithm is widely used in deep reinforcement learning because it combines the advantages of the Actor-Critic approach and exhibits good performance in both efficiency and execution.

[0056] The node deployment scheme problem in cross-regional distributed simulation experiments is essentially a graph search problem. The network connections between the nodes can be summarized into an undirected fully connected graph, and determining the node deployment scheme involves selecting nodes and edges from this graph to form a subgraph.

[0057] To solve this type of problem, we first need to abstract the network topology, representing it as a graph G = (V, E), where the set of nodes V = {1, 2, ..., N} and the set of edges E = {(m, n), m, n V}. We define the adjacency matrix x of graph G as a connectivity variable, i.e., if i and j are connected, then (x)ij = 1; otherwise, (x)ij = 0.

[0058] The following metrics are used to measure network topology performance:

[0059] Efficiency(x)=αPressure(x)+βCost(x,x_0)+Delay(x) (1)

[0060] Distance(e)≤ (2)

[0061] Load(e)≤ (3)

[0062] M(x) = True (4)

[0063] In (1), x represents the optimized network structure, Efficiency(x) represents the efficiency index of the distributed network, and Pressure(x) represents the maximum data transmission volume of a single link under the topology, used to describe the pressure situation of high-concurrency nodes. x0 represents the network structure before optimization, Cost(x,x_0) is the number of nodes that need to be moved when the graph is transformed from the initial topology x0 to x, which is the cost of the optimization process, Delay(x) refers to the average delay between nodes in the network, and α and β are the weights of the two when evaluating the network operating efficiency.

[0064] The network structure needs to satisfy the constraints shown in formulas (2), (3), and (4). In constraint (2), Distance(e) represents the distance between the two end nodes of a given edge e in the network topology. D represents the maximum acceptable physical distance between two nodes under the condition that the physical delay does not interfere with the accuracy of the experiment. For the optimized new graph, for each link (i; j), the two end nodes vi and vj must be within the range of the connection distance D. In constraint (3), Load(e) represents the load of the link e between nodes. This is the maximum allowable load. In actual experiments, due to the high concurrency of the test data, it is necessary to limit the link load between the two nodes to prevent data congestion. Condition (4) is an abstract feasibility requirement of the topology, which is judged according to the specific requirements during the experiment, such as the limit on the number of times data is forwarded.

[0065] The algorithm flow used in this embodiment is as follows: Figure 1As shown, the algorithm consists of three main parts: a representation layer, a DRL agent, and a validator. In the figure, s represents the state, a represents the action, and r represents the reward function. T(s,a) represents the transition function. π is the target value calculated by the validator. The A2C algorithm uses the representation layer to compress the state and action into s0 and a0, and uses a GNN classifier ~f to learn the true target function f.

[0066] A specific implementation process is as follows:

[0067] Construct a graph data structure to represent the overall structure of the distributed network. The content includes the set of nodes V of the distributed network and the set of communication relationships between the nodes E. The two are represented and stored using the vertices and edges in the data structure, respectively.

[0068] Initialize the graph neural network with appropriate parameters, construct an evaluation function using the policy function parameter theta to quantify the optimization scheme, input the current state of the network and the operation performed, output the quantified evaluation value of the operation, and initialize the buffer R.

[0069] Initialize constants such as time steps T, number of epochs N, and threshold Q.

[0070] Set the reward calculation selection parameter C_r=1. When C_r = 1, it means the reward parameter c is obtained from the Verifier in the A2V algorithm. When C_r = 2, it means it is obtained from the GNN.

[0071] Initialize the action and state selection C_as. When C_as = 1, it represents full space. When C_as = 2, it represents compressed space.

[0072] Repeat the following steps:

[0073] Based on parameters C_r and C_as, execute the A2C algorithm for T steps, calculate the policy gradient, and adjust the policy function parameter theta. If C_r = 1, then record the data {a^t, s^t, r^t}^T _{t=1} into the buffer R.

[0074] Perform GNN training operations. Input the data in buffer R and the threshold Q into the existing graph neural network f(t), and then output a new graph neural network f(t+1).

[0075] The action sequence {a_1, a_2, …, a_n} and the associated network structure {s_1, s_2, …, s_n} are generated using a policy function. Then, a verifier is executed to check the validity of the graph, the corresponding target values ​​{r_1, r_2, …, r_n} are calculated, and the data is added to the buffer R.

[0076] Perform the GNN training operation. Input the trained GNN neural network and the data generated in the previous step. Then output the accuracy; if it is greater than 90%, C_r = 1; otherwise, C_r = 2.

[0077] The GNN training process in sections 6.2 and 6.4 is as follows:

[0078] 1. A randomized buffer R is used to collect data pairs n, denoted as D = {(s_i, r_i)}^n _{i=1}

[0079] 2. Add labels to the chart data, using 1 and 0 to represent good and bad network conditions.

[0080] 3. Update the input GNN layer parameters for each pair of data (s_i, r_i) using the optimizer.

[0081] 4. Repeat the above steps N times.

[0082] Compared with the prior art, this embodiment has at least the following advantages:

[0083] This invention proposes a method for local network optimization based on deep reinforcement learning to avoid errors in large-scale distributed simulation experiments due to network performance issues in poor network environments.

[0084] This invention proposes a scheme to optimize the deployment of node networks in large-scale experiments by replacing manual labor with deep reinforcement learning. This method effectively saves manpower, can quickly recalculate according to field conditions, and can optimize the handling of changes in the situation of a few nodes in large-scale computations.

[0085] A second embodiment of the present invention provides an electronic device, such as... Figure 3 As shown, it can be understood as a physical device, including a processor and a memory storing processor-executable instructions. When the instructions are executed by the processor, the following operations are performed:

[0086] Step S1: Construct a graph data structure, including the set of nodes in the distributed network and the set of communication relationships between the nodes;

[0087] Step S2: Initialize the graph neural network and the buffer, wherein the graph neural network is used for quantitative evaluation;

[0088] Step S3, initialize constants;

[0089] Step S4: Configure reward calculation selection parameters and initialize action and state selection parameters, wherein the value of the reward calculation selection parameter represents the source of the parameter acquisition, and the value of the action and state selection parameters represents whether the current state is a full space or a compressed space.

[0090] Step S5: Based on the reward calculation selection parameters and the action and state selection parameters, execute the first algorithm for a preset number of steps to determine the policy gradient and adjust the policy function parameters. If the reward calculation selection parameters indicate that the source of the parameters is the verifier of the first algorithm, then record the corresponding data in the buffer.

[0091] Step S6: Execute the second algorithm, input the data and threshold of the buffer into the preset graph neural network, and output the updated graph neural network;

[0092] Step S7: Use the configured policy function to generate action sequences and related network structures, execute Verifier to check the effectiveness of the current graph neural network, calculate the corresponding target value, and add the target value to the buffer;

[0093] Step S8: Execute the operation of the second algorithm, input the trained network of the second algorithm and the generated target value data, and output the accuracy. If the accuracy is greater than the preset value, the reward calculation selection parameter is represented by the first algorithm; otherwise, it is the second algorithm. Repeat steps S5 to S8.

[0094] In the third embodiment of the present invention, the process of the distributed experimental network topology optimization method based on deep reinforcement learning is the same as that in the first and second embodiments. The difference lies in the engineering implementation: this embodiment can be implemented using software plus necessary general-purpose hardware platforms. While hardware implementation is also possible, the former is often a better approach. Based on this understanding, the method of the present invention can be embodied in the form of a computer software product stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including several instructions to cause a device to execute the method described in the embodiments of the present invention.

[0095] Through the description of specific embodiments, a more in-depth and specific understanding should be gained of the technical means and effects adopted by the present invention to achieve the intended purpose. However, the accompanying drawings are only provided for reference and illustration and are not intended to limit the present invention.

Claims

1. A method for topology optimization of distributed experimental networks based on deep reinforcement learning, characterized in that, include: Step S1: Construct a graph data structure, including the set of nodes in the distributed network and the set of communication relationships between the nodes; Step S2: Initialize the graph neural network and the buffer, wherein the graph neural network is used for quantitative evaluation; Step S3, initialize constants; Step S4: Configure reward calculation selection parameters and initialize action and state selection parameters, wherein the value of the reward calculation selection parameter represents the source of the parameter acquisition, and the value of the action and state selection parameters represents whether the current state is a full space or a compressed space. Step S5: Based on the reward calculation selection parameters and the action and state selection parameters, execute the first algorithm for a preset number of steps to determine the policy gradient and adjust the policy function parameters. If the reward calculation selection parameters indicate that the source of the parameters is the verifier of the first algorithm, then record the corresponding data in the buffer. Step S6: Execute the second algorithm, input the data and threshold of the buffer into the preset graph neural network, and output the updated graph neural network; Step S7: Use the configured policy function to generate action sequences and related network structures, execute Verifier to check the effectiveness of the current graph neural network, calculate the corresponding target value, and add the target value to the buffer; Step S8: Execute the operation of the second algorithm, input the trained network of the second algorithm and the generated target value data, and output the accuracy. If the accuracy is greater than the preset value, the reward calculation selection parameter is represented by the first algorithm; otherwise, it is the second algorithm. Repeat steps S5 to S8.

2. The method for topology optimization of distributed experimental networks based on deep reinforcement learning according to claim 1, characterized in that, In step S3, the constants include: preset number of steps, number of epochs, and threshold.

3. The method for topology optimization of distributed experimental networks based on deep reinforcement learning according to claim 2, characterized in that, The first algorithm is A2C, and the second algorithm is GNN.

4. The method for topology optimization of distributed experimental networks based on deep reinforcement learning according to claim 1, characterized in that, The accuracy preset value is specifically configured to be at least 90%.

5. The method for topology optimization of distributed experimental networks based on deep reinforcement learning according to claim 3, characterized in that, The training process for the network of the second algorithm, which is input into the network, includes: Randomize the buffer, collect data pairs, and represent them as chart data; Add labels to the chart data; the labels are used to indicate the current network status. The optimizer is used to update the input GNN layer parameters for each pair of data in the graph data; Repeat the above steps for a total of the specified number of epochs.

6. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the distributed experimental network topology optimization method based on deep reinforcement learning as described in any one of claims 1 to 5.

7. A computer storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the distributed experimental network topology optimization method based on deep reinforcement learning as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Traffic light control method based on deep reinforcement learning and inverse reinforcement learning

    CN115762199A

  • Deep reinforcement learning SDN intelligent routing optimization method based on graph neural network

    CN116938810A