OTN private network operation process optimization method based on deterministic strategy gradient reinforcement learning algorithm

By building a digital twin model of OTN private network and using DDPG reinforcement learning algorithm, the operation process of OTN private network is optimized, and the problem of insufficient automation and intelligence in the existing technology is solved, and more efficient network management and maintenance is achieved.

CN120238781APending Publication Date: 2025-07-01CHINA YANGTZE POWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510211688.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The optimization of the existing OTN private network operation process depends on expert experience, and the degree of automation and intelligence is insufficient, resulting in inefficient network management and maintenance.

Method used

The digital twin model of OTN private network is built, and reinforcement learning method based on deep deterministic strategy gradient algorithm (DDPG) is adopted to optimize the operation process of OTN private network through state space, action space and reward functions, including bandwidth adjustment, traffic routing changes, device switching and topology optimization.

Benefits of technology

It has improved the intelligent operation and maintenance level of OTN private network, improved network transmission efficiency, reduced latency and optimized resource configuration, and achieved more efficient network management and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238781A_ABST
    Figure CN120238781A_ABST
Patent Text Reader

Abstract

The invention discloses an OTN private network operation process optimization method based on a deterministic strategy gradient reinforcement learning algorithm. The method comprises the following steps: 1) constructing a digital twinborn model of an OTN private network operation process; 2) adopting a reinforcement learning method based on a depth deterministic policy gradient (DDPG) algorithm to realize optimal configuration of the OTN private network operation process; and constructing an environment model of reinforcement learning, and realizing OTN private network operation process optimization configuration based on the deep deterministic strategy gradient algorithm reinforcement learning method based on a state space, an action space and a reward function of a reinforcement learning algorithm constructed on the basis of the environment model. On the basis of the digital twinborn model, the OTN private network operation process optimization method based on reinforcement learning is provided, and the intelligent operation and maintenance degree of the OTN private network can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to smart grid technologies, and particularly to an OTN private network operation process optimization method based on a deterministic policy gradient reinforcement learning algorithm. Background Art

[0002] An optical transport network (OTN) refers to a transport network that realizes the transmission, multiplexing, routing selection, and monitoring of service signals in the optical domain and ensures its performance indicators and survivability. OTN technology is a compromise between electrical networks and all-optical networks, transplanting the powerful and perfect monitoring and management functions of the synchronous digital hierarchy (SDH) into the wavelength division multiplexing (WDM) optical network, effectively making up for the deficiencies of existing WDM systems in performance monitoring and maintenance management. OTN provides a bandwidth much higher than that of mobile communication networks, and is often used to provide data transmission capabilities of 10 Gbps or even higher, and is suitable for the interconnection of data centers, large-capacity video transmission, and the connection of internal networks of large enterprises. Since the transmission speed of optical signals is much faster than that of electrical signals, this makes OTN particularly suitable for those communication applications with strict real-time requirements, such as financial transactions and telemedicine. OTN uses wavelength division multiplexing technology to simultaneously transmit multiple independent data streams in the same optical fiber, enabling it to carry multiple services (such as voice, video, data), and realizing efficient resource utilization. Since the transmission distance of optical fibers is much longer than that of copper cables and optical signals are less susceptible to interference than electrical signals, OTN private networks are more popular in large enterprises with geographically dispersed locations and application scenarios that require the protection of data confidentiality. Through these characteristics, OTN private networks have become an ideal choice for enterprises, government agencies, and large organizations that require high data transmission quality and stability, not just limited to the mobile communication field.

[0003] The existing optimization of the OTN private network operation process includes network design optimization, optical fiber management and maintenance, optical forwarding unit tuning, monitoring and data analysis, and maintenance and update. These optimization methods rely on expert experience, with poor automation and intelligence levels, mainly because of the insufficient digitalization of the OTN private network operation. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an OTN private network operation process optimization method based on a deterministic policy gradient reinforcement learning algorithm for the defects in the prior art.

[0005] The technical solution adopted by the present invention to solve its technical problems is: an OTN private network operation process optimization method based on a deterministic policy gradient reinforcement learning algorithm, including the following steps:

[0006] 1) Build a digital twin model for the operation process of the OTN private network;

[0007] 1.1) Collect data on the operation process of the OTN private network, including:

[0008] Network traffic data, including: packet rate, latency, packet loss rate, bandwidth utilization;

[0009] Node status data, including: temperature, optical power;

[0010] Device operation load data, including: CPU usage, memory usage, throughput;

[0011] 1.2) Data preprocessing; clean and standardize the collected raw data;

[0012] Data standardization for bandwidth, latency, jitter, and packet loss rate

[0013] 1.3) Build a digital twin model for the OTN private network, including a bandwidth model, a latency model, a jitter model, and a packet loss rate model;

[0014] Among them, the specific construction methods of the bandwidth model, latency model, jitter model, and packet loss rate model are shown in the following formula:

[0015] B link = min(B link1 , B link2 ,..., B linkN )

[0016]

[0017] J path = max(Δt i - min(Δt i ))

[0018]

[0019] Among them, B link is the actual bandwidth of the link, and the minimum bandwidth among all links is selected as the bandwidth of the current link; D path is the total latency from the source node to the destination node, D linki is the latency of the i-th link, and Process_Delay is the processing latency of the network device; J path is the jitter of the network path, and Δt i is the change in the time interval between packets; PLR is the packet loss rate, N is the number of lost packets, and T is the total number of packets sent;

[0020] 2) Adopt a reinforcement learning method based on the Deep Deterministic Policy Gradient algorithm (DDPG) to achieve optimal configuration during the operation of the OTN private network;

[0021] Construct an environment model for reinforcement learning, and based on this, construct the state space, action space, and reward function of the reinforcement learning algorithm to achieve optimal configuration during the operation of the OTN private network based on the reinforcement learning method of the Deep Deterministic Policy Gradient algorithm;

[0022] 2.1) Construct the state space of the OTN private network digital twin model; the state space should include all possible states that the agent can observe in the environment, and the state space includes: Link traffic status: the packet rate of each link, the bandwidth utilization rate of each link, the delay of each link, the packet loss rate of each link; Node status information: the temperature of each node, the optical power of each node, the optical loss of each node; Device performance metrics: the CPU usage rate of each node, the memory usage of each node, the throughput of each node;

[0023] Integrate this information into a state vector to represent all the information of the network at the current time step;

[0024] 2.2) Construct the action space of the reinforcement learning algorithm so that the agent has the ability to make corresponding behaviors according to the state space parameters;

[0025] The action space defines the executable operations, including the following operations: Bandwidth adjustment: increase or decrease the bandwidth of the link; Traffic routing change: adjust the transmission path of the traffic in the network; Device switch: control the on / off state of the device to save power consumption; Topology optimization: modify the network topology: enable or disable the standby link;

[0026] The design of the action space is a continuous vector, where each element represents a specific operation, as shown in the following formula:

[0027] action = [bandwidth_change_1, bandwidth_change_2, route_change_1,..., device_on / off_n]

[0028] 2.3) Construct the reward function of the reinforcement learning algorithm;

[0029] The reward function is designed as:

[0030] reward = α * (network_utilization - target_utilization) - β * (power_consumption - target_power_consumption)

[0031] Where: network_utilization represents the current utilization rate of the network; target_utilization represents the target network utilization rate; power_consumption represents the total power consumption of the network; target_power_consumption represents the target power consumption level; α and β are weight coefficients used to balance utilization optimization and power consumption control.

[0032] 2.4) DDPG reinforcement learning network training; Initialize the policy network and value network, and conduct the learning of the agent;

[0033] During the learning process, the agent obtains environmental information through the digital twin model of the OTN private network. The policy network will select the corresponding action a based on the current environmental state and the value evaluation of the value network, and then update the state s, reward value reward, and determine whether to end the task of the agent at the next moment according to the current action a. At the same time, the three are stored in the data buffer for the learning of the agent;

[0034] When each round ends, perform discount decay on the reward value, extract the data in the buffer and input it into the policy optimization algorithm for update, including calculating the policy loss, calculating the value loss, and updating the network parameters according to the loss; Use the state s, action a, and reward r to update the network parameter weights in each update iteration, and adjust the hyperparameters of the DDPG algorithm;

[0035] 2.5) Use a network simulation tool for preliminary testing to verify the effectiveness of the DDPG algorithm model;

[0036] 2.6) Deploy the optimized model into the actual system. The agent selects the optimal action according to the current state to achieve the optimization of the operation process of the OTN private network.

[0037] The beneficial effects produced by the present invention are:

[0038] Based on the digital twin model, the present invention proposes an optimization method for the operation process of the OTN private network based on reinforcement learning, which can effectively improve the intelligent operation and maintenance level of the OTN private network. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The following will further illustrate the present invention in conjunction with the drawings. In the drawings:

[0040] Figure 1 is the flowchart of the method of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] To make the objectives, technical solutions, and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0042] As Figure 1 shown, an OTN private network operation process optimization method based on a deterministic policy gradient reinforcement learning algorithm includes the following steps:

[0043] 1) Construct a digital twin model of the OTN private network operation process;

[0044] 1.1) Collect data on the OTN private network operation process, including:

[0045] Network traffic data, including: packet rate, latency, packet loss rate, bandwidth utilization;

[0046] Node status data, including: temperature, optical power;

[0047] Device operation load data, including: CPU usage, memory usage, throughput;

[0048] 1.2) Data preprocessing; clean and normalize the collected raw data;

[0049] Include data standardization of bandwidth, latency, jitter, and packet loss rate

[0050] 1.3) Construct a digital twin model of the OTN private network, including a bandwidth model, a latency model, a jitter model, and a packet loss rate model;

[0051] Among them, the specific construction methods of the bandwidth model, the latency model, the jitter model, and the packet loss rate model are shown in the following formulas:

[0052] B link = min(B link1 , B link2 ,..., B linkN )

[0053]

[0054] J path = max(Δt i - min(Δt i ))

[0055]

[0056] Among them, B link is the actual bandwidth of the link, and the minimum bandwidth among all links is selected as the bandwidth of the current link; D path is the total latency from the source node to the destination node, Dlinki is the delay of the i-th link, and Process_Delay is the processing delay of the network device; J path is the jitter of the network path, and Δt i is the change in the time interval between data packets; PLR is the packet loss rate, N is the number of lost packets, and T is the total number of data packets sent;

[0057] 2) Adopt a reinforcement learning method based on the Deep Deterministic Policy Gradient algorithm DDPG to optimize the configuration during the operation of the OTN private network;

[0058] Construct an environmental model for reinforcement learning, and based on this, construct the state space, action space, and reward function of the reinforcement learning algorithm to optimize the configuration during the operation of the OTN private network based on the reinforcement learning method of the Deep Deterministic Policy Gradient algorithm;

[0059] 2.1) Construct the state space of the OTN private network digital twin model; The state space should include all possible states that the agent can observe in the environment, and the state space includes: Link traffic status: The data packet rate of each link, the bandwidth utilization rate of each link, the delay of each link, and the packet loss rate of each link; Node status information: The temperature of each node, the optical power of each node, and the optical loss of each node; Device performance metrics: The CPU usage rate of each node, the memory usage of each node, and the throughput of each node;

[0060] Integrate this information into a state vector to represent all the information of the network at the current time step;

[0061] 2.2) Construct the action space of the reinforcement learning algorithm to enable the agent to have the ability to make corresponding behaviors according to the state space parameters;

[0062] The action space defines the executable operations, including the following operations: Bandwidth adjustment: Increase or decrease the bandwidth of the link; Traffic routing change: Adjust the transmission path of the traffic in the network; Device on / off: Control the on / off state of the device to save power; Topology optimization: Modify the network topology: Enable or disable standby links;

[0063] The design of the action space is a continuous vector, where each element represents a specific operation, as shown in the following formula:

[0064] action = [bandwidth_change_1, bandwidth_change_2, route_change_1,..., device_on / off_n]

[0065] 2.3) Construct the reward function of the reinforcement learning algorithm;

[0066] The reward function is designed as follows:

[0067] reward = α * (network_utilization - target_utilization) - β * (power_consumption - target_power_consumption)

[0068] where: network_utilization represents the current utilization rate of the network; target_utilization represents the target network utilization rate; power_consumption represents the total power consumption of the network; target_power_consumption represents the target power consumption level; α and β are weight coefficients used to balance utilization optimization and power consumption control.

[0069] 2.4) Training of the DDPG reinforcement learning network; initializing the policy network and the value network, and performing the learning of the agent;

[0070] During the learning process, the agent obtains environmental information through the digital twin model of the OTN private network. The policy network will select the corresponding action a according to the current environmental state and the value evaluation of the value network, and then update the state s, the reward value reward, and determine whether to end the task at the next moment of the agent according to the current action a. At the same time, the three are stored in the data buffer for the learning of the agent;

[0071] When each round ends, the reward value is discounted and decayed, and the data in the buffer is extracted and input into the policy optimization algorithm for update, including calculating the policy loss, calculating the value loss, and updating the network parameters according to the loss; in each update iteration, the network parameter weights are updated using the state s, the action a, and the reward r, and the hyperparameters of the DDPG algorithm are adjusted;

[0072] 2.5) Conduct preliminary tests using a network simulation tool to verify the effectiveness of the DDPG algorithm model;

[0073] 2.6) Deploy the optimized model into the actual system. The agent selects the optimal action according to the current state to achieve the optimization of the OTN private network operation process.

[0074] The present invention relates to an OTN private network operation process optimization method based on reinforcement learning, which forms a comprehensive digital twin model of the OTN private network. By deeply analyzing the operation data of the OTN private network, the optimization objectives are clarified. The data includes the network traffic load, link status, device performance indicators, etc., and the objectives are to improve the network transmission efficiency, reduce latency, and optimize resource allocation. Then, a DDPG environment suitable for the OTN private network is designed. This environment simulates the operation state of the network and provides the infrastructure for training and evaluating the policy. The environment includes aspects such as the network state representation, action space definition, and reward mechanism design, ensuring that the DDPG algorithm can effectively learn and optimize in this environment. Subsequently, the DDPG algorithm is proposed, and a deep learning method is used to construct an actor-critic network architecture for policy optimization. Through the training of the DDPG algorithm, it can continuously explore and improve the network operation policy, optimizing the overall performance of the OTN private network. The optimized policy obtained from the training is applied to the actual operation of the OTN private network, and further optimization adjustments are made according to the feedback data.

[0075] It is expected that the present invention can effectively improve the operation effect of the OTN private network and provide advanced technical support for network management and maintenance. The present invention can realize the optimized configuration of the OTN private network operation process under different states.

[0076] It should be understood that those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations shall fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for optimizing the operation process of an OTN private network based on a deterministic policy gradient reinforcement learning algorithm, characterized in that: The following steps are involved: 1) Build a digital twin model of the OTN private network operation process; 2) Adopt the reinforcement learning method based on the deep deterministic policy gradient algorithm DDPG to achieve the optimal configuration of the OTN private network operation process; Construct an environment model for reinforcement learning, and based on it, construct the state space, action space, and reward function of the reinforcement learning algorithm to achieve the optimal configuration of the OTN private network operation process based on the deep deterministic policy gradient algorithm reinforcement learning method; 2.1) Construct the state space of the OTN private network digital twin model; The state space should contain all possible states that the agent can observe in the environment. The state space includes: link traffic state: packet rate of each link, bandwidth utilization of each link, delay of each link, packet loss rate of each link; node state information: temperature of each node, optical power of each node, optical loss of each node; device performance indicators: CPU utilization of each node, memory usage of each node, throughput of each node; Integrate this information into a state vector, which represents all the information of the network at the current time step; 2.2) Construct the action space of the reinforcement learning algorithm so that the agent has the ability to make corresponding behaviors according to the state space parameters; The action space defines the executable operations, including the following operations: bandwidth adjustment: increase or decrease the bandwidth of the link; traffic routing change: adjust the transmission path of traffic in the network; device switch: control the switch state of the device to save power consumption; topology optimization: modify the network topology structure: enable or disable backup links; 2.3) Construct the reward function of the reinforcement learning algorithm; The reward function is designed as: reward=α*(network_utilization-target_utilization)-β*(power_consumption-target_power_consumption) Where: network_utilization represents the current utilization of the network; target_utilization represents the target network utilization; power_consumption represents the total power consumption of the network; target_power_consumption represents the target power consumption level; α and β are weight coefficients used to balance utilization optimization and power consumption control. 2.4) DDPG reinforcement learning network training; initialization of policy network and value network, and learning of intelligent agent; During the learning process, the agent obtains environmental information through the digital twin model of the OTN private network. The policy network selects the corresponding action a based on the current environmental state and the value evaluation of the value network, and then updates the state s and reward value of the agent at the next moment based on the current action a, and determines whether to end the task. At the same time, the three are stored in the data cache area for the learning of the agent. At the end of each round, the reward value is discounted and decayed, and the data in the cache is extracted and input into the strategy optimization algorithm for updating, including calculating the strategy loss, calculating the value loss, and updating the network parameters according to the loss; in each update iteration, the state s, action a, and reward r are used to update the network parameter weights and adjust the hyperparameters of the DDPG algorithm; 2.5) Use network simulation tools to conduct preliminary tests to verify the effectiveness of the DDPG algorithm model; 2.6) The optimized model is deployed to the actual system, and the intelligent agent selects the optimal action according to the current state to optimize the operation process of the OTN private network.

2. The method for optimizing the operation process of an OTN private network based on a deterministic policy gradient reinforcement learning algorithm according to claim 1 is characterized in that: The step 1) constructs a digital twin model of the OTN private network operation process, as follows: 1.1) Collect data on the operation of the OTN private network, including: Network traffic data, including: packet rate, latency, packet loss rate, bandwidth utilization; Node status data, including: temperature, optical power; Equipment operation load data, including: CPU usage, memory usage, and throughput; 1.2) Data preprocessing: cleaning and normalizing the collected raw data; Including data standardization of bandwidth, delay, jitter, and packet loss rate 1.3) Build a digital twin model of the OTN private network, including bandwidth model, delay model, jitter model and packet loss rate model; The specific construction methods of the bandwidth model, delay model, jitter model and packet loss rate model are as follows: B link =min(B link1 ,B link2 ,...,B linkN ) J path =max(Δt i -min(Δt i )) Among them, B link is the actual bandwidth of the link, and the link with the smallest bandwidth among all links is selected as the bandwidth of the current link; D path is the total delay from the source node to the destination node, is the delay of the ith link, Process_Delay is the processing delay of the network device; J path is the jitter of the network path, Δt i is the change in the time interval between data packets; PLR is the packet loss rate, N is the number of lost packets, and T is the total number of data packets sent.

3. An electronic device, characterized in that: include: one or more processors; as well as a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 2.

4. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 2 is implemented.