OTN network electric layer service dynamic route distribution method and device and storage medium

By using a main neural network to train and generate a target model in the OTN network and dynamically allocating bandwidth resources, the problems of uneven service transmission latency and channel utilization in the OTN network are solved, thereby improving network resource utilization efficiency and service transmission quality.

CN121967298APending Publication Date: 2026-05-01FIBERHOME TELECOMMUNICATION TECHNOLOGIES CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FIBERHOME TELECOMMUNICATION TECHNOLOGIES CO LTD
Filing Date
2026-01-20
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing OTN networks struggle to simultaneously optimize service transmission latency, channel utilization, and channel resource balance during service transmission. Traditional methods cannot meet the core demands of multi-service carrying requirements and resource optimization.

Method used

A dynamic routing allocation method based on a main neural network is adopted. The main neural network is trained with training set data to generate a target model. The target value of the output path is determined according to the new service request, and bandwidth resources are allocated. The latency, channel utilization and load balancing are comprehensively considered.

Benefits of technology

It optimizes service transmission latency, improves channel utilization, and balances resources in OTN networks, thereby enhancing network resource utilization efficiency and service transmission quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967298A_ABST
    Figure CN121967298A_ABST
Patent Text Reader

Abstract

An OTN network electric layer service dynamic route allocation method, apparatus and device, and a computer readable storage medium, the method comprising: training a main neural network according to constructed training set data to obtain target parameters of the trained main neural network, the training set data comprising multiple groups of training data, and the training set data comprising multiple groups of training data; the training data comprises a current global network state, a service request action, a return value and a next-moment global network state; updating parameters of a target neural network in an OTN network controller according to the obtained target parameters to generate a target model; outputting a target value of each path according to an obtained new service request based on the target model; and determining a target path according to the target value of each path, and allocating bandwidth resources to the target path, thereby solving the technical problem that the balance of service transmission delay, channel utilization rate and channel resources is difficult to optimize at the same time in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

OTN Network Electrical Layer Service Dynamic Routing Allocation Methods, Equipment, and Storage Media Technical Field

[0001] This application relates to the field of communications, specifically to a method, apparatus, device, and computer-readable storage medium for dynamic routing allocation of electrical layer services in an OTN network. Background Technology

[0002] In the Optical Transport Network (OTN) architecture, backbone optical channel planning and access-side service aggregation and transmission constitute the core networking logic. Operators use backbone optical channels to carry high-capacity services across regions, while diverse customer services on the access side need to be aggregated and then accessed through optical channels to complete end-to-end transmission. Under this architecture, the core requirement for service routing is to achieve efficient scheduling of customer services to suitable optical channels, ensuring accurate and stable transmission of services to their destinations. This efficiency directly determines network resource utilization and service transmission quality.

[0003] With the explosive growth of emerging services such as 5G, cloud computing, and high-definition video, customer services are characterized by surging bandwidth demands, sensitivity to transmission latency, and diverse service types, placing higher demands on the accuracy and multi-objective optimization capabilities of OTN network service routing. Traditional service routing solutions are mostly based on the K-shortest path algorithm, which calculates multiple candidate paths and selects the path with sufficient resources to complete service deployment. While this method can initially meet service connectivity requirements, it has significant limitations.

[0004] Specifically, traditional methods struggle to balance multi-dimensional optimization objectives: on the one hand, relying solely on path resource sufficiency as the core selection criterion easily overlooks the crucial indicator of service transmission latency, failing to meet the needs of latency-sensitive services; on the other hand, the lack of a holistic consideration of channel utilization may lead to an imbalance where some channels are overloaded while others are idle, reducing the overall network resource utilization efficiency. This single-objective-oriented routing model is no longer suitable for the current OTN network's multi-service carrying requirements and core demands for resource optimization, necessitating the exploration of efficient service routing technologies that balance multiple objectives. Summary of the Invention

[0005] This application provides a method, apparatus, device, and computer-readable storage medium for dynamic routing allocation of electrical layer services in an OTN network, which can solve the technical problem in the prior art that it is difficult to simultaneously optimize service transmission delay, channel utilization, and channel resource balance.

[0006] In a first aspect, embodiments of this application provide a method for dynamic routing allocation of electrical layer services in an OTN network. The method includes: training a main neural network based on constructed training set data to obtain target parameters of the trained main neural network, wherein the training set data includes multiple sets of training data, including the current global network state, service request actions, reward values, and the global network state at the next moment; updating the parameters of a target neural network located in the OTN network controller based on the obtained target parameters to generate a target model; based on the target model, outputting target values ​​for each path according to newly obtained service requests; determining target paths based on the target values ​​of each path, and allocating bandwidth resources to the target paths.

[0007] In conjunction with the first aspect, in one implementation, the step of outputting target values ​​for each path based on the target model and the acquired new business request includes: inputting the acquired new business request into the target model; and performing forward calculation based on the new business request through the target model to output multi-dimensional vector values, wherein each vector value is the target value for the corresponding path.

[0008] In conjunction with the first aspect, in one implementation, determining a target path based on the target values ​​of each of the paths and allocating bandwidth resources to the target path includes: comparing the target values ​​of each of the paths; determining the maximum target value from the plurality of target values ​​and using the path corresponding to the maximum target value as the target path; and allocating bandwidth resources to the target path.

[0009] In conjunction with the first aspect, in one implementation, training the main neural network based on the constructed training set data to obtain the target parameters of the trained main neural network includes: sequentially inputting the constructed training set data into the main neural network; obtaining a target prediction vector for a service request action at the next time step based on the global network state at the next time step using the main neural network; obtaining a target value of the main neural network based on the target prediction vector and the reward value; obtaining a prediction vector based on the global network state using the main neural network, wherein the prediction vector includes the global network state and the service request action; determining whether the main neural network is in a convergent state based on the target value and the prediction vector, and obtaining the target parameters of the main neural network after it is in a convergent state.

[0010] In conjunction with the first aspect, in one implementation, determining whether the main neural network is in a convergent state based on the target value and the prediction vector includes: calculating the mean squared error of the target value and the prediction vector to obtain a corresponding loss value; determining whether the main neural network is in a convergent state based on the loss value; and determining that the main neural network is in a convergent state if the loss value is less than or equal to a preset loss value.

[0011] In conjunction with the first aspect, in one implementation, before determining that the main neural network is in a convergent state, the method further includes: obtaining the number of training iterations of the main neural network; if the number of training iterations of the main neural network is greater than or equal to a preset number, then the main neural network is determined to be in a convergent state.

[0012] In conjunction with the first aspect, in one implementation, before training the main neural network based on the constructed training set data to obtain the target parameters of the trained main neural network, the method further includes: initializing the global network state and service requests based on the network topology generated by the simulator, and generating the current global network state; generating a service request action based on the current global network state and the obtained service request instructions; determining the target optical path based on the feasibility verification of the service request action, and obtaining the remaining bandwidth and wavelength utilization of the target optical path; determining the reward value and the global network state at the next moment based on the remaining bandwidth and wavelength utilization of the target optical path; and constructing training set data based on the current global network state, the service request action, the reward value, and the global network state at the next moment.

[0013] Secondly, embodiments of this application provide an OTN network electrical layer service dynamic routing allocation device, comprising: an acquisition module, configured to train a main neural network based on constructed training set data to obtain target parameters of the trained main neural network, wherein the training set data includes multiple sets of training data, including the current global network state, service request actions, reward values, and the next-time global network state; a generation module, configured to update the parameters of a target neural network located in the OTN network controller based on the acquired target parameters to generate a target model; an output module, configured to output target values ​​for each path based on the target model and the acquired new service requests; and a determination and allocation module, configured to determine a target path based on the target values ​​of each path and allocate bandwidth resources to the target path.

[0014] Thirdly, embodiments of this application provide an OTN network electrical layer service dynamic routing allocation device, the OTN network electrical layer service dynamic routing allocation device including a processor, a memory, and an OTN network electrical layer service dynamic routing allocation program stored in the memory and executable by the processor, wherein when the OTN network electrical layer service dynamic routing allocation program is executed by the processor, it implements the steps of the OTN network electrical layer service dynamic routing allocation method as described above.

[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing an OTN network electrical layer service dynamic routing allocation program, wherein when the OTN network electrical layer service dynamic routing allocation program is executed by a processor, it implements the steps of the OTN network electrical layer service dynamic routing allocation method as described above.

[0016] The beneficial effects of the technical solution provided in this application include: training the main neural network based on the constructed training set data to obtain the target parameters of the trained main neural network, wherein the training set data includes multiple sets of training data, including the current global network state, service request actions, reward values, and the global network state at the next moment; updating the parameters of the target neural network located in the OTN network controller according to the obtained target parameters to generate a target model; based on the target model, outputting the target values ​​of each path according to the obtained new service requests; determining the target path according to the target values ​​of each path, and allocating bandwidth resources to the target path, thereby solving the technical problem in the prior art that it is difficult to simultaneously optimize service transmission latency, channel utilization, and channel resource balance. Attached Figure Description

[0017] Figure 1 is a flowchart illustrating the first embodiment of the OTN network electrical layer service dynamic routing allocation method of this application; Figure 2 is a functional module diagram illustrating an embodiment of the OTN network electrical layer service dynamic routing allocation device of this application; Figure 3 is a hardware structure diagram illustrating the OTN network electrical layer service dynamic routing allocation device involved in the embodiment scheme of this application. Detailed Implementation

[0018] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0019] First, some of the technical terms used in this application will be explained to help those skilled in the art understand this application.

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0021] In a first aspect, embodiments of this application provide a method for dynamic routing allocation of electrical layer services in an OTN network.

[0022] In one embodiment, referring to Figure 1, which is a flowchart of the first embodiment of the OTN network electrical layer service dynamic routing allocation method of this application, the method includes: Step S10: Training the main neural network based on the constructed training set data to obtain the target parameters of the trained main neural network, wherein the training set data includes multiple sets of training data, and the training data includes the current global network state, service request actions, reward values, and the global network state at the next moment; exemplaryly, the training set data includes multiple sets of training data, and each set of training data includes the current global network state. Business request actions Return value and the global network state at the next moment The constructed training data is input into the main neural network to train it. Once the main neural network is determined to be in a convergent state after training, the target parameters of the main neural network after convergence are obtained. Before training begins, a main neural network for function approximation needs to be constructed. Network structure definition: This invention constructs two neural networks with identical structures but different parameter update strategies, both employing a fully connected feedforward neural network structure. Input layer: The input to the neural network is a state feature vector. This vector consists of two parts: features of the current business request (source node number, destination node number, requested bandwidth); and network state features (normalized remaining bandwidth, utilization, latency, etc. of critical links in the network). The input layer has a dimension of 100. Hidden layers: The network contains two hidden layers. The first hidden layer contains 128 neurons, and the second hidden layer contains 64 neurons. Each neuron is followed by a ReLU activation function. Output layer: The dimension of the output layer is equal to the total number of actions the agent can choose in all states. The value of each input encoding point represents the Q-value estimate for choosing the corresponding action in the input state.

[0023] Specifically, training the main neural network based on the constructed training set data to obtain the target parameters of the trained main neural network includes: sequentially inputting the constructed training set data into the main neural network; obtaining the target prediction vector of the service request action at the next time step based on the global network state at the next time step through the main neural network; obtaining the target value of the main neural network based on the target prediction vector and the reward value; obtaining the prediction vector based on the global network state through the main neural network, wherein the prediction vector includes the global network state and the service request action; determining whether the main neural network is in a convergent state based on the target value and the prediction vector, and obtaining the target parameters of the main neural network after it is in a convergent state.

[0024] As an example, a small batch of samples is uniformly and randomly drawn from the experience replay pool. , , , Random sampling breaks the correlation between samples, which can significantly improve the stability and efficiency of training.

[0025] Compute the main neural network: for each sample in the mini-batch ( , , , Perform the following operations in sequence to... Input the main neural network to obtain all possible actions. Given the target Q-value prediction vector, perform a maximum value operation on this vector and select the largest target Q-value. Substitute into the equation to calculate The target value is calculated, where It's the return value. The discount factor is between 0 and 1, used to balance the importance of immediate rewards and long-term returns, and 'a' is the weighting coefficient for bandwidth utilization.

[0026] Calculate the loss value and the current state. Input the main neural network to obtain its response to the action. Original Q-value prediction Calculate the target value and predicted value of the main neural network. The mean squared error between them is used as the loss function: Update the main neural network using gradient descent, updating only the parameters w of the main Q-network to minimize the loss L calculated in the previous step. The goal is to make the main neural network more adaptable to different network conditions. The predicted values ​​gradually approach a better and more stable target value. The convergence condition determination continues the above iterative training cycle until one of the following convergence conditions is met: Condition 1: Performance Convergence: The prediction performance of the main Q-network tends to stabilize. Specifically, the value of the loss function L drops below a certain low threshold and no longer decreases significantly, or key indicators such as average service congestion rate and resource utilization evaluated on the validation set no longer improve for several consecutive cycles. Condition 2: Reaching the Maximum Number of Iterations: The training reaches the pre-set maximum number of iterations. Once the model converges, the training process ends.

[0027] Specifically, before training the main neural network based on the constructed training set data to obtain the target parameters of the trained main neural network, the process further includes: initializing the global network state and service requests based on the network topology generated by the simulator, and generating the current global network state; generating service request actions based on the current global network state and the obtained service request instructions; determining the target optical path by verifying the feasibility of the service request actions, and obtaining the remaining bandwidth and wavelength utilization of the target optical path; determining the reward value and the global network state at the next moment based on the remaining bandwidth and wavelength utilization of the target optical path; and constructing training set data based on the current global network state, service request actions, reward value, and the global network state at the next moment.

[0028] As an example, in an OTN optical layer topology, the optical channel (OCH) is the carrier of signal transmission, and the nodes are reconfigurable optical add-drop multiplexers (ROADMs). For diagram compression, the channel between two nodes containing an OCH is called a link. The number of OCHs in a link is the number of available optical channels in the link. Let V represent the set of nodes, and use... Let |I| represent a link between nodes o and d, where the number of optical paths is denoted by |I|. This indicates the wavelength number of an optical path. This indicates the remaining bandwidth of this optical path. This indicates the utilization rate of this wavelength across the entire network. This represents the transmission latency of a link. All such links constitute the entire network's resource set, and set E is called the current state of the entire network. A service is represented by Service(s,d,b), where s represents the source node, d represents the destination node, and b represents the bandwidth required by the service. Each time an electrical layer service is established, it requires bandwidth in the optical path and updates the link state. The overall network state E consists of the current states of all links. Each service expands outward from the source node along the links, and the set of available links constitutes the action space for that state. The reward value for each action is defined. Taking into account link latency, node latency, bandwidth utilization, and load balancing, the reward function is defined as follows: .in Indicates bandwidth utilization. Indicates load balancing, and The weighting coefficients representing time delay. For node delay, since the three targets have different dimensions, Min-Max is used here to normalize each dimension to the range of [0, 1]. This indicates that the remaining bandwidth is normalized; a larger value indicates that the link is more idle. The wavelength utilization rate is normalized; the larger the value, the more congested the link is, so a negative sign is added in front of it. The link delay is normalized; a larger value indicates higher delay, hence the negative sign.

[0029] For example, initial state generation: When the simulator starts, it initializes the global network state according to the predefined network topology. . It contains the initial state of all links, initializing all optical paths for each link. Remaining bandwidth Wavelength utilization The system calculates the transmission delay for each optical path, taking into account distance and the speed of light. Service request generation: The simulator generates a service request Service(s,d,b) according to a preset random distribution, where the source node s, destination node d, and required bandwidth b are all randomly sampled. Status Build: Set the current global network state Together with the currently arriving business requests, they constitute the environment state for reinforcement learning. and pass it to the intelligent agent. .

[0030] Receive actions and simulate execution (Step), receive business request actions. The agent is based on And its strategy to select a business request action This business request action Select a specific end-to-end OCH path for the current business request. Feasibility verification: The simulator first verifies the action. The legality of the selected path. Does the chosen path exist in the topology? At least one optical path on the path has a remaining bandwidth greater than or equal to a threshold b, i.e. If the action is illegal, the path does not exist in the topology, or resources are insufficient, then give... A value of -1000 indicates a decision error. Then, proceed to the first step to reset and begin the next training round. Resource allocation and state update: If the action is valid, select the optimal optical path on the chosen path. Update the remaining bandwidth of this optical path. Recalculate the utilization rate of this wavelength across the entire network. Based on this change, the state of the global network at the next time step is calculated. Calculate the return value. : Calculate the action to be performed based on the reward function. The instant reward obtained afterward .

[0031] Step S20: Update the parameters of the target neural network located in the OTN network controller according to the obtained target parameters to generate the target model; exemplary, the obtained target parameters are... Update the parameters of the target neural network in the OTN network controller, where the target neural network has the same network structure as the main application network, so that the updated target neural network is in a convergent state and the target model is generated.

[0032] Step S30: Based on the target model, output the target values ​​for each path according to the acquired new service request; exemplary, input the acquired new service request action into the target model; perform forward computation based on the new service request action through the target model, and output multi-dimensional vector values, where each vector value is the target value for the corresponding path. For example, when a new service request arrives, the controller constructs a 100-dimensional feature vector from the current real-time network state and the service request, and inputs it into the target model. Perform forward computation through the target model, and output a 20-dimensional vector, where each value represents the Q-value of selecting the corresponding path.

[0033] Step S40: Determine the target path based on the target value of each path, and allocate bandwidth resources to the target path.

[0034] As an example, the target values ​​of each path are compared; the maximum target value is determined from multiple target values, and the path corresponding to the maximum target value is taken as the target path; thus, bandwidth resources are allocated to the target path.

[0035] In this embodiment, the main neural network is trained based on the constructed training set data to obtain the target parameters of the trained main neural network. The training set data includes multiple sets of training data, including the current global network state, service request actions, reward values, and the global network state at the next moment. The parameters of the target neural network located in the OTN network controller are updated according to the obtained target parameters to generate a target model. Based on the target model, the target values ​​of each path are output according to the obtained new service requests. The target paths are determined according to the target values ​​of each path, and bandwidth resources are allocated to the target paths. This solves the technical problem in the prior art that it is difficult to simultaneously optimize service transmission latency, channel utilization, and channel resource balance.

[0036] Secondly, embodiments of this application also provide an OTN network electrical layer service dynamic routing allocation device.

[0037] In one embodiment, referring to Figure 2, which is a functional block diagram of an embodiment of the OTN network electrical layer service dynamic routing allocation device of this application, the OTN network electrical layer service dynamic routing allocation device includes: an acquisition module 10, used to train the main neural network according to the constructed training set data to obtain the target parameters of the trained main neural network, wherein the training set data includes multiple sets of training data, including the current global network state, service request actions, reward values, and the global network state at the next moment; a generation module 20, used to update the parameters of the target neural network located in the OTN network controller according to the acquired target parameters to generate a target model; an output module 30, used to output the target values ​​of each path based on the target model and the acquired new service requests; and a determination and allocation module 40, used to determine the target path according to the target values ​​of each path and allocate bandwidth resources to the target path.

[0038] Furthermore, the output module 30 is used to: input the acquired new business request into the target model; perform forward calculation based on the new business request through the target model, and output multi-dimensional vector values, wherein each vector value is the target value of the corresponding path.

[0039] Furthermore, the determination and allocation module 40 is used to: compare the target values ​​of each of the acquired paths; determine the maximum target value from the multiple target values, and take the path corresponding to the maximum target value as the target path; and allocate bandwidth resources to the target path.

[0040] Further, the acquisition module 10 is used to: sequentially input the constructed training set data into the main neural network; obtain the target prediction vector of the service request action at the next time step based on the global network state at the next time step through the main neural network; obtain the target value of the main neural network according to the target prediction vector and the reward value; obtain the prediction vector based on the global network state through the main neural network, wherein the prediction vector includes the global network state and the service request action; determine whether the main neural network is in a convergent state according to the target value and the prediction vector, and obtain the target parameters of the main neural network after it is in a convergent state.

[0041] Furthermore, in one embodiment, the OTN network electrical layer service dynamic routing allocation device further includes a new module for: calculating the mean square error of the target value and the prediction vector to obtain the corresponding loss value; determining whether the main neural network is in a convergent state based on the loss value; and determining that the main neural network is in a convergent state if the loss value is less than or equal to a preset loss value.

[0042] Furthermore, in one embodiment, the OTN network electrical layer service dynamic routing allocation device further includes a new module for: obtaining the number of training iterations of the main neural network; if the number of training iterations of the main neural network is greater than or equal to a preset number, then determining that the main neural network is in a convergent state.

[0043] Furthermore, in one embodiment, the OTN network electrical layer service dynamic routing allocation device further includes a new module, used for: initializing the global network state and service requests based on the network topology generated by the simulator, and generating the current global network state; generating a service request action based on the current global network state and the obtained service request instruction; determining the target optical path by verifying the feasibility of the service request action, and obtaining the remaining bandwidth and wavelength utilization of the target optical path; determining the reward value and the global network state at the next moment based on the remaining bandwidth and wavelength utilization of the target optical path; and constructing training set data based on the current global network state, service request action, reward value, and the global network state at the next moment.

[0044] The functions of each module in the aforementioned OTN network electrical layer service dynamic routing allocation device correspond to the steps in the aforementioned OTN network electrical layer service dynamic routing allocation method embodiment, and their functions and implementation processes will not be described in detail here.

[0045] Thirdly, embodiments of this application provide an OTN network electrical layer service dynamic routing allocation device, which can be a personal computer (PC), laptop computer, server, or other device with data processing capabilities.

[0046] Referring to Figure 3, which is a schematic diagram of the hardware structure of the OTN network electrical layer service dynamic routing allocation device involved in the embodiment of this application, the OTN network electrical layer service dynamic routing allocation device may include a processor, a memory, a communication interface, and a communication bus.

[0047] The communication bus can be of any type and is used to interconnect the processor, memory, and communication interface.

[0048] The communication interface includes input / output (I / O) interfaces, physical interfaces, and logical interfaces used for interconnecting devices within the OTN network electrical layer service dynamic routing allocation device, as well as interfaces used for interconnecting the OTN network electrical layer service dynamic routing allocation device with other devices (such as other computing devices or user equipment). Physical interfaces can be Ethernet interfaces, fiber optic interfaces, ATM interfaces, etc.; user equipment can be displays, keyboards, etc.

[0049] Memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical storage, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.

[0050] The processor can be a general-purpose processor, which can call the OTN network electrical layer service dynamic routing allocation program stored in the memory and execute the OTN network electrical layer service dynamic routing allocation method provided in the embodiments of this application. For example, the general-purpose processor can be a central processing unit (CPU). The method executed when the OTN network electrical layer service dynamic routing allocation program is called can be referred to the various embodiments of the OTN network electrical layer service dynamic routing allocation method of this application, and will not be repeated here.

[0051] Those skilled in the art will understand that the hardware structure shown in Figure 3 does not constitute a limitation of this application, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0052] Fourthly, embodiments of this application also provide a computer-readable storage medium.

[0053] The present application stores an OTN network electrical layer service dynamic routing allocation program on a computer-readable storage medium, wherein when the OTN network electrical layer service dynamic routing allocation program is executed by a processor, it implements the steps of the OTN network electrical layer service dynamic routing allocation method as described above.

[0054] The method implemented when the OTN network electrical layer service dynamic routing allocation procedure is executed can be referred to in the various embodiments of the OTN network electrical layer service dynamic routing allocation method of this application, and will not be repeated here.

[0055] It should be noted that the sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0056] The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus. The terms "first," "second," and "third," etc., are used to distinguish different objects, etc., and do not indicate a sequence, nor do they limit "first," "second," and "third" to different types.

[0057] In the description of the embodiments of this application, terms such as "exemplary," "for example," or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary," "for example," or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary," "for example," or "for instance" is intended to present the relevant concepts in a concrete manner.

[0058] In the description of the embodiments of this application, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in the text is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more.

[0059] In some processes described in the embodiments of this application, multiple operations or steps are included in a specific order. However, it should be understood that these operations or steps may not be executed in the order they appear in the embodiments of this application, or they may be executed in parallel. The sequence number of the operation is only used to distinguish different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed sequentially or in parallel, and these operations or steps may be combined.

[0060] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device to execute the methods described in the various embodiments of this application.

[0061] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for dynamic routing allocation of electrical layer services in an OTN network, characterized in that, The OTN network electrical layer service dynamic routing allocation method includes: training a main neural network based on a constructed training set of data to obtain target parameters of the trained main neural network, wherein the training set of data includes multiple sets of training data, including the current global network state, service request actions, reward values, and the global network state at the next moment; updating the parameters of the target neural network located in the OTN network controller based on the obtained target parameters to generate a target model; based on the target model, outputting target values ​​for each path according to the obtained new service requests; determining target paths based on the target values ​​of each path, and allocating bandwidth resources to the target paths.

2. The OTN network electrical layer service dynamic routing allocation method as described in claim 1, characterized in that, The step of outputting target values ​​for each path based on the target model and the acquired new business request includes: inputting the acquired new business request action into the target model; and performing forward calculation based on the new business request action through the target model to output multi-dimensional vector values, wherein each vector value is the target value for the corresponding path.

3. The OTN network electrical layer service dynamic routing allocation method as described in claim 1, characterized in that, The step of determining a target path based on the target values ​​of each path and allocating bandwidth resources to the target path includes: comparing the target values ​​of each path; determining the maximum target value from the multiple target values ​​and using the path corresponding to the maximum target value as the target path; and allocating bandwidth resources to the target path.

4. The OTN network electrical layer service dynamic routing allocation method as described in claim 1, characterized in that, The step of training the main neural network based on the constructed training set data to obtain the target parameters of the trained main neural network includes: sequentially inputting the constructed training set data into the main neural network; obtaining the target prediction vector of the service request action at the next time step based on the global network state at the next time step through the main neural network; obtaining the target value of the main neural network based on the target prediction vector and the reward value; obtaining the prediction vector based on the global network state through the main neural network, wherein the prediction vector includes the global network state and the service request action; determining whether the main neural network is in a convergent state based on the target value and the prediction vector, and obtaining the target parameters of the main neural network after it is in a convergent state.

5. The OTN network electrical layer service dynamic routing allocation method as described in claim 4, characterized in that, The step of determining whether the main neural network is in a convergent state based on the target value and the prediction vector includes: calculating the mean square error of the target value and the prediction vector to obtain the corresponding loss value; determining whether the main neural network is in a convergent state based on the loss value; and determining that the main neural network is in a convergent state if the loss value is less than or equal to a preset loss value.

6. The OTN network electrical layer service dynamic routing allocation method as described in claim 5, characterized in that, Before determining that the main neural network is in a convergent state, the method further includes: obtaining the number of training iterations of the main neural network; if the number of training iterations of the main neural network is greater than or equal to a preset number, then the main neural network is determined to be in a convergent state.

7. The OTN network electrical layer service dynamic routing allocation method as described in claim 1, characterized in that, Before training the main neural network based on the constructed training set data to obtain the target parameters of the trained main neural network, the method further includes: initializing the global network state and service requests based on the network topology generated by the simulator, and generating the current global network state; generating service request actions based on the current global network state and the obtained service request instructions; determining the target optical path based on the feasibility verification of the service request actions, and obtaining the remaining bandwidth and wavelength utilization of the target optical path; determining the reward value and the global network state at the next moment based on the remaining bandwidth and wavelength utilization of the target optical path; and constructing training set data based on the current global network state, service request actions, reward value, and the global network state at the next moment.

8. A dynamic routing allocation device for electrical layer services in an OTN network, characterized in that, The OTN network electrical layer service dynamic routing allocation device includes: an acquisition module, used to train the main neural network based on the constructed training set data to obtain the target parameters of the trained main neural network, wherein the training set data includes multiple sets of training data, including the current global network state, service request actions, reward values, and the global network state at the next moment; a generation module, used to update the parameters of the target neural network located in the OTN network controller based on the acquired target parameters to generate a target model; an output module, used to output the target values ​​of each path based on the target model and the acquired new service requests; and a determination and allocation module, used to determine the target path based on the target values ​​of each path and allocate bandwidth resources to the target path.

9. A dynamic routing allocation device for electrical layer services in an OTN network, characterized in that, The OTN network electrical layer service dynamic routing allocation device includes a processor, a memory, and an OTN network electrical layer service dynamic routing allocation program stored in the memory and executable by the processor, wherein when the OTN network electrical layer service dynamic routing allocation program is executed by the processor, it implements the steps of the OTN network electrical layer service dynamic routing allocation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an OTN network electrical layer service dynamic routing allocation program, wherein when the OTN network electrical layer service dynamic routing allocation program is executed by a processor, it implements the steps of the OTN network electrical layer service dynamic routing allocation method as described in any one of claims 1 to 7.