An autonomous routing planning method, device and system for an electrically elastic optical network
By using an autonomous routing planning method to generate and update routing planning schemes, the problem of poor autonomy in traditional power communication optical networks is solved, enabling rapid adaptation to changes in power services in routing and spectrum planning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-23
- Publication Date
- 2026-03-24
AI Technical Summary
Traditional power communication optical network routing and spectrum planning methods lack autonomy and cannot adapt to the time-varying nature of network resources and the dynamic accessibility of power services, resulting in slow convergence speed and poor service assurance.
An autonomous routing planning method is adopted. By acquiring available routing data and frequency slot data at the physical network layer, a routing planning scheme is generated, a preference value is calculated and simulated transmission is performed, the preference value is updated using the learning reward value, and the optimal routing planning scheme is dynamically selected.
It enables autonomous planning of network routing and spectrum, improves the service reliability and convergence speed of power services, and adapts to the needs of power services with large-scale concurrent access.
Smart Images

Figure CN116647494B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of communication, in particular to an autonomous routing planning method, device and system of a power elastic optical network. BACKGROUND
[0002] The traditional power communication optical network adopts wavelength division multiplexing technology for communication, however, the wavelength unit of the wavelength division multiplexing optical network is fixed, the transmission rate and bandwidth size of the optical transponder are fixed, and the modulation mode of services of different transmission distances is fixed, and the coarse-grained spectrum division mode limits the flexibility of the routing planning of the communication network, and cannot adapt to the transmission requirements of differentiated power services.
[0003] The power elastic optical network can realize smaller-grained communication bandwidth scheduling through dynamic allocation of spectrum according to differentiated requirements of power services, and improve the resource utilization rate of the network, however, the core problem of the power elastic optical network is the routing and spectrum planning problem, that is, according to the bandwidth size of the service request, an optical path is established between the source node and the destination node, and continuous spectrum resources are allocated, at present, the routing and spectrum planning problem of the power elastic optical network is mainly solved by using static algorithms and dynamic algorithms, wherein the static algorithms include exact algorithms, static heuristic algorithms and intelligent optimization algorithms, and the dynamic algorithms include dynamic heuristic algorithms and learning algorithms.
[0004] The traditional static power elastic optical network routing and spectrum planning method needs to comprehensively obtain the source node, the destination node and the bandwidth demand information, and obtains the optical routing and spectrum planning scheme of the service request in the offline state, this method lacks the autonomous planning capability of the network routing and spectrum, and cannot adapt to the time-varying nature of the network resources and the dynamic arrival of the power services, and the power elastic optical network routing and spectrum planning method based on the traditional learning algorithm does not consider the time-varying nature of the communication network state and the randomness of the service request, and the autonomy of the routing and spectrum planning is poor, which leads to slow convergence speed, a large number of service requests are needed for testing, and the service guarantee is poor. SUMMARY
[0005] The technical problem to be solved by the present application is to provide an autonomous routing planning method, device and system of a power elastic optical network, which can realize autonomous planning of network routing and spectrum, and effectively adapt to large-scale concurrent access of power services.
[0006] In order to solve the above technical problems, the present application provides an autonomous routing planning method of a power elastic optical network, comprising:
[0007] When a service request is detected to arrive, all available routing data and all available frequency slot data of a physical network layer are acquired, all routing planning schemes corresponding to the service request are generated based on the all available routing data and the all available frequency slot data, and a preference value of the service request for each routing planning scheme is calculated;
[0008] A first routing planning scheme corresponding to the maximum preference value is selected, and the service request is simulated to be transmitted according to the first routing planning scheme, a first routing and frequency slot selection indication variable and a first learning reward value of the service request are calculated, and a current first simulation transmission frequency is recorded;
[0009] Based on the first routing and frequency slot selection indication variable, a first learning reward sample mean is calculated, and the preference value of the first routing planning scheme is updated based on the first learning reward sample mean;
[0010] Based on the first learning reward value, when the successful transmission indication variable is 1, the simulation transmission of the service request is stopped, the first routing planning scheme is taken as an optimal routing planning scheme corresponding to the service request, and the optimal routing planning scheme is sent to the physical network layer, so that the physical network layer transmits the service request according to the optimal routing planning scheme.
[0011] The autonomous routing planning method of the power elastic optical network provided by the application further comprises:
[0012] Based on the first learning reward value, when the successful transmission indication variable is not 1 and the current first simulation transmission frequency is not equal to a preset simulation transmission threshold, a second routing planning scheme corresponding to the maximum preference value is selected again, and the second routing planning scheme is simulated to be transmitted until the successful transmission indication variable is 1 or the current first simulation transmission frequency is equal to the preset simulation transmission threshold, and the second routing planning scheme is taken as the optimal routing planning scheme corresponding to the service request.
[0013] In a possible implementation manner, based on the all available routing data and the all available frequency slot data, the all routing planning schemes corresponding to the service request are generated, and specifically comprising:
[0014] The number of service request frequency slots of the service request is calculated, all candidate routes of the service request are determined according to the number of service request frequency slots, and the number of routing frequency slots of the service request on each candidate route is calculated;
[0015] The all routing planning schemes corresponding to the service request are generated according to the all candidate routes and the number of routing frequency slots.
[0016] In a possible implementation, the preference value of the service request for each routing plan is calculated based on a preset preference value calculation formula, where the preset preference value calculation formula is as follows:
[0017] ;
[0018] wherein, is the preference value of the tth service request, is the mean value of the learning reward sample corresponding to the tth service request, is a trade-off coefficient between exploration and utilization, is the selection frequency of the kth route and the fth frequency slot in the tth service request until the current time slot.
[0019] In a possible implementation, the first route and frequency slot selection indicator variable and the first learning reward value of the service request are calculated by simulating transmission of the service request according to the first routing plan, and the calculation specifically includes the following steps.
[0020] The first preference value corresponding to the first routing plan is obtained, the first preference value is substituted into a preset first route and frequency slot selection indicator variable calculation formula, and the first route and frequency slot selection indicator variable of the first routing plan is calculated, where the preset first route and frequency slot selection indicator variable calculation formula is as follows:
[0021] ;
[0022] wherein, is the first route and frequency slot selection indicator variable corresponding to the tth service request,
[0023] When the service request is simulated to be transmitted according to the first routing plan, the data transmission rate, the transmission power consumption and the successful transmission indicator variable of the service request are obtained, the data transmission rate, the transmission power consumption and the successful transmission indicator variable are substituted into a preset learning reward value calculation formula, and the first learning reward value of the first routing plan is calculated, where the preset learning reward value calculation formula is as follows:
[0024] ;
[0025] wherein, is the first learning reward value corresponding to the tth service request, is the data transmission rate, is the transmission power consumption corresponding to the tth service request, is the successful transmission indicator variable corresponding to the tth service request, .
[0026] In a possible implementation, the first learning reward sample mean is calculated based on the first route and frequency slot selection indicator variable, and the preference value of the first route planning scheme is updated based on the first learning reward sample mean, specifically comprising:
[0027] The selection times of the kth route and the fth frequency slot until the current time slot are obtained, the selection times are input into a preset selection times updating formula, the first selection times of the kth route and the fth frequency slot corresponding to the first route planning scheme until the current time slot are obtained, and the selection times are updated to the first selection times;
[0028] ;
[0029] In the formula, is the first selection times of the kth route and the fth frequency slot until the current time slot corresponding to the tth service request, is the selection times of the kth route and the fth frequency slot until the current time slot corresponding to the t-1th service request, is the first route and frequency slot selection indicator variable corresponding to the tth service request;
[0030] The first route and frequency slot selection indicator variable and the first learning reward value of the first route planning scheme, and the updated selection times are obtained, the first route and frequency slot selection indicator variable, the first learning reward value and the selection times are substituted into a preset learning reward sample mean calculation formula, a first learning reward sample mean is obtained, the learning reward sample mean is updated to the first learning reward sample mean, wherein the preset learning reward sample mean calculation formula is as follows:
[0031] ;
[0032] In the formula, is the first learning reward sample mean corresponding to the tth service request, is the learning reward sample mean corresponding to the t-1th service request, is the first route and frequency slot selection indicator variable corresponding to the tth service request, is the first learning reward value corresponding to the tth service request, is the first selection times of the kth route and the fth frequency slot until the current time slot corresponding to the tth service request;
[0033] The updated selection times of the kth route and the fth frequency slot until the current time slot and the updated learning reward sample mean are substituted into a preset preference value calculation formula, the preference value of the first route planning scheme is calculated and updated.
[0034] The application provides an autonomous route planning device of a power elastic optical network, comprising a route planning scheme preference value calculation module, a service request simulation transmission module, a route planning scheme preference value updating module and an optimal route planning scheme determination module.
[0035] The route planning scheme preference value calculation module is configured to, when detecting arrival of a service request, acquire all available route data and all available frequency slot data of a physical network layer, generate all route planning schemes corresponding to the service request based on the all available route data and the all available frequency slot data, and calculate a preference value of each route planning scheme for the service request.
[0036] The service request simulation transmission module is configured to select a first route planning scheme corresponding to a maximum preference value, simulate transmission of the service request according to the first route planning scheme, calculate a first route and frequency slot selection indicator variable and a first learning reward value of the service request, and record a current first simulation transmission frequency.
[0037] The route planning scheme preference value updating module is configured to calculate a first learning reward sample mean based on the first route and frequency slot selection indicator variable, and update the preference value of the first route planning scheme based on the first learning reward sample mean.
[0038] The optimal route planning scheme determination module is configured to, based on the first learning reward value, determine that the simulation transmission of the service request is stopped when a successful transmission indicator variable is 1, take the first route planning scheme as an optimal route planning scheme corresponding to the service request, and send the optimal route planning scheme to the physical network layer so that the physical network layer transmits the service request according to the optimal route planning scheme.
[0039] The optimal route planning scheme determination module is further configured to, based on the first learning reward value, determine that a second route planning scheme corresponding to a maximum preference value is reselected when the successful transmission indicator variable is not 1 and the current first simulation transmission frequency is not equal to a preset simulation transmission threshold, and simulate transmission of the second route planning scheme until the successful transmission indicator variable is 1 or the current first simulation transmission frequency is equal to the preset simulation transmission threshold, and take the second route planning scheme as the optimal route planning scheme corresponding to the service request.
[0040] In a possible implementation, the route planning scheme preference value calculation module is configured to generate all route planning schemes corresponding to the service request based on the all available route data and the all available frequency slot data, and specifically comprises:
[0041] calculating a service request frequency slot number of the service request, determining all candidate routes of the service request according to the service request frequency slot number, and calculating a route frequency slot number of the service request on each candidate route;
[0042] generating all route planning schemes corresponding to the service request according to the all candidate routes and the route frequency slot number.
[0043] In a possible implementation, the route planning scheme preference value calculation module is configured to calculate a preference value of each route planning scheme of the service request based on a preset preference value calculation formula, where the preset preference value calculation formula is as follows:
[0044] ;
[0045] In the formula, preference value of the tth service request, is a preference value of the tth service request, is a learning reward sample mean corresponding to the tth service request, is a compromise coefficient between exploration and utilization, is a selection number of the kth route and the fth frequency slot of the tth service request until the current time slot.
[0046] In a possible implementation, the service request simulation transmission module is configured to perform simulation transmission on the service request according to the first route planning scheme, and calculate a first route and frequency slot selection indication variable and a first learning reward value of the service request, and specifically includes the following steps.
[0047] obtaining a first preference value corresponding to the first route planning scheme, substituting the first preference value into a preset first route and frequency slot selection indication variable calculation formula, and calculating a first route and frequency slot selection indication variable of the first route planning scheme, where the preset first route and frequency slot selection indication variable calculation formula is as follows:
[0048] ;
[0049] In the formula, first route and frequency slot selection indication variable corresponding to the tth service request, is a first route and frequency slot selection indication variable corresponding to the tth service request,
[0050] When the service request is simulated to be transmitted according to the first routing planning scheme, a data transmission rate, a transmission power consumption and a successful transmission indication variable of the service request are obtained, the data transmission rate, the transmission power consumption and the successful transmission indication variable are substituted into a preset learning reward value calculation formula, and a first learning reward value of the first routing planning scheme is calculated.
[0051] ;
[0052] In the formula, is a first learning reward value corresponding to the tth service request, is the data transmission rate, is the transmission power consumption corresponding to the tth service request, is a successful transmission indication variable corresponding to the tth service request, .
[0053] In a possible implementation manner, the routing planning scheme preference value updating module is configured to calculate a first learning reward sample mean based on the first route and frequency slot selection indication variable, and update the preference value of the first routing planning scheme based on the first learning reward sample mean, and specifically includes the following steps.
[0054] The selection times of the kth route and the fth frequency slot until the current time slot are obtained, the selection times are input into a preset selection times updating formula, the first selection times of the kth route and the fth frequency slot corresponding to the first routing planning scheme until the current time slot are obtained, and the selection times are updated to the first selection times.
[0055] ;
[0056] In the formula, is the first selection times of the kth route and the fth frequency slot corresponding to the tth service request until the current time slot, is the selection times of the kth route and the fth frequency slot corresponding to the (t-1)th service request until the current time slot, is the first route and frequency slot selection indication variable corresponding to the tth service request;
[0057] The first route and frequency slot selection indicator variable and the first learning reward value of the first route planning scheme are obtained, and the first route and frequency slot selection indicator variable, the first learning reward value and the selection times are substituted into a preset learning reward sample mean calculation formula to obtain a first learning reward sample mean, and the learning reward sample mean is updated as the first learning reward sample mean, wherein the preset learning reward sample mean calculation formula is as follows:
[0058] ;
[0059] In the formula, is a first learning reward sample mean corresponding to the tth service request, is a learning reward sample mean corresponding to the (t-1) th service request, is a first route and frequency slot selection indicator variable corresponding to the tth service request, is a first learning reward value corresponding to the tth service request, is a first selection times of the kth route and the fth frequency slot up to the current time slot corresponding to the tth service request;
[0060] The updated selection times of the kth route and the fth frequency slot up to the current time slot and the updated learning reward sample mean are substituted into a preset preference value calculation formula, and the preference value of the first route planning scheme is calculated and updated.
[0061] The application provides an electric power elastic optical network, comprising a physical network layer, a twin network layer and a service application layer.
[0062] The twin network layer comprises a north interface and a south interface.
[0063] The twin network layer receives a demand input by the service application layer through the north interface.
[0064] The twin network layer acquires network state data in the physical network layer through the south interface, and sends an optimal route planning scheme to the physical network layer through the south interface after determining the optimal route planning scheme.
[0065] The twin network layer is used for executing the autonomous route planning method of the electric power elastic optical network according to any one of the preceding embodiments.
[0066] The application provides a terminal device, comprising a processor, a memory and a computer program stored in the memory and configured to be executed by the processor, and the processor executes the computer program to implement the autonomous route planning method of the electric power elastic optical network according to any one of the preceding embodiments.
[0067] The application provides a computer readable storage medium comprising a stored computer program, wherein the computer readable storage medium controls a device in which the computer readable storage medium is located to perform the autonomous routing planning method of the power elastic optical network according to any one of the preceding embodiments when the computer program is executed.
[0068] Compared with the prior art, the power elastic optical network autonomous routing planning method, device and system provided by the embodiments of the application have the following beneficial effects:
[0069] When a service request is detected, all available routing data and all available frequency slot data of the physical network layer are acquired, all routing planning schemes corresponding to the service request are generated, the preference value of each routing planning scheme is calculated, the first routing planning scheme corresponding to the maximum preference value is used to simulate transmission of the service request, and the first routing and frequency slot selection indicator variable and the first learning reward value of the service request are calculated; based on the first routing and frequency slot selection indicator variable, the preference value of the first routing planning scheme is updated based on the first learning reward sample mean value; based on the first learning reward value, when the successful transmission indicator variable is 1, the first routing planning scheme is determined as the optimal routing planning scheme corresponding to the service request, and is delivered to the physical network layer; compared with the prior art, the technical scheme of the application accelerates the convergence speed of learning by simulating experiments on routing planning schemes in the twin network layer, dynamically selects the routing planning scheme with the maximum preference value, realizes autonomous planning of network routing and spectrum, and can effectively solve the problems that the traditional machine learning method has poor autonomy and a large number of service request experiments lead to poor service guarantee, and can realize autonomous planning of network routing and spectrum. BRIEF DESCRIPTION OF DRAWINGS
[0070] Figure 1 is a flowchart of an embodiment of the power elastic optical network autonomous routing planning method provided by the application;
[0071] Figure 2 is a structural schematic diagram of an embodiment of the power elastic optical network autonomous routing planning device provided by the application;
[0072] Figure 3 is a structural schematic diagram of an embodiment of the power elastic optical network provided by the application;
[0073] Figure 4 is a structural schematic diagram of the power elastic optical network of an embodiment provided by the application;
[0074] Figure 5 is a routing planning principle diagram of the twin network layer of an embodiment provided by the application. DETAILED DESCRIPTION
[0075] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0076] Embodiment 1
[0077] Referring to Figure 1 , Figure 1 is a flowchart of an embodiment of an autonomous routing planning method of an electric power elastic optical network provided by the present application, as shown in the figure, the method comprises steps 101-104, and the details are as follows: Figure 1
[0078] Step 101: When detecting that a service request arrives, all available routing data of a physical network layer and all available frequency slot data are acquired, all routing planning schemes corresponding to the service request are generated based on the all available routing data and the all available frequency slot data, and a preference value of the service request for each routing planning scheme is calculated.
[0079] In an embodiment, the electric power elastic optical network comprises a physical network layer, a twin network layer and a service application layer, as shown in the figure, Figure 4 is a structural diagram of the electric power elastic optical network. Figure 4
[0080] In an embodiment, the twin network layer comprises a southbound interface and a northbound interface.
[0081] In an embodiment, the physical network layer contains physical entities of the electric power elastic optical network such as routers and optical fiber links, and various network elements interact with the twin network through the southbound interface of the twin network layer to exchange network data and network control information, so as to realize reliable bearing of specific services of the service application layer.
[0082] In an embodiment, the service application layer inputs requirements to the twin network through the northbound interface of the twin network layer, so as to support deployment of electric power services in the twin network layer.
[0083] In an embodiment, the physical network layer and the twin network layer both use sets , denote a node set, denote an optical link set, , denote a frequency slot set on each link, wherein , denotes that there are at most A frequency slot; the frequency slot is the minimum unit of spectrum resource planning in the power optical network, unlike the traditional wavelength division multiplexing optical network, it needs to meet the continuity, consistency, and non-reusable constraints in the route autonomous planning.
[0084] Preferably, the continuity, consistency, and non-reusable constraints met by the frequency slot in the route autonomous planning are illustrated as follows: Figure 5 Figure 5 is a twin network layer routing planning schematic diagram; when a service request needs two frequency slots to be transmitted from the routing node 1 to the routing node 4, although both the link 1→2 and the link 1→3 have two idle frequency slots, but since the two frequency slots on the link 1→3 do not meet the continuity constraint, the link 1→2 is selected to reach the routing node 2; when the service request reaches the routing node 2, although both the link 2→3 and the link 2→4 have two continuous idle frequency slots, but since the frequency slot position on the link 2→4 is different from that on the link 1→2, it does not meet the frequency slot consistency constraint, so the link 2→3 is selected to transmit to the routing node 4; finally, the planning route of the service request is routing node 1→routing node 2→routing node 3→routing node 4, and during the service request duration, the frequency slots 2 and 3 are prohibited from being planned for other service requests.
[0085] In an embodiment, during a service request duration period, the physical network layer continuously receives a plurality of service requests, and generates a service request set based on the plurality of service requests, wherein each service request in the service request set includes a source node, a destination node, a bandwidth requirement, and a duration of the service request.
[0086] Specifically, during the service request duration period, a total of service requests continuously arrive, and the service request set is , wherein ; each service request is represented by a set , wherein represents the source node of the service request, represents the destination node, and , and represent the bandwidth requirement and the duration of the service request, respectively.
[0087] In an embodiment, the twin network layer constructs a digital network that can reflect the physical network state in real time through data interaction with the physical network layer, and performs dynamic autonomous planning of network routing and spectrum; therefore, when a service request is detected to arrive, the twin network layer generates all routing planning schemes corresponding to the service request based on all available routing data and all available frequency slot data of the physical network layer.
[0088] In an embodiment, a service request frequency slot number of the service request is calculated, all candidate routes of the service request are determined according to the service request frequency slot number, and a route frequency slot number of the service request on each candidate route is calculated; all route planning schemes corresponding to the service request are generated according to the all candidate routes and the route frequency slot number.
[0089] Specifically, when the service request frequency slot number of the service request is calculated, the bandwidth requirement, the transmission rate, and the modulation level corresponding to the modulation mode adopted of the service request are obtained, the bandwidth requirement, the transmission rate, and the modulation level are substituted into a preset service request frequency slot number calculation formula to obtain the service request frequency slot number of the service request; wherein the preset service request frequency slot number calculation formula is as follows:
[0090] ;
[0091] In the formula, is the service request frequency slot number, is the bandwidth requirement of the service request, represents the transmission rate supported by a single frequency slot when the BPSK modulation mode is adopted, respectively represent the modulation levels corresponding to the BPSK, QPSK, 8QAM, 16QAM, and 32QAM modulation modes.
[0092] In order to serve the service request, an end-to-end route needs to be established, and the route has a frequency slot block satisfying the constraints of frequency slot continuity, consistency, and non-reusability.
[0093] Specifically, all candidate routes of the service request are determined according to the service request frequency slot number, and a candidate route set is generated.
[0094] Preferably, the candidate route set from the source node to the destination node generated by the service request is K={k}, and k=[1,K], the route frequency slot number of the service request on each candidate route is calculated based on a preset frequency slot number calculation formula, wherein the preset frequency slot number calculation formula is as follows:
[0095] ;
[0096] In the formula, represents the maximum frequency slot number required by the lowest modulation level in the optical link included in the route, is a guard frequency slot.
[0097] In an embodiment, the preference value of each route planning scheme of the service request is calculated based on a preset preference value calculation formula, wherein the preset preference value calculation formula is as follows:
[0098] ;
[0099] In the formula, Let t be the preference value for the t-th business request. Let be the mean of the learning reward samples corresponding to the t-th business request. As a compromise between exploration and utilization, The larger they are, the more they tend to explore. This represents the number of times the k-th route and f-th frequency slot have been selected in the t-th service request up to the current time slot, and when t=0, Initialize to 0, Initialize to 0.
[0100] Step 102: Select the first routing planning scheme corresponding to the largest preference value, simulate the transmission of the service request according to the first routing planning scheme, calculate the first route and frequency slot selection indicator variable and the first learning reward value of the service request, and record the current first simulated transmission count.
[0101] In one embodiment, since the twin network synchronizes available routes and frequency slot information with the physical network when each service request arrives, the maximum number of route planning scheme simulations in the twin network layer is set to I, i.e., the simulation number threshold is set to I, and the actual number of simulations is i. The route and frequency slot with the highest preference value are selected to simulate the transmission process in the twin network.
[0102] In one embodiment, before simulating the transmission of the service request according to the first routing planning scheme, the routing and frequency slot selection indicator variables are first set to... ,when Indicates a business request Select candidate routes From the first Starting from a frequency gap Transmission is performed using consecutive frequency slots; otherwise... Indicates a business request Unable to select candidate routes From the first Starting from a frequency gap Transmission is carried out using consecutive frequency slots.
[0103] Due to the indivisibility of a single service request, meaning that a single service request can only select one frequency slot block on one route for transmission at a time, It should meet the following requirements: .
[0104] In an embodiment, the first route planning scheme is used to simulate transmission of the service request, and a first route and frequency slot selection indicator of the service request is calculated, specifically including: a first preference value corresponding to the first route planning scheme is obtained, the first preference value is substituted into a preset first route and frequency slot selection indicator calculation formula, and a first route and frequency slot selection indicator of the first route planning scheme is calculated.
[0105] ;
[0106] In the formula, is the first route and frequency slot selection indicator corresponding to the tth service request.
[0107] In an embodiment, due to concurrent access of a large number of service requests, a service request may be transmitted to a node according to a preset route and frequency slot, but the selected frequency slot is occupied by other service requests, resulting in transmission failure of the service request. Therefore, in the embodiment, a successful transmission indicator is also set as , wherein indicates that the service request can reach the destination node through the preset route and frequency slot, otherwise .
[0108] In an embodiment, before the service request is simulated to be transmitted according to the first route planning scheme, a service request transmission completion time is also set as , and a transmission power consumption corresponding to the transmission completion or transmission failure is , wherein a transmission power consumption calculation formula of the service request is as follows:
[0109] .
[0110] In an embodiment, before the service request is simulated to be transmitted according to the first route planning scheme, a frequency slot availability indicator is also set as , when indicates that the candidate route has continuous frequency slots from the first frequency slot of the candidate route in an idle state, otherwise .
[0111] In an embodiment, based on the transmission power consumption of the service request, the successful transmission indicator, the route and frequency slot selection indicator, the frequency slot availability indicator, the bandwidth requirement and the duration, an optimization objective function is set to autonomously find a proper route planning scheme, wherein the optimization objective function is as follows:
[0112]
[0113] wherein, denotes a service request data transmission rate; denotes a service request data transmission latency; then denotes a service request data transmission amount per unit energy consumption. representing the availability constraints of routes and frequency slots, representing the value range constraints of the route and frequency slot selection indicator variables, denotes the indivisibility constraint of a single service request.
[0114] In an embodiment, the service request is simulated to be transmitted according to the first route planning scheme, and a first learning reward value of the service request is calculated, specifically including: when the service request is simulated to be transmitted according to the first route planning scheme, the data transmission rate, transmission power consumption and successful transmission indicator variable of the service request are obtained, the data transmission rate, transmission power consumption and successful transmission indicator variable are substituted into a preset learning reward value calculation formula, and a first learning reward value of the first route planning scheme is calculated, wherein the preset learning reward value calculation formula is as follows:
[0115] ;
[0116] wherein, is a first learning reward value corresponding to the tth service request, is a data transmission rate, is transmission power consumption corresponding to the tth service request, is a successful transmission indicator variable corresponding to the tth service request, .
[0117] Preferably, the learning reward value definition is consistent with the optimization objective function.
[0118] In an embodiment, before the service request is simulated to be transmitted according to the first route planning scheme, the selection number of the kth route and the fth frequency slot until the current time slot is also set.
[0119] In an embodiment, before the service request is simulated to be transmitted according to the first route planning scheme, the route and frequency slot selection indicator variable, the selection number of the kth route and the fth frequency slot until the current time slot and the learning reward value are also initialized; specifically, when t=0, the route and frequency slot selection indicator variable , the selection number of the kth route and the fth frequency slot until the current time slot and the learning reward value are set to 0.
[0120] Step 103: calculating a first learning reward sample mean based on the first route and frequency slot selection indicator variable, and updating a preference value of the first route planning scheme based on the first learning reward sample mean.
[0121] In an embodiment, a selection number of the kth route and the fth frequency slot until the current time slot is obtained, the selection number is input into a preset selection number updating formula, a first selection number of the kth route and the fth frequency slot corresponding to the first route planning scheme until the current time slot is obtained, and the selection number is updated to the first selection number.
[0122] ;
[0123] In the formula, is the first selection number of the kth route and the fth frequency slot until the current time slot corresponding to the tth service request, is a selection number of the kth route and the fth frequency slot until the current time slot corresponding to the (t-1)th service request, is the first route and frequency slot selection indicator variable corresponding to the tth service request.
[0124] In an embodiment, the first route and frequency slot selection indicator variable and the first learning reward value of the first route planning scheme, and the updated selection number are obtained, the first route and frequency slot selection indicator variable, the first learning reward value and the selection number are substituted into a preset learning reward sample mean calculation formula, a first learning reward sample mean is obtained, and the learning reward sample mean is updated to the first learning reward sample mean, wherein the preset learning reward sample mean calculation formula is as follows:
[0125] ;
[0126] In the formula, is the first learning reward sample mean corresponding to the tth service request, is a learning reward sample mean corresponding to the (t-1)th service request, is the first route and frequency slot selection indicator variable corresponding to the tth service request, is the first learning reward value corresponding to the tth service request, is the first selection number of the kth route and the fth frequency slot until the current time slot corresponding to the tth service request.
[0127] In an embodiment, the updated selection times of the kth route and the fth frequency slot up to the current time slot and the updated learning reward sample mean are substituted into a preset preference value calculation formula to calculate and update the preference value of the first routing scheme. By updating the preference value list, the routing and spectrum planning strategy can be quickly converged.
[0128] In step 104, based on the first learning reward value, when the successful transmission indication variable is 1, the simulation transmission of the service request is stopped, the first routing scheme is taken as the optimal routing scheme corresponding to the service request, and the optimal routing scheme is sent to the physical network layer so that the physical network layer transmits the service request according to the optimal routing scheme.
[0129] In an embodiment, based on the first learning reward value, when the successful transmission indication variable is not 1 and the current first simulation transmission times is not equal to the preset simulation transmission threshold, a second routing scheme corresponding to the maximum preference value is selected again, and the second routing scheme is simulated until the successful transmission indication variable is 1 or the current first simulation transmission times is equal to the preset simulation transmission threshold, and the second routing scheme is taken as the optimal routing scheme corresponding to the service request.
[0130] Specifically, in the twin network layer, each service request is simulated by at most routing schemes to determine the final optimal routing scheme. If , the simulation process is terminated, the routing scheme corresponding to the current simulation transmission is taken as the optimal routing scheme, and the optimal routing scheme is sent to the physical network layer, otherwise, the current simulation times is increased by , and the above simulation selection process is repeated until .
[0131] In this embodiment, different energy consumption feedbacks are obtained by selecting different actions for each service request, which are used as experience values for future decision-making, thereby realizing the joint optimization of service request data transmission delay and energy consumption, and learning and converging to the optimal routing and frequency slot planning strategy under the condition of not mastering global information.
[0132] In summary, the application provides a kind of autonomous routing planning method of power elastic optical network, by combining digital twin technology with confidence upper limit algorithm, the available route and frequency gap information of twin network layer and physical network layer are synchronized, the convergence speed of simulation experiment of routing planning scheme in twin network layer is accelerated, the routing planning scheme with maximum preference value is dynamically selected, the problems of poor autonomy of traditional machine learning method and poor service guarantee caused by a large number of service request experiments can be effectively solved;And by estimating the confidence upper limit of the preference value of each routing planning scheme, the optimal routing planning scheme can be found under the condition that network state information dynamically changes, network routing and spectrum autonomous planning are realized, and large-scale concurrent access of power service is effectively adapted.
[0133] Embodiment 2
[0134] Reference Figure 2 , Figure 2 is a kind of structure schematic diagram of the embodiment of the autonomous routing planning device of power elastic optical network provided by the application, as Figure 2 The device includes routing planning scheme preference value calculation module 201, service request simulation transmission module 202, routing planning scheme preference value updating module 203 and optimal routing planning scheme determination module 204, as follows:
[0135] The routing planning scheme preference value calculation module 201 is used to obtain all available route data and all available frequency gap data of physical network layer when detecting that service request arrives, generate all routing planning schemes corresponding to the service request based on the all available route data and the all available frequency gap data, and calculate the preference value of each routing planning scheme for the service request.
[0136] The service request simulation transmission module 202 is used to select the first routing planning scheme corresponding to the maximum preference value, simulate transmission according to the first routing planning scheme for the service request, calculate the first route and frequency gap selection indication variable and the first learning reward value of the service request, and record the current first simulation transmission frequency.
[0137] The routing planning scheme preference value updating module 203 is used to calculate the first learning reward sample mean based on the first route and frequency gap selection indication variable, and update the preference value of the first routing planning scheme based on the first learning reward sample mean.
[0138] The optimal routing scheme determination module 204 is configured to determine, based on the first learning reward value, that the successful transmission indication variable is 1, stop simulating transmission of the service request, take the first routing scheme as the optimal routing scheme corresponding to the service request, and send the optimal routing scheme to the physical network layer, so that the physical network layer transmits the service request according to the optimal routing scheme.
[0139] In an embodiment, the optimal routing scheme determination module 204 is further configured to, based on the first learning reward value, determine that the successful transmission indication variable is not 1, and the current first simulation transmission number is not equal to the preset simulation transmission threshold, reselect a second routing scheme corresponding to the maximum preference value, and simulate transmission of the second routing scheme until the successful transmission indication variable is 1, or the current first simulation transmission number is equal to the preset simulation transmission threshold, take the second routing scheme as the optimal routing scheme corresponding to the service request.
[0140] In an embodiment, the routing scheme preference value calculation module 201 is configured to generate all routing schemes corresponding to the service request based on the all available routing data and the all available frequency slot data, specifically including: calculating a service request frequency slot number of the service request, determining all candidate routes of the service request according to the service request frequency slot number, and calculating a routing frequency slot number of the service request on each candidate route; and generating all routing schemes corresponding to the service request according to the all candidate routes and the routing frequency slot number.
[0141] In an embodiment, the routing scheme preference value calculation module 201 is configured to calculate the preference value of each routing scheme of the service request based on a preset preference value calculation formula, wherein the preset preference value calculation formula is as follows:
[0142] ;
[0143] In the formula, preference value of the tth service request, is a learning reward sample mean corresponding to the tth service request, is a compromise coefficient between exploration and utilization, is the selection number of the kth route and the fth frequency slot of the tth service request until the current time slot.
[0144] In an embodiment, the service request simulation transmission module 202 is configured to simulate transmission of the service request according to the first routing plan, and calculate the first routing and frequency slot selection indicator variable and the first learning reward value of the service request, specifically including: obtaining the first preference value corresponding to the first routing plan, substituting the first preference value into a preset first routing and frequency slot selection indicator variable calculation formula, and calculating the first routing and frequency slot selection indicator variable of the first routing plan, wherein the preset first routing and frequency slot selection indicator variable calculation formula is as follows:
[0145] ;
[0146] wherein, is the first routing and frequency slot selection indicator variable corresponding to the tth service request.
[0147] In an embodiment, when the service request simulation transmission module 202 simulates transmission of the service request according to the first routing plan, the data transmission rate, the transmission power consumption and the successful transmission indicator variable of the service request are obtained, the data transmission rate, the transmission power consumption and the successful transmission indicator variable are substituted into a preset learning reward value calculation formula, and the first learning reward value of the first routing plan is calculated, wherein the preset learning reward value calculation formula is as follows:
[0148] ;
[0149] wherein, is the first learning reward value corresponding to the tth service request, is the data transmission rate, is the transmission power consumption corresponding to the tth service request, is the successful transmission indicator variable corresponding to the tth service request, .
[0150] In an embodiment, the routing plan preference value updating module 203 is configured to calculate the first learning reward sample mean based on the first routing and frequency slot selection indicator variable, and update the preference value of the first routing plan based on the first learning reward sample mean, specifically including: obtaining the selection times of the kth routing and the fth frequency slot until the current time slot, inputting the selection times into a preset selection times updating formula, obtaining the first selection times of the kth routing and the fth frequency slot corresponding to the first routing plan until the current time slot, and updating the selection times to the first selection times;
[0151] ;
[0152] wherein, the first selection number of the kth route and the fth frequency slot corresponding to the tth service request until the current time slot, the selection number of the kth route and the fth frequency slot corresponding to the t-1th service request until the current time slot, the first route and frequency slot selection indicator variable corresponding to the tth service request.
[0153] In an embodiment, the route planning scheme preference value updating module 203 is configured to obtain the first route and frequency slot selection indicator variable, the first learning reward value and the updated selection number of the first route planning scheme, substitute the first route and frequency slot selection indicator variable, the first learning reward value and the selection number into a preset learning reward sample mean calculation formula, obtain a first learning reward sample mean, and update the learning reward sample mean to the first learning reward sample mean. The preset learning reward sample mean calculation formula is as follows:
[0154] ;
[0155] In the formula, the first learning reward sample mean corresponding to the tth service request, the learning reward sample mean corresponding to the t-1th service request, the first route and frequency slot selection indicator variable corresponding to the tth service request, the first learning reward value corresponding to the tth service request, the first selection number of the kth route and the fth frequency slot corresponding to the tth service request until the current time slot;
[0156] In an embodiment, the route planning scheme preference value updating module 203 is configured to substitute the updated selection number of the kth route and the fth frequency slot until the current time slot and the updated learning reward sample mean into a preset preference value calculation formula, calculate and update the preference value of the first route planning scheme.
[0157] In summary, the application provides an autonomous routing planning device for a power elastic optical network, which combines digital twin technology with a confidence upper limit algorithm, synchronizes the available routing and frequency gap information of the network layer and the physical network layer, accelerates the convergence speed of learning through simulation experiments on routing planning schemes in the twin network layer, dynamically selects the routing planning scheme with the largest preference value, effectively solves the problems of poor autonomy of traditional machine learning methods and poor service guarantee caused by a large number of service request experiments, estimates the confidence upper limit of the preference value of each routing planning scheme, finds the optimal routing planning scheme under the condition of dynamic changes in network state information, realizes autonomous planning of network routing and spectrum, and effectively adapts to large-scale concurrent access of power services.
[0158] Embodiment 3
[0159] Reference Figure 3 , Figure 3 is a structural schematic diagram of an embodiment of the power elastic optical network provided by the application, as Figure 2 shown, the network includes a physical network layer 301, a twin network layer 302 and a service application layer 303, and specifically as follows:
[0160] The twin network layer 302 includes a northbound interface and a southbound interface.
[0161] The twin network layer 302 receives the demand input by the service application layer 303 through the northbound interface.
[0162] The twin network layer 302 obtains the network state data in the physical network layer 301 through the southbound interface, and after determining the optimal routing planning scheme, transmits the optimal routing planning scheme to the physical network layer 301 through the southbound interface.
[0163] The twin network layer 302 executes the autonomous routing planning method of the power elastic optical network as described in the above embodiment.
[0164] In an embodiment, the physical network layer 301 includes power elastic optical network physical entities such as routers and fiber links, and various network elements interact with the twin network layer 302 through the southbound interface of the twin network layer 302 to exchange network data and network control information, and realize reliable bearing of specific services of the service application layer 303.
[0165] In an embodiment, the service application layer 303 inputs the demand to the twin network layer 302 through the northbound interface of the twin network layer 302, and supports the deployment of power services in the twin network layer 302.
[0166] In an embodiment, the twin network layer 302 builds a digital network that can reflect the network state of the physical network layer 301 in real time through data interaction with the physical network layer 301, performs dynamic autonomous planning of network routing and spectrum, and learns the optimal routing planning scheme according to the current network state. After sufficient verification, the optimal routing planning scheme is issued to the physical network layer 301 through the southbound interface.
[0167] To sum up, the power elastic optical network provided by the application realizes smaller granularity of communication bandwidth scheduling, improves the resource utilization rate of the network, and supports autonomous routing planning of the network to adapt to differentiated power service requirements, by setting the power elastic optical network architecture based on digital twin assistance.
[0168] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0169] It should be noted that the embodiments of the autonomous routing planning device of the power elastic optical network described above are only illustrative, and the modules described as separate components can or can not be physically separated, and the components displayed as modules can or can not be physical units, that is, they can be located in one place or distributed on multiple network units. According to actual needs, part or all of the modules can be selected to achieve the purpose of the embodiment.
[0170] On the basis of the foregoing embodiments of the autonomous routing planning method of the power elastic optical network, another embodiment of the application provides an autonomous routing planning terminal device of a power elastic optical network. The autonomous routing planning terminal device of the power elastic optical network comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the power elastic optical network autonomous routing planning method of any one of the embodiments of the application is implemented.
[0171] For example, in this embodiment, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the application. The one or more modules can be a series of computer program instruction segments that can complete a specific function, which are used to describe the execution process of the computer program in the autonomous routing planning terminal device of the power elastic optical network.
[0172] The autonomous routing planning terminal device of the power elastic optical network can be a desktop computer, a notebook computer, a palm computer, a cloud server, and the like. The autonomous routing planning terminal device of the power elastic optical network can include, but is not limited to, a processor and a memory.
[0173] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, and the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor and the like. The processor is a control center of the autonomous routing planning terminal device of the power elastic optical network, and connects various parts of the autonomous routing planning terminal device of the power elastic optical network through various interfaces and lines.
[0174] The memory can be used to store the computer program and / or the module, and the processor realizes various functions of the autonomous routing planning terminal device of the power elastic optical network by running or executing the computer program and / or the module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area. The program storage area can store an operating system, at least one application required by a function, and the like; and the data storage area can store data created according to the use of the mobile phone, and the like. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, for example, a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state memory devices.
[0175] On the basis of the above-mentioned embodiments of the power elastic optical network autonomous routing planning method, another embodiment of the present application provides a storage medium including a stored computer program, wherein when the computer program runs, the device where the storage medium is located is controlled to execute the power elastic optical network autonomous routing planning method of any one of the embodiments of the present application.
[0176] In this embodiment, the storage medium described above is a computer readable storage medium, the computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.
[0177] To sum up, the application provides an autonomous routing planning method, device and system of an electric power elastic optical network. When detecting that a service request arrives, all available routing data and all available frequency slot data of a physical network layer are acquired, all routing planning schemes corresponding to the service request are generated, the preference value of each routing planning scheme is calculated, the first routing planning scheme corresponding to the maximum preference value is used to simulate transmission of the service request, the first routing and frequency slot selection instruction variable and the first learning reward value of the service request are calculated, the preference value of the first routing planning scheme is updated based on the first routing and frequency slot selection instruction variable and the mean value of the first learning reward sample, and when the successful transmission instruction variable is 1, the first routing planning scheme is determined as the optimal routing planning scheme corresponding to the service request and is issued to the physical network layer. Compared with the prior art, the technical scheme of the application can realize autonomous planning of network routing and spectrum.
[0178] The above only describes the preferred embodiments of the application. It should be noted that for those skilled in the art, without departing from the technical principles of the application, some improvements and replacements can be made, which should also be considered as the protection scope of the application.
Claims
1. An autonomous routing planning method for a power-elastic optical network, characterized in that, include: When a service request is detected, all available routing data and all available frequency slot data of the physical network layer are obtained. Based on all available routing data and all available frequency slot data, all routing planning schemes corresponding to the service request are generated, and the preference value of the service request for each routing planning scheme is calculated. Select the first routing planning scheme corresponding to the largest preference value, simulate the transmission of the service request according to the first routing planning scheme, calculate the first route and frequency slot selection indicator variable and the first learning reward value of the service request, and record the current first simulated transmission count, including: obtaining the first preference value corresponding to the first routing planning scheme, substituting the first preference value into the preset calculation formula of the first route and frequency slot selection indicator variable, and calculating the first route and frequency slot selection indicator variable of the first routing planning scheme; Obtain the data transmission rate, transmission power consumption, and successful transmission indication variable of the service request, and substitute the data transmission rate, transmission power consumption, and successful transmission indication variable into the preset learning reward value calculation formula to calculate the first learning reward value of the first routing planning scheme. Based on the first route and frequency slot selection indicator variables, calculate the mean of the first learning reward sample, and update the preference value of the first route planning scheme based on the mean of the first learning reward sample. Based on the first learning reward value, when the successful transmission indicator variable is determined to be 1, the simulated transmission of the service request is stopped, the first routing plan is taken as the optimal routing plan corresponding to the service request, and the optimal routing plan is sent to the physical network layer so that the physical network layer can transmit the service request according to the optimal routing plan.
2. The autonomous routing planning method for a power-elastic optical network as described in claim 1, characterized in that, Also includes: Based on the first learning reward value, when the successful transmission indicator variable is not 1 and the current first simulated transmission count is not equal to the preset simulated transmission threshold, the second routing plan scheme corresponding to the largest preference value is reselected, and the second routing plan scheme is simulated for transmission until the successful transmission indicator variable is 1 or the current first simulated transmission count is equal to the preset simulated transmission threshold. Then, the second routing plan scheme is taken as the optimal routing plan scheme corresponding to the service request.
3. The autonomous routing planning method for a power-elastic optical network as described in claim 1, characterized in that, Based on all available routing data and all available frequency slot data, all routing planning schemes corresponding to the service request are generated, specifically including: Calculate the number of service request slots for the service request, determine all candidate routes for the service request based on the number of service request slots, and calculate the number of route slots for the service request on each candidate route; Based on all candidate routes and the number of route slots, generate all routing planning schemes corresponding to the service request.
4. The autonomous routing planning method for a power-elastic optical network as described in claim 1, characterized in that, Based on a preset preference value calculation formula, the preference value of the service request for each routing plan is calculated, wherein the preset preference value calculation formula is as follows: ; In the formula, Let t be the preference value for the t-th business request. Let be the mean of the learning reward samples corresponding to the t-th business request. As a compromise between exploration and utilization, This represents the number of times the k-th route and f-th frequency slot have been selected in the t-th service request up to the current time slot.
5. The autonomous routing planning method for a power-elastic optical network as described in claim 4, characterized in that, The preset formulas for calculating the first route and frequency slot selection indicator variables and the preset formulas for calculating the learning reward value specifically include: The preset formula for calculating the first route and frequency slot selection indicator variable is as follows: ; In the formula, Select indicator variables for the first route and frequency slot corresponding to the t-th service request; The preset formula for calculating the learning reward value is as follows: ; In the formula, Let t be the first learning reward value corresponding to the t-th business request. For data transmission rate, Let be the transmission power consumption corresponding to the t-th service request. This is the successful transmission indication variable corresponding to the t-th service request. .
6. The autonomous routing planning method for a power-elastic optical network as described in claim 5, characterized in that, Based on the first route and frequency slot selection indicator variables, the mean of the first learning reward samples is calculated, and based on the mean of the first learning reward samples, the preference value of the first route planning scheme is updated, specifically including: Obtain the selection count of the k-th route and the f-th frequency slot up to the current time slot, input the selection count into a preset selection count update formula, obtain the first selection count of the k-th route and the f-th frequency slot corresponding to the first routing planning scheme up to the current time slot, and update the selection count to the first selection count; ; In the formula, This represents the number of times the k-th route and the f-th frequency slot are selected for the t-th service request up to the current time slot. This represents the number of times the k-th route and f-th frequency slot have been selected up to the current time slot for the (t-1)-th service request. Select indicator variables for the first route and frequency slot corresponding to the t-th service request; Obtain the first route and frequency slot selection indicator variable and the first learning reward value of the first route planning scheme, as well as the updated selection count. Substitute the first route and frequency slot selection indicator variable, the first learning reward value, and the selection count into a preset learning reward sample mean calculation formula to obtain the first learning reward sample mean. Update the learning reward sample mean to the first learning reward sample mean. The preset learning reward sample mean calculation formula is as follows: ; In the formula, Let be the mean of the first learning reward sample corresponding to the t-th business request. Let be the mean of the learning reward samples corresponding to the (t-1)th business request. Select indicator variables for the first route and frequency slot corresponding to the t-th service request. Let t be the first learning reward value corresponding to the t-th business request. This represents the number of times the k-th route in the current time slot and the first selection count in the f-th frequency slot corresponding to the t-th service request; The updated selection counts of the kth route up to the current time slot and the fth frequency slot, along with the updated average of the learning reward samples, are substituted into the preset preference value calculation formula to calculate and update the preference value of the first route planning scheme.
7. An autonomous routing planning device for a power-resilient optical network, characterized in that, include: The system includes a routing planning scheme preference value calculation module, a service request simulation transmission module, a routing planning scheme preference value update module, and an optimal routing planning scheme determination module. The routing planning scheme preference value calculation module is used to obtain all available routing data and all available frequency slot data of the physical network layer when a service request is detected, generate all routing planning schemes corresponding to the service request based on the all available routing data and all available frequency slot data, and calculate the preference value of the service request for each routing planning scheme. The service request simulation transmission module is used to select the first routing planning scheme corresponding to the largest preference value, simulate the transmission of the service request according to the first routing planning scheme, and calculate the first route and frequency slot selection indicator variable and the first learning reward value of the service request, including: obtaining the first preference value corresponding to the first routing planning scheme, substituting the first preference value into the preset calculation formula of the first route and frequency slot selection indicator variable, and calculating the first route and frequency slot selection indicator variable of the first routing planning scheme; Obtain the data transmission rate, transmission power consumption, and successful transmission indication variable of the service request; substitute the data transmission rate, transmission power consumption, and successful transmission indication variable into the preset learning reward value calculation formula to calculate the first learning reward value of the first routing planning scheme; and record the current first simulated transmission count. The routing planning scheme preference value update module is used to calculate the mean of the first learning reward sample based on the first route and the frequency slot selection indicator variable, and update the preference value of the first routing planning scheme based on the mean of the first learning reward sample. The optimal routing plan determination module is used to stop simulating the transmission of the service request when the successful transmission indication variable is 1 based on the first learning reward value, take the first routing plan as the optimal routing plan corresponding to the service request, and send the optimal routing plan to the physical network layer so that the physical network layer transmits the service request according to the optimal routing plan.
8. A power-elastic optical network, characterized in that, include: Physical network layer, twin network layer, and business application layer; The twin network layer includes a northbound interface and a southbound interface; The twin network layer receives the requirements input by the service application layer through the northbound interface; The twin network layer obtains network status data from the physical network layer through the southbound interface, and after determining the optimal routing plan, sends the optimal routing plan to the physical network layer through the southbound interface. The twin network layer is used to execute the autonomous routing planning method for power elastic optical networks as described in any one of claims 1-6.
9. A terminal device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the autonomous routing planning method for a power resilient optical network as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device containing the computer-readable storage medium to perform the autonomous routing planning method for a power-resilient optical network as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Routing and spectrum resource allocation method and system for resource awareness in elastic optical path network
CN103051547A
Optical network path planning method and device
CN114885236A