Tidal flow network optical path planning method, manager, controller and system based on meta-learning
By adopting a meta-learning-based optical path planning method in the optical network under tidal service traffic, the problem of high retraining time and energy consumption of the agent is solved, and more efficient optical network control and network service quality improvement is achieved.
Patent Information
- Application Number
- CN202510156395.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-02-12
AI Technical Summary
Under tidal service flow, the existing optical network optical path planning algorithm is difficult to effectively reduce the time and energy consumption of agents for retraining, resulting in a low level of optical network management and control.
The optical path planning method of tidal flow network based on meta-learning is adopted to retrain the reinforcement learning agent according to the traffic distribution data in the current time step, and the optimal optical path configuration strategy is obtained, and the initial configuration strategy is updated using the meta-learning algorithm to reduce the training needs of future time steps.
This method can significantly reduce the time and energy consumption of retraining of the agent, improve the level of optical network control, provide effective solutions for traffic diversion of tidal services, and improve network service quality.
Smart Images

Figure CN120186501A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of software-defined optical networks (SDON), and particularly to a method, a manager, a controller, and a system for optical path planning of a tidal traffic network based on meta-learning. Background Art
[0002] Communication networks serve humans. Due to the periodic work and rest of humans, the demand for communication services also shows periodic changes. For example, the demand for map services increases during the morning and evening rush hours, and the demand for takeaway services increases during meal times. The network load also shows relatively regular periodic tidal changes in units of days, weeks, months, and years, thus forming a tidal traffic network. Optical networks need to provide physical resources for these tidal service communications.
[0003] Currently, optical path planning in software-defined optical networks (SDON, abbreviated as optical networks) is an important issue, and similar RSA (routing and resource allocation) problems have been proven to be non-deterministic polynomial (NP-hard) problems, and it is a consensus in the industry that their solution is difficult and the process is complex. Optical network path planning under tidal service traffic is more complex than traditional path planning. Optical path planning also faces challenges such as a sharp increase in traffic volume and distribution drift, and the disadvantages of traditional optical path planning algorithms are obvious. Therefore, how to design an optical network optical path planning algorithm based on serving tidal services to solve the tidal problem of services is crucial for improving network service quality.
[0004] Existing optical network optical path planning algorithms include greedy methods, linear programming-based methods, heuristic-based methods, and reinforcement learning-based methods. Among them, the reinforcement learning-based method is the most popular solution in the current academic field and has relatively high optimization degree and convergence speed compared with other algorithms. However, the disadvantage of the reinforcement learning-based method is that when the service distribution changes, the decision model needs to be retrained, so it will consume relatively high costs such as time and energy consumption.
[0005] Based on this, how to reduce the costs such as time and energy consumption of agent retraining under tidal traffic services to improve the optical network management and control level is an urgent problem to be solved currently. Summary of the Invention
[0006] In view of this, embodiments of the present application provide a method, a manager, a controller, and a system for optical path planning of a tidal traffic network based on meta-learning to eliminate or improve one or more defects existing in the prior art.
[0007] One aspect of the present application provides a method for optical path planning of a tidal traffic network based on meta-learning, including:
[0008] According to the traffic distribution data of the tidal traffic network at the current time step, retrain the reinforcement learning agent for the initial optical path configuration strategy corresponding to the current time step to obtain the optimal optical path configuration strategy corresponding to the current time step; and, within the current time step, update the initial optical path configuration strategy corresponding to the current time step based on the meta-learning algorithm to obtain the initial optical path configuration strategy corresponding to the next time step;
[0009] Based on the optimal optical path configuration strategy, perform optical path reconstruction on the optical network layer corresponding to the tidal traffic network to correspondingly update the topology of the IP layer corresponding to the tidal traffic network.
[0010] In some embodiments of the present application, if the current time step is the first time step within a preset time period, the initial optical path configuration strategy corresponding to the current time step is randomly generated in advance;
[0011] If the current time step is a non-first time step within a preset time period, the initial optical path configuration strategy corresponding to the current time step is obtained by updating the optimal optical path configuration strategy corresponding to the previous time step based on the meta-learning algorithm within the previous time step.
[0012] In some embodiments of the present application, the updating the initial optical path configuration strategy corresponding to the current time step based on the meta-learning algorithm within the current time step to obtain the initial optical path configuration strategy corresponding to the next time step includes:
[0013] Within the current time step, execute a preset meta-learning algorithm according to the initial optical path configuration strategy corresponding to the current time step, the trajectory data set generated through reinforcement learning interaction during the retraining process of the reinforcement learning agent in the current time step, a preset reinforcement learning learning rate, and a meta-learning learning rate, to obtain the updated result data of the initial optical path configuration strategy corresponding to the current time step, and use this updated result data as the initial optical path configuration strategy corresponding to the next time step.
[0014] In some embodiments of the present application, the performing optical path reconstruction on the optical network layer corresponding to the tidal traffic network based on the optimal optical path configuration strategy to correspondingly update the topology of the IP layer corresponding to the tidal traffic network includes:
[0015] Within the current time step, asynchronously send the optimal optical path configuration strategy corresponding to the current time step to a controller, so that the controller performs optical path reconstruction on the optical network layer corresponding to the tidal traffic network based on the optimal optical path configuration strategy corresponding to the current time step to correspondingly update the topology of the IP layer corresponding to the tidal traffic network.
[0016] In some embodiments of the present application, the controller is used to perform the following:
[0017] Receive the optimal optical path configuration strategy corresponding to the current time step;
[0018] If it is determined that the optimal optical path configuration strategy corresponding to the previous time step is stored locally, then according to the comparison result between the optimal optical path configuration strategy corresponding to the previous time step and the optimal optical path configuration strategy corresponding to the current time step, among the various optical cross-connect devices in the optical network layer corresponding to the tidal flow network, determine the optical cross-connect device to be reconfigured for the optical path currently as the target device, and extract the optical cross-connect reconfiguration configuration data for each of the target devices respectively from the optimal optical path configuration strategy corresponding to the current time step;
[0019] Send the corresponding optical cross-connect reconfiguration configuration data to each of the target devices respectively, so that each of the target devices respectively performs optical cross-connect reconfiguration according to the received optical cross-connect reconfiguration configuration data and returns the corresponding optical cross-connect reconfiguration completion information;
[0020] Receive the optical cross-connect reconfiguration completion information respectively returned by each of the target devices. If it is determined that all the optical cross-connect reconfiguration completion information corresponding to each of the target devices is received, then send out the global reconfiguration completion information.
[0021] In some embodiments of the present application, the optical path planning method for the tidal flow network based on meta-learning further includes:
[0022] If the global reconfiguration completion information sent by the controller is received and the initial optical path configuration strategy corresponding to the next time step has been obtained currently, then it is determined that the optical path planning for the tidal flow network corresponding to the current time step has been completed.
[0023] In some embodiments of the present application, the traffic distribution data includes the traffic matrix of the tidal flow network, wherein each element in the traffic matrix is respectively used to represent the traffic between two nodes in the tidal flow network, and the communication frequencies corresponding to each of the elements are represented by different colors or different numerical parameters of the same color, wherein the numerical parameters include at least one of a hue value, a saturation value, and a lightness value;
[0024] Correspondingly, before re-training the reinforcement learning agent for the initial optical path configuration strategy corresponding to the current time step according to the traffic distribution data of the tidal flow network at the current time step, it further includes:
[0025] Obtain the current traffic matrix of the tidal flow network.
[0026] Another aspect of the present application provides an optical path planning device for a tidal flow network based on meta-learning, including:
[0027] A reinforcement learning and meta-learning parallel module, which is used to retrain the reinforcement learning agent for the initial optical path configuration policy corresponding to the current time step according to the traffic distribution data of the tidal traffic network at the current time step, so as to obtain the optimal optical path configuration policy corresponding to the current time step; and, within the current time step, update the initial optical path configuration policy corresponding to the current time step based on the meta-learning algorithm to obtain the initial optical path configuration policy corresponding to the next time step;
[0028] An optical path reconstruction module, which is used to reconstruct the optical path of the optical network layer corresponding to the tidal traffic network based on the optimal optical path configuration policy to correspondingly update the topology structure of the IP layer corresponding to the tidal traffic network.
[0029] The third aspect of the present application provides an SDON manager, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the optical path planning method for the tidal traffic network based on meta-learning.
[0030] The fourth aspect of the present application provides an SDON controller, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The SDON controller is communicatively connected to an SDON manager. The SDON manager is used to execute the optical path planning method for the tidal traffic network based on meta-learning; and the SDON manager is used to asynchronously send the optimal optical path configuration policy corresponding to the current time step to the SDON controller within the current time step and receive the global reconfiguration completion information sent by the SDON controller;
[0031] When the processor executes the computer program, the following content is implemented:
[0032] Receive the optimal optical path configuration policy corresponding to the current time step;
[0033] If it is determined that the optimal optical path configuration policy corresponding to the previous time step is locally stored, then according to the comparison result between the optimal optical path configuration policy corresponding to the previous time step and the optimal optical path configuration policy corresponding to the current time step, in each optical cross-connect device in the optical network layer corresponding to the tidal traffic network, determine the optical cross-connect device to be reconstructed currently as the target device, and extract the optical cross-connect reconstruction configuration data for each of the target devices from the optimal optical path configuration policy corresponding to the current time step;
[0034] Send the corresponding optical cross-connect reconstruction configuration data to each of the target devices respectively, so that each of the target devices respectively performs optical cross-connect reconstruction according to the received optical cross-connect reconstruction configuration data and returns the corresponding optical cross-connect reconfiguration completion information;
[0035] Receive the optical cross-connection reconfiguration completion information respectively returned by each of the target devices. If it is determined that all the optical cross-connection reconfiguration completion information corresponding to the respective target devices has been received, then send out the global reconfiguration completion information.
[0036] The fifth aspect of this application provides an optical path planning system for a tidal flow network based on meta-learning, the SDON manager mentioned in the third aspect above, and the SDON controller mentioned in the fourth aspect above;
[0037] The SDON manager is used to obtain the traffic distribution data of the tidal flow network at the current time step from the traffic layer of the tidal flow network;
[0038] The SDON controller is communicatively connected to each optical cross-connection device in the optical network layer corresponding to the tidal flow network respectively.
[0039] The sixth aspect of this application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the optical path planning method for the tidal flow network based on meta-learning described above.
[0040] The seventh aspect of this application provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the optical path planning method for the tidal flow network based on meta-learning as described above.
[0041] The optical path planning method for the tidal flow network based on meta-learning provided by this application, according to the traffic distribution data of the tidal flow network at the current time step, re-trains the reinforcement learning agent for the initial optical path configuration strategy corresponding to the current time step to obtain the optimal optical path configuration strategy corresponding to the current time step; and, within the current time step, updates the initial optical path configuration strategy corresponding to the current time step based on the meta-learning algorithm to obtain the initial optical path configuration strategy corresponding to the next time step; based on the optimal optical path configuration strategy, performs optical path reconstruction on the optical network layer corresponding to the tidal flow network to correspondingly update the topology of the IP layer corresponding to the tidal flow network, which can reduce the costs such as the time and energy consumption of the agent's re-training, thereby improving the optical network management and control level, providing an effective solution for the traffic diversion of tidal services, and improving the network service quality.
[0042] The additional advantages, objectives, and features of this application will be partially elaborated in the following description, and will become partially obvious to those of ordinary skill in the art after studying the following text, or can be learned from the practice of this application. The objectives and other advantages of this application can be achieved and obtained through the structures specifically pointed out in the description and the drawings.
[0043] Those skilled in the art will understand that the objectives and advantages achievable with the present application are not limited to those specifically described above, and the above and other objectives achievable with the present application will be more clearly understood from the following detailed description. Description of the Drawings
[0044] The drawings described herein are used to provide a further understanding of the present application, form a part of the present application, and do not limit the present application. The components in the drawings are not drawn to scale, but are only for showing the principles of the present application. For the convenience of showing and describing some parts of the present application, the corresponding parts in the drawings may be enlarged, that is, they may become larger relative to other components in the exemplary device actually manufactured according to the present application. In the drawings:
[0045] Figure 1 Schematic diagrams for examples of existing greedy methods, linear programming-based methods, heuristic-based methods, and reinforcement learning-based methods.
[0046] Figure 2 Schematic diagram of the first process of the meta-learning-based tidal flow network optical path planning method in an embodiment of the present application.
[0047] Figure 3 Schematic diagram of the second process of the meta-learning-based tidal flow network optical path planning method in an embodiment of the present application.
[0048] Figure 4 Schematic diagram of the logical solution process of formulas (1) and (2) in an embodiment of the present application.
[0049] Figure 5 Schematic diagram of the parallel execution process of reinforcement learning and meta-learning for steps 110 and 120 in an example of the present application.
[0050] Figure 6 Schematic diagram of the process executed by the SDON controller in an embodiment of the present application.
[0051] Figure 7 Schematic diagram of the structure of the SDON manager in an embodiment of the present application.
[0052] Figure 8 Schematic diagram of the execution logic of the meta-learning-based tidal flow network optical path planning system in an application example of the present application.
[0053] Figure 9 Schematic diagram of the interaction of the meta-learning-based tidal flow network optical path planning method executed by the meta-learning-based tidal flow network optical path planning system in an application example of the present application. Detailed Embodiments
[0054] To make the objectives, technical solutions and advantages of this application more clear and understandable, the following further elaborates on this application in combination with the embodiments and the accompanying drawings. Herein, the illustrative embodiments of this application and their descriptions are used to explain this application, but do not limit this application.
[0055] Herein, it should also be noted that in order to avoid obscuring this application with unnecessary details, only the structures and / or processing steps closely related to the solution of this application are shown in the drawings, while other details less related to this application are omitted.
[0056] It should be emphasized that the term "including / containing" as used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0057] Herein, it should also be noted that if not otherwise specified, the term "connection" in this text can refer not only to direct connection, but also to indirect connection with an intermediate.
[0058] In the following, embodiments of this application will be described with reference to the drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0059] See Figure 1 , existing optical network optical path planning algorithms include the greedy method, the linear programming-based method, the heuristic-based method, and the reinforcement learning-based method.
[0060] The greedy algorithm is mainly characterized by simple rules and rapid solution, and is a conventional solution for delay-sensitive communication systems. However, its main disadvantage is the low optimization level.
[0061] The greatest advantage of the linear programming scheme is the high optimization level, and its solution is the optimal solution. However, its main disadvantages are difficult modeling, difficult transformation of constraint conditions, and extremely slow solution speed under large-scale variables. Therefore, it is not a conventional solution.
[0062] The optimization level and convergence speed of the heuristic method are both between the greedy algorithm and the linear programming algorithm, and it is usually used as a compromise solution. However, it has the disadvantages of poor interpretability and unstable convergence speed.
[0063] As the most popular solution in the current academic field, for the qualitative analysis of the optimization performance and convergence speed of the reinforcement learning-based method, see Table 1, where the number of stars is positively correlated with the advantages; compared with other algorithms, the reinforcement learning-based method is characterized by a relatively high optimization level and convergence speed. The disadvantage is that when the service distribution changes, the decision model of the reinforcement learning-based method needs to be retrained.
[0064] Qualitative Analysis of Optimization Performance and Convergence Speed of Reinforcement Learning-Based Methods
[0065]
[0066]
[0067] Based on this, in order to solve the problems such as the high costs of the time and energy consumption for the retraining of the agent under the tidal flow service in the existing network optical path planning methods based on reinforcement learning, the embodiments of the present application respectively provide a meta-learning-based tidal flow network optical path planning method, a meta-learning-based tidal flow network optical path planning method device for executing the meta-learning-based tidal flow network optical path planning method, an SDON manager, an SDON controller, a meta-learning-based tidal flow network optical path planning system, a computer-readable storage medium, and a computer program product, which can apply meta-learning to the optical network optical path planning under the background of tidal services for traffic grooming.
[0068] Specific details are described in detail through the following embodiments.
[0069] Based on this, the embodiments of the present application provide a meta-learning-based tidal flow network optical path planning method that can be implemented by an SDON manager. Refer to Figure 2 , the meta-learning-based tidal flow network optical path planning method specifically includes the following content:
[0070] Step 100: According to the traffic distribution data of the tidal flow network at the current time step, retrain the reinforcement learning agent for the initial optical path configuration strategy corresponding to the current time step to obtain the optimal optical path configuration strategy corresponding to the current time step; and, within the current time step, update the initial optical path configuration strategy corresponding to the current time step based on the meta-learning algorithm to obtain the initial optical path configuration strategy corresponding to the next time step.
[0071] In one or more embodiments of the present application, the time step t generally refers to the time interval between one observation point and the next observation point in the sequence data in the time series, which can be represented by t, and it defines the sampling frequency and time scale of the time series data. For example, if the data is collected once a day, then each time step represents one day; if it is recorded once an hour, then each time step is one hour. Among them, the total length of the time series is represented by T, and T can be any time length or positive infinity, and can be specifically set according to actual application requirements.
[0072] It can be understood that a tidal flow network refers to a network in which the current flow exhibits obvious periodic fluctuations over time, similar to the periodic changes of ocean tides. This phenomenon is particularly evident in 5G networks, especially in areas such as universities, industrial parks, CBD business districts, and large residential areas.
[0073] In one or more embodiments of the present application, both the initial optical path configuration strategy and the optimal optical path configuration strategy include configuration data corresponding to each optical cross-connect device in the optical network layer of the current tidal flow network, and the configuration data includes the cross-connection relationships between each of the optical cross-connect devices and other optical cross-connect devices.
[0074] Among them, the optical cross-connect device OXC (optical cross-connect) is an important network unit in the optical wave network. Its function can be analogous to that of a switch in a time-division multiplexing network. It is mainly used to complete the cross-connection between multi-wavelength ring networks. As a node of a mesh optical network, its purpose is to achieve the automatic configuration, protection / restoration, and reconstruction of the optical wave network. Each optical cross-connect device in the tidal flow network can be abbreviated as OXCs.
[0075] In step 100, according to the traffic distribution data of the tidal flow network at the current time step, the retraining of the reinforcement learning agent for the initial optical path configuration strategy corresponding to the current time step can be performed using existing retraining methods for reinforcement learning agents. For example, at the beginning of each time step, when it is considered that the traffic distribution at the current time step is inconsistent with that of the previous step, that is, a distribution drift has occurred and the decision of the reinforcement learning agent fails and needs to be retrained. The SDON manager performs DRL for reinforcement learning training again. Using the current traffic matrix as the environmental feature and maximizing the throughput as the goal, a Markov decision process is modeled. After the reinforcement learning agent has undergone several rounds of training, it can output the optimal optical path configuration strategy.
[0076] It can be understood that the software-defined optical network SDON is a centralized management and control separation scheduling architecture. The optical path refers to an end-to-end determined wavelength optical channel configured by an optical switch in the optical network. After the optical path reconstruction of the optical network layer corresponding to the tidal flow network is performed based on the optimal optical path configuration strategy to correspondingly update the topology of the IP layer corresponding to the tidal flow network, the re-planning of the optical path is completed.
[0077] In step 100, within the current time step, the execution of updating the initial optical path configuration policy corresponding to the current time step based on the meta-learning algorithm to obtain the initial optical path configuration policy corresponding to the next time step and the execution of performing optical path reconstruction on the optical network layer corresponding to the tidal flow network based on the optimal optical path configuration policy to correspondingly update the topology of the IP layer corresponding to the tidal flow network can be executed synchronously or asynchronously. Specifically, the execution process of the meta-learning algorithm can be executed synchronously with the reinforcement learning process without waiting to obtain the optimal optical path configuration policy corresponding to the current time step before execution, thereby effectively improving the efficiency of the agent's retraining and reducing the energy consumption cost.
[0078] It can be understood that meta-learning, also known as "learning how to learn", is an important algorithm framework in the field of artificial intelligence. The goal of traditional machine learning methods is to learn the distribution model of data or a specific strategy, while meta-learning hopes to learn the distribution model of data distributions or learn the learning ability of a certain strategy, which is a learning model with a higher level of intention. Meta-learning has good performance in the fields of images and languages. It is usually used in the field of few-shot learning, using the learning results of the model in the multi-sample domain, retaining knowledge, and achieving rapid adaptation in the few-shot domain.
[0079] Step 200: Perform optical path reconstruction on the optical network layer corresponding to the tidal flow network based on the optimal optical path configuration policy to correspondingly update the topology of the IP layer corresponding to the tidal flow network.
[0080] From the above description, it can be seen that the optical path planning method for the tidal flow network based on meta-learning provided by the embodiments of the present application can reduce costs such as the time and energy consumption of the agent's retraining, improve the efficiency of the agent's retraining, thereby improving the optical network management and control level, providing an effective solution for the traffic diversion of tidal services, and improving the network service quality.
[0081] In order to further improve the efficiency and reliability of the retraining of the reinforcement learning agent in the optical path planning process of the tidal flow network based on meta-learning, in an optical path planning method for the tidal flow network based on meta-learning provided by the embodiments of the present application, if the current time step is the first time step within the preset time period, that is, t = 0, the initial optical path configuration policy corresponding to the current time step is randomly generated in advance; if the current time step is a non-first time step within the preset time period, the initial optical path configuration policy corresponding to the current time step is obtained by updating the optimal optical path configuration policy corresponding to the previous time step based on the meta-learning algorithm in the previous time step.
[0082] To further improve the effectiveness and reliability of updating the initial optical path configuration strategy corresponding to the current time step based on the meta-learning algorithm in the optical path planning process of the tidal flow network based on meta-learning, in a method for optical path planning of a tidal flow network based on meta-learning provided in an embodiment of the present application, refer to Figure 3 Step 100 in the method for optical path planning of the tidal flow network based on meta-learning specifically includes the following content:
[0083] Step 110: According to the traffic distribution data of the tidal flow network at the current time step, retrain the reinforcement learning agent for the initial optical path configuration strategy corresponding to the current time step to obtain the optimal optical path configuration strategy corresponding to the current time step.
[0084] And, Step 120: Within the current time step, execute a preset meta-learning algorithm according to the initial optical path configuration strategy corresponding to the current time step, the trajectory data set generated through reinforcement learning interaction during the retraining of the reinforcement learning agent at the current time step, the preset reinforcement learning rate, and the meta-learning rate, to obtain the updated result data of the initial optical path configuration strategy corresponding to the current time step, and use this updated result data as the initial optical path configuration strategy corresponding to the next time step.
[0085] In an example of Step 120, the specific meta-learning algorithm is shown in Formulas (1) and (2). Among them, after executing Formula (1), the policy parameter update needs to be executed for w components with a fixed length in a preset window. In this embodiment of the present application, only the first update of the policy parameter corresponding to Formula (2) is taken as an example. In practical applications, after executing Formula (2), there may be more iterative update processes such as the second update of the policy parameter.
[0086]
[0087] In the above formulas, represents a temporary variable, used to represent the influence of the traffic at the current time step t on the first policy update;
[0088] α represents the preset reinforcement learning rate;
[0089] β represents the preset meta-learning rate;
[0090] represents the policy parameter of the initial optical path configuration strategy corresponding to the current time step;
[0091] represents the policy parameter of the first updated optical path configuration strategy;
[0092] π represents the optical path configuration strategy;
[0093] represents the initial optical path configuration strategy corresponding to the current time step;
[0094] represents the updated optical path configuration strategy for the first update. If w = 1, the updated optical path configuration strategy for the first update is the updated result data of the initial optical path configuration strategy corresponding to the current time step, and is also the initial optical path configuration strategy for the next time step t + 1; if w = 2, it is necessary to use as the current optical path configuration strategy to be updated, and modify the subscript of the parameter with subscript 1 in formula (2) to the value of the current component w, that is, 2, and modify the subscript of the parameter with subscript 0 to w - 1, that is, 1, and then execute the modified formula (2) until the calculation of formula (2) corresponding to all components of the window is completed, and the final updated result data of the initial optical path configuration strategy corresponding to the current time step t, which is also the initial optical path configuration strategy for the next time step t + 1.
[0095] represents the set of trajectory data generated for the first time through reinforcement learning interaction during the retraining process of the reinforcement learning agent for the initial optical path configuration strategy corresponding to the current time step, that is, using as the policy and M 0 as the traffic environment to generate a set of trajectory data through reinforcement learning interaction;
[0096] represents using as the policy and M 1 as the traffic environment to generate a set of trajectory data for the second time through reinforcement learning interaction;
[0097] represents using as the policy and as the reinforcement learning optimization objective function for the data;
[0098] represents using as the policy and as the reinforcement learning optimization objective function for the data;
[0099] M 0 represents an example traffic environment;
[0100] M 1 represents another example traffic environment;
[0101] represents taking the partial derivative of ;
[0102] t' represents a temporary variable used to represent the accumulation of t.
[0103] Among them, Figure 4 Taking the current time steps \(t = 0\) and \(t'=0\) as examples, the logical solution processes of the above formulas (1) and (2) are illustrated, where the blue dashed arrow represents formula (1) and the black dashed arrow represents formula (2).
[0104] Based on this, the parallel execution process diagram of the reinforcement learning and meta - learning in step 110 and step 120 provided by the embodiments of the present application is as Figure 5 shown.
[0105] Among them, represents the policy parameters of the initialization policy optimized by the meta - learning algorithm at time \(t\).
[0106] In order to further improve the execution effectiveness and reliability of the optical path planning for the tidal flow network based on meta - learning, in a method for optical path planning of a tidal flow network based on meta - learning provided by the embodiments of the present application, see Figure 3 , step 200 in the method for optical path planning of a tidal flow network based on meta - learning specifically includes the following contents:
[0107] Step 210: Asynchronously send the optimal optical path configuration policy corresponding to the current time step to a controller within the current time step, so that the controller reconstructs the optical path for the optical network layer corresponding to the tidal flow network based on the optimal optical path configuration policy corresponding to the current time step to correspondingly update the topology of the IP layer corresponding to the tidal flow network.
[0108] Among them, the controller in the embodiments of the present application can also be called an SDON controller. In order to further improve the execution effectiveness and reliability of the optical path planning for the tidal flow network based on meta - learning, see Figure 6 , this controller is used to execute the following steps:
[0109] Step 211: Receive the optimal optical path configuration policy corresponding to the current time step.
[0110] Step 212: If it is determined that the optimal optical path configuration policy corresponding to the previous time step is locally stored, then according to the comparison result between the optimal optical path configuration policy corresponding to the previous time step and the optimal optical path configuration policy corresponding to the current time step, in each optical cross - connect device in the optical network layer corresponding to the tidal flow network, determine the optical cross - connect device to be reconstructed for the optical path at the current time as the target device, and extract the optical cross - connect reconstruction configuration data for each of the target devices from the optimal optical path configuration policy corresponding to the current time step.
[0111] Step 213: Send the corresponding optical cross-connection reconstruction configuration data to each of the target devices, so that each of the target devices performs optical cross-connection reconstruction according to the received optical cross-connection reconstruction configuration data and returns the corresponding optical cross-connection reconfiguration completion information.
[0112] Step 214: Receive the optical cross-connection reconfiguration completion information returned by each of the target devices respectively. If it is determined that all the optical cross-connection reconfiguration completion information corresponding to each of the target devices has been received, then send out the global reconfiguration completion information.
[0113] In order to further improve the application reliability of the optical path planning process of the tidal flow network based on meta-learning, in a method for optical path planning of a tidal flow network based on meta-learning provided in an embodiment of the present application, see Figure 3 , after step 200 in the method for optical path planning of a tidal flow network based on meta-learning, the following specific content is further included:
[0114] Step 300: If the global reconfiguration completion information sent by the controller is received and the initial optical path configuration strategy corresponding to the next time step has been obtained currently, it is determined that the optical path planning of the tidal flow network corresponding to the current time step has been completed.
[0115] And, in order to further improve the accuracy and effectiveness of the optical path planning of the tidal flow network, in a method for optical path planning of a tidal flow network based on meta-learning provided in an embodiment of the present application, the traffic distribution data includes the traffic matrix of the tidal flow network, where each element in the traffic matrix is respectively used to represent the traffic between two nodes in the tidal flow network, and the communication frequencies corresponding to each of the elements are represented by different colors or different numerical parameters of the same color, where the numerical parameters include at least one of hue value, saturation value, and lightness value; based on this, see Figure 3 , before step 110 and step 120 in the method for optical path planning of a tidal flow network based on meta-learning, the following specific content is further included:
[0116] Step 010: Obtain the current traffic matrix of the tidal flow network.
[0117] Specifically, an SDON traffic awareness component can be used to obtain the current traffic matrix of the tidal flow network and transmit the current traffic matrix of the tidal flow network to the SDON manager.
[0118] From a software level, the present application also provides a device for optical path planning of a tidal flow network based on meta-learning for executing all or part of the content in the method for optical path planning of a tidal flow network based on meta-learning, seeFigure 7 , the optical path planning device for tidal flow network based on meta - learning specifically includes the following:
[0119] The reinforcement learning and meta - learning parallel module 10 is used to retrain the reinforcement learning agent for the initial optical path configuration strategy corresponding to the current time step according to the traffic distribution data of the tidal flow network at the current time step, so as to obtain the optimal optical path configuration strategy corresponding to the current time step; and, within the current time step, update the initial optical path configuration strategy corresponding to the current time step based on the meta - learning algorithm to obtain the initial optical path configuration strategy corresponding to the next time step;
[0120] The optical path reconstruction module 20 is used to reconstruct the optical path of the optical network layer corresponding to the tidal flow network based on the optimal optical path configuration strategy, so as to correspondingly update the topology structure of the IP layer corresponding to the tidal flow network.
[0121] The embodiment of the optical path planning device for tidal flow network based on meta - learning provided by this application can specifically be used to execute the processing flow of the embodiment of the optical path planning method for tidal flow network based on meta - learning in the above - mentioned embodiment. Its functions will not be elaborated here, and reference can be made to the detailed description of the embodiment of the optical path planning method for tidal flow network based on meta - learning above.
[0122] The part of the optical path planning for tidal flow network based on meta - learning by the optical path planning device for tidal flow network based on meta - learning can be completed in a server or a client device. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario, etc. This application does not make a limitation in this regard. If all operations are completed in the client device, the client device may further include a processor for the specific processing of the optical path planning for tidal flow network based on meta - learning.
[0123] The above - mentioned client device may have a communication module (i.e., communication unit), which can communicate with a remote server to realize data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, it may also include a server of an intermediate platform, such as a server of a third - party server platform having a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed device.
[0124] Any suitable network protocol can be used for communication between the above-mentioned server and the client device, including network protocols that have not been developed as of the filing date of this application. The network protocol can, for example, include TCP / IP protocol, UDP / IP protocol, HTTP protocol, HTTPS protocol, etc. Of course, the network protocol can also, for example, include the RPC protocol (Remote Procedure Call Protocol) and the REST protocol (Representational State Transfer) used on top of the above-mentioned protocols, etc.
[0125] As can be seen from the above description, the meta-learning-based optical path planning device for tidal flow networks provided by the embodiments of this application can reduce costs such as the time and energy consumption for the agent to retrain, thereby improving the optical network management level, providing an effective solution for the traffic diversion of tidal services, and improving the network service quality.
[0126] The embodiments of this application also provide an electronic device, which may include a processor, a memory, a receiver, and a transmitter. The processor is used to execute the meta-learning-based optical path planning method for tidal flow networks mentioned in the above embodiments. The processor and the memory can be connected through a bus or other means. Taking the bus connection as an example, the receiver can be connected to the processor and the memory in a wired or wireless manner.
[0127] The processor can be a Central Processing Unit (CPU). The processor can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., or a combination of the above types of chips.
[0128] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the meta-learning-based optical path planning method for tidal flow networks in the embodiments of this application. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory, that is, to implement the meta-learning-based optical path planning method in the above method embodiments.
[0129] The memory may include a program storage area and a data storage area. Among them, the program storage area can store the operating system and application programs required for at least one function; the data storage area can store data created by the processor and the like. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely provided with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0130] The one or more modules are stored in the memory and, when executed by the processor, execute the meta-learning-based tidal flow network optical path planning method in the embodiments.
[0131] In some embodiments of the present application, the user equipment may include a processor, a memory, and a transceiver unit. The transceiver unit may include a receiver and a transmitter. The processor, the memory, the receiver, and the transmitter may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to transmit and receive signals.
[0132] As an implementation manner, the functions of the receiver and the transmitter in the present application may be considered to be implemented through a transceiver circuit or a dedicated transceiver chip, and the processor may be considered to be implemented through a dedicated processing chip, a processing circuit, or a general-purpose chip.
[0133] As another implementation manner, it may be considered to use a general-purpose computer to implement the server provided in the embodiments of the present application. That is, the program codes for implementing the functions of the processor, the receiver, and the transmitter are stored in the memory, and the general-purpose processor implements the functions of the processor, the receiver, and the transmitter by executing the codes in the memory.
[0134] In an example, the electronic device may be an SDON manager, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the aforementioned meta-learning-based tidal flow network optical path planning method is implemented.
[0135] In addition, an embodiment of the present application further provides an SDON controller, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the SDON controller is communicatively connected to an SDON manager, and the SDON manager is used to execute the aforementioned meta-learning-based tidal traffic network optical path planning method; and the SDON manager is used to asynchronously send the optimal optical path configuration strategy corresponding to the current time step to the SDON controller within the current time step and receive the global reconfiguration completion information sent by the SDON controller; when the processor executes the computer program, the aforementioned steps 211 to 214 are implemented.
[0136] On this basis, see Figure 8 , an embodiment of the present application also provides a tidal traffic network optical path planning system based on meta-learning, the system includes the above-mentioned SDON manager and SDON controller;
[0137] The SDON manager is used to obtain the traffic distribution data of the tidal traffic network at the current time step from the traffic layer of the tidal traffic network; the SDON controller is respectively communicated with each optical cross-connection device in the optical network layer corresponding to the tidal traffic network.
[0138] like Figure 8 As shown, in a cross-layer communication system, as the service distribution changes, relying solely on the regulation of the IP layer cannot solve the problem of poor global delay stability caused by the unstable average number of service hops under a fixed IP virtual topology. The degree of adaptation between the service distribution and the topology determines the average number of hops for global services. Therefore, this application starts with the management and control system of SDON, with the goal of simplifying global cross-layer communications, and changes the topology of the IP layer through optical path reconstruction of the optical network layer, so that the service distribution and topology at each moment are well adapted, thereby stabilizing the average number of service hops on the time scale and improving the quality of network services.
[0139] To further illustrate the above embodiments, the present application also provides a specific application example of a tidal traffic network optical path planning method based on meta-learning implemented by a tidal traffic network optical path planning system based on meta-learning. Under tidal traffic services, the reinforcement learning agent fails when the service distribution drifts. The algorithm proposed in the application example of this application aims to reduce the cost (time, energy consumption, etc.) of retraining the agent, thereby improving the level of optical network management and control, providing an effective solution for traffic diversion of tidal services, and improving the quality of network services. See Figure 9 , the application example specifically includes the following steps:
[0140] Step 1: If Figure 9As shown, at the beginning of each time step, it is considered that when the traffic distribution of the time step is inconsistent with that of the previous step, i.e., distribution drift occurs, the decision-making of the reinforcement learning agent fails and retraining is required. The SDON manager re-performs reinforcement learning training DRL. Using the current traffic matrix as the environmental feature and maximizing throughput as the goal, a Markov decision process is modeled. After the reinforcement learning agent undergoes several rounds of training, it can output the optimal optical path configuration policy as the current optical path configuration policy solution.
[0141] Step 2: After the SDON manager re-performs reinforcement learning training, it outputs the current optical path configuration policy solution (i.e., the global reconfiguration policy) and sends it to the SDON controller in the form of asynchronous messages. After receiving the optical path configuration policy solution, the SDON controller analyzes the optical cross-connect (OXC) reconstruction solution and sends reconfiguration instructions to the OXCs involved in the reconstruction. The reconfiguration policies for all OXCs in the network are synchronized to ensure the reconstruction efficiency. After receiving the reconstruction instructions, the OXC executes the optical cross-connect reconstruction solution and sends a reconstruction completion confirmation message to the SDON controller after the reconstruction is completed.
[0142] Step 3: After the SDON controller receives the reconstruction confirmation messages from all OXCs, it sends a reconstruction completion confirmation message to the SDON manager to inform the SDON manager that the current global reconfiguration is completed.
[0143] Step 4: It can be executed in parallel with Step 1. Without waiting for the SDON manager to send the global optical path configuration policy to the SDON controller or for the asynchronous callback instruction of this instruction, the meta-learning algorithm such as formulas (1) and (2) above can be executed in parallel.
[0144] Step 5: The SDON manager confirms the completion of the meta-learning algorithm and receives the completion instructions from all OXCs, that is, the reconfiguration plan for the current time step is completed.
[0145] In summary, the optical path planning method for tidal flow networks based on meta-learning provided by the application example of this application applies meta-learning to the optical network path planning in the context of tidal services for traffic grooming, and has the following beneficial effects:
[0146] 1. Reduce the average number of hops of services;
[0147] 2. Simplify cross-layer communication and reduce the global service delay;
[0148] 3. Accelerate the network configuration policy calculation process and improve the network configuration efficiency;
[0149] 4. Improve the network service quality and user satisfaction.
[0150] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing optical path planning method for a tidal flow network based on meta-learning are implemented. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the art.
[0151] The embodiments of the present application also provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the foregoing optical path planning method for a tidal flow network based on meta-learning are implemented.
[0152] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.
[0153] It should be clear that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.
[0154] In the present application, the features described and / or illustrated for one embodiment can be used in the same manner or in a similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0155] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A tidal traffic network optical path planning method based on meta-learning, characterized in that: include: According to the flow distribution data of the tidal flow network at the current time step, the reinforcement learning agent is retrained for the initial light path configuration strategy corresponding to the current time step to obtain the optimal light path configuration strategy corresponding to the current time step; and, within the current time step, the initial light path configuration strategy corresponding to the current time step is updated based on the meta-learning algorithm to obtain the initial light path configuration strategy corresponding to the next time step; Based on the optimal optical path configuration strategy, the optical network layer corresponding to the tidal traffic network is reconstructed to correspondingly update the topological structure of the IP layer corresponding to the tidal traffic network.
2. According to claim 1, the method for tidal flow network optical path planning based on meta-learning is characterized in that: If the current time step is the first time step in a preset time period, the initial light path configuration strategy corresponding to the current time step is randomly generated in advance; If the current time step is not the first time step in the preset time period, the initial light path configuration strategy corresponding to the current time step is obtained by updating the optimal light path configuration strategy corresponding to the previous time step based on the meta-learning algorithm in the previous time step.
3. According to claim 1, the method for tidal flow network optical path planning based on meta-learning is characterized in that: The step of updating the initial light path configuration strategy corresponding to the current time step based on the meta-learning algorithm in the current time step to obtain the initial light path configuration strategy corresponding to the next time step includes: In the current time step, a preset meta-learning algorithm is executed according to the initial light path configuration strategy corresponding to the current time step, the trajectory data set generated by reinforcement learning interaction during the retraining of the reinforcement learning agent in the current time step, the preset reinforcement learning learning rate and the meta-learning learning rate to obtain updated result data of the initial light path configuration strategy corresponding to the current time step, and the updated result data is used as the initial light path configuration strategy corresponding to the next time step.
4. According to claim 1, the method for tidal flow network optical path planning based on meta-learning is characterized in that: The optical path reconstruction of the optical network layer corresponding to the tidal traffic network based on the optimal optical path configuration strategy to correspondingly update the topology structure of the IP layer corresponding to the tidal traffic network includes: In the current time step, the optimal optical path configuration strategy corresponding to the current time step is asynchronously sent to a controller, so that the controller reconstructs the optical path of the optical network layer corresponding to the tidal traffic network based on the optimal optical path configuration strategy corresponding to the current time step to update the topology structure of the IP layer corresponding to the tidal traffic network.
5. According to claim 4, the method for tidal flow network optical path planning based on meta-learning is characterized in that: The controller is used to execute the following: Receive the optimal light path configuration strategy corresponding to the current time step; If it is determined that the optimal optical path configuration strategy corresponding to the previous time step is stored locally, then according to the comparison result between the optimal optical path configuration strategy corresponding to the previous time step and the optimal optical path configuration strategy corresponding to the current time step, the optical cross-connection device currently to be optically reconstructed is determined as the target device in each optical cross-connection device in the optical network layer corresponding to the tidal flow network, and the optical cross-connection reconstruction configuration data for each of the target devices is extracted from the optimal optical path configuration strategy corresponding to the current time step; Sending the corresponding optical cross-connection reconstruction configuration data to each of the target devices respectively, so that each of the target devices respectively performs optical cross-connection reconstruction according to the received optical cross-connection reconstruction configuration data and returns corresponding optical cross-connection reconfiguration completion information; The optical cross-connection reconfiguration completion information returned by each of the target devices is received respectively, and if it is determined that the optical cross-connection reconfiguration completion information corresponding to all the target devices is received, a global reconfiguration completion information is issued.
6. According to claim 5, the method for tidal flow network optical path planning based on meta-learning is characterized in that: Also includes: If the global reconfiguration completion information sent by the controller is received and the initial optical path configuration strategy corresponding to the next time step has been obtained, it is determined that the tidal flow network optical path planning corresponding to the current time step has been completed.
7. The method for tidal flow network optical path planning based on meta-learning according to claim 1 is characterized in that: The traffic distribution data includes a traffic matrix of the tidal traffic network, wherein each element in the traffic matrix is used to represent the traffic between two nodes in the tidal traffic network, and the communication frequency corresponding to each of the elements is represented by different colors or different numerical parameters of the same color, wherein the numerical parameters include at least one of a hue value, a saturation value, and a brightness value; Correspondingly, before retraining the reinforcement learning agent according to the flow distribution data of the tidal flow network at the current time step and the initial light path configuration strategy corresponding to the current time step, it also includes: Obtain the current flow matrix of the tidal flow network.
8. An SDON manager, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the meta-learning based tidal flow network optical path planning method as described in any one of claims 1 to 7.
9. An SDON controller, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The SDON controller is communicatively connected to an SDON manager, and the SDON manager is used to execute the tidal flow network optical path planning method based on meta-learning according to any one of claims 1 to 7; and the SDON manager is used to asynchronously send the optimal optical path configuration strategy corresponding to the current time step to the SDON controller within the current time step and receive the global reconfiguration completion information sent by the SDON controller; When the processor executes the computer program, the following contents are achieved: Receive the optimal light path configuration strategy corresponding to the current time step; If it is determined that the optimal optical path configuration strategy corresponding to the previous time step is stored locally, then according to the comparison result between the optimal optical path configuration strategy corresponding to the previous time step and the optimal optical path configuration strategy corresponding to the current time step, the optical cross-connection device currently to be optically reconstructed is determined as the target device in each optical cross-connection device in the optical network layer corresponding to the tidal flow network, and the optical cross-connection reconstruction configuration data for each of the target devices is extracted from the optimal optical path configuration strategy corresponding to the current time step; Sending the corresponding optical cross-connection reconstruction configuration data to each of the target devices respectively, so that each of the target devices respectively performs optical cross-connection reconstruction according to the received optical cross-connection reconstruction configuration data and returns corresponding optical cross-connection reconfiguration completion information; The optical cross-connection reconfiguration completion information returned by each of the target devices is received respectively, and if it is determined that the optical cross-connection reconfiguration completion information corresponding to all the target devices is received, a global reconfiguration completion information is issued.
10. A tidal flow network optical path planning system based on meta-learning, characterized in that: include: The SDON manager as claimed in claim 8 and the SDON controller as claimed in claim 9; The SDON manager is used to obtain the traffic distribution data of the tidal traffic network at the current time step from the traffic layer of the tidal traffic network; The SDON controller is communicatively connected to each optical cross-connection device in the optical network layer corresponding to the tidal traffic network.
Citation Information
Patent Citations
Intelligent network path optimization method and system based on deep reinforcement learning
CN116527567A
Network traffic engineering method and device based on meta-learning
CN118194981A
Multi-objective reinforcement learning method and system for adaptive network flow optimization
CN119071236A