Tidal flow network optical path planning method and system based on meta learning

By applying meta-learning algorithms to update optical path configuration strategies and topology in optical networks, the high cost of agent retraining under tidal traffic services is solved, thereby improving network service quality and management efficiency.

CN120186501BActive Publication Date: 2026-03-31BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing optical network path planning algorithms have high time and energy costs for agent retraining under tidal traffic services, making it difficult to effectively cope with changes in service distribution and leading to a decline in network service quality.

Method used

A meta-learning-based approach is adopted, which involves retraining the reinforcement learning agent on the initial optical path configuration strategy at the current time step, updating the optical path configuration strategy using meta-learning algorithms, and combining the topology reconstruction of the optical network layer and IP layer to reduce the time and energy consumption of agent retraining.

Benefits of technology

It improves the control level of optical networks, reduces the time and energy consumption for retraining agents, enhances network service quality, and provides an effective solution for traffic diversion of tidal services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186501B_ABST
    Figure CN120186501B_ABST
Patent Text Reader

Abstract

The application provides a tidal flow network optical path planning method, manager, controller and system based on meta-learning, the method comprises the following steps: according to the traffic distribution data of the tidal flow network at the current time step, the retraining of the reinforcement learning intelligent agent is carried out for the initial optical path configuration strategy corresponding to the current time step, so as to obtain the optimal optical path configuration strategy corresponding to the current time step; and in the current time step, the initial optical path configuration strategy corresponding to the current time step is updated based on the meta-learning algorithm to obtain the initial optical path configuration strategy corresponding to the next time step; the optical path of the optical network layer corresponding to the tidal flow network is reconstructed based on the optimal optical path configuration strategy to update the topology structure of the IP layer corresponding to the tidal flow network. The application can reduce the time and energy consumption of the retraining of the intelligent agent, thereby improving the optical network management and control level, providing an effective solution for the traffic diversion of the tidal service, and improving the network service quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of software-defined optical networks (SDON), and in particular to a method and system for optical path planning in tidal flow networks based on meta-learning. Background Technology

[0002] Communication networks serve humanity, and due to human cyclical work and rest schedules, communication service demands also exhibit cyclical changes. For example, the demand for map services increases during morning and evening rush hours, while the demand for food delivery services increases during mealtimes. Network load also shows relatively regular tidal variations on a daily, weekly, monthly, and yearly basis, thus forming tidal traffic networks. Optical networks need to provide physical resources for these tidal communication services.

[0003] Currently, optical path planning in Software-Defined Optical Networks (SDON) is a crucial issue. Similar RSA (Routing and Resource Allocation) problems have been proven to be nondeterministic polynomial (NP-hard) problems, with significant difficulty and complexity being widely acknowledged in the industry. Optical network path planning under tidal traffic conditions is even more complex than traditional path planning. Optical path planning also faces challenges such as surging traffic volumes and distribution drift, highlighting the shortcomings of traditional optical path planning algorithms. Therefore, designing an optical network path planning algorithm based on tidal traffic services to address the tidal nature of services is essential for improving network service quality.

[0004] Existing optical network path planning algorithms include greedy methods, linear programming-based methods, heuristic-based methods, and reinforcement learning-based methods. Among these, reinforcement learning-based methods are currently the most popular solution in academia, offering both high optimization efficiency and fast convergence speed compared to other algorithms. However, a drawback of reinforcement learning-based methods is that the decision model needs to be retrained when the service distribution changes, resulting in significant time and energy costs.

[0005] Therefore, how to reduce the time and energy consumption costs of agent retraining under tidal traffic services in order to improve the level of optical network management is an urgent problem to be solved. Summary of the Invention

[0006] In view of this, embodiments of this application provide a method, manager, controller and system for optical path planning in tidal flow networks based on meta-learning, in order to eliminate or improve one or more defects existing in the prior art.

[0007] One aspect of this application provides a meta-learning-based optical path planning method for tidal flow networks, including:

[0008] Based on the traffic distribution data of the tidal flow network at the current time step, the reinforcement learning agent is retrained for the initial optical path configuration strategy corresponding to the current time step to obtain the optimal optical path configuration strategy corresponding to the current time step; and within the current time step, the initial optical path configuration strategy corresponding to the current time step is updated based on the meta-learning algorithm to obtain the initial optical path configuration strategy corresponding to the next time step.

[0009] Based on the optimal optical path configuration strategy, the optical network layer corresponding to the tidal flow network is reconstructed to update the topology of the IP layer corresponding to the tidal flow network.

[0010] In some embodiments of this application, if the current time step is the first time step within a preset time period, the initial optical path configuration strategy corresponding to the current time step is generated randomly in advance.

[0011] If the current time step is not the first time step within the preset time period, the initial optical path configuration strategy corresponding to the current time step is obtained by updating the optimal optical path configuration strategy corresponding to the previous time step based on the meta-learning algorithm in the previous time step.

[0012] In some embodiments of this application, the step of updating the initial optical path configuration strategy corresponding to the current time step based on the meta-learning algorithm to obtain the initial optical path configuration strategy corresponding to the next time step includes:

[0013] Within the current time step, a preset meta-learning algorithm is executed based on the initial optical path configuration strategy corresponding to the current time step, the trajectory data set generated by reinforcement learning interaction during the retraining process of the reinforcement learning agent in the current time step, the preset reinforcement learning learning rate, and the meta-learning learning rate, in order to obtain the updated result data of the initial optical path configuration strategy corresponding to the current time step, and the updated result data is used as the initial optical path configuration strategy for the next time step.

[0014] In some embodiments of this application, the step of reconstructing the optical network layer corresponding to the tidal traffic network based on the optimal optical path configuration strategy to update the topology of the IP layer corresponding to the tidal traffic network includes:

[0015] Within the current time step, the optimal optical path configuration strategy corresponding to the current time step is asynchronously sent to a controller, so that the controller can reconstruct the optical network layer corresponding to the tidal traffic network based on the optimal optical path configuration strategy corresponding to the current time step, so as to update the topology of the IP layer corresponding to the tidal traffic network.

[0016] In some embodiments of this application, the controller is used to perform the following:

[0017] Receive the optimal optical path configuration strategy corresponding to the current time step;

[0018] If the optimal optical path configuration strategy corresponding to the previous time step is determined to be stored locally, then based on the comparison result between the optimal optical path configuration strategy corresponding to the previous time step and the optimal optical path configuration strategy corresponding to the current time step, the optical cross-connect device to be reconfigured is determined as the target device among the optical cross-connect devices in the optical network layer corresponding to the tidal flow network, and the optical cross-connect reconfiguration configuration data for each target device is extracted from the optimal optical path configuration strategy corresponding to the current time step.

[0019] The corresponding optical cross-connect reconfiguration configuration data is sent to each of the target devices, so that each of the target devices performs optical cross-connect reconfiguration according to the received optical cross-connect reconfiguration configuration data and returns the corresponding optical cross-connect reconfiguration completion information;

[0020] The system receives optical cross-connect reconfiguration completion information returned by each of the target devices. If it is determined that the optical cross-connect reconfiguration completion information corresponding to each of the target devices has been received, a global reconfiguration completion information is sent.

[0021] In some embodiments of this application, the meta-learning-based optical path planning method for tidal flow networks further includes:

[0022] If the controller sends a global reconfiguration completion message and the initial optical path configuration strategy for the next time step has been obtained, then it is determined that the tidal flow network optical path planning for the current time step has been completed.

[0023] In some embodiments of this application, the traffic distribution data includes a traffic matrix of the tidal traffic network, wherein each element in the traffic matrix is ​​used to represent the traffic between two nodes in the tidal traffic network, and the communication frequency corresponding to each element is represented by different colors or different numerical parameters of the same color, wherein the numerical parameters include at least one of hue value, saturation value and brightness value.

[0024] Correspondingly, before retraining the reinforcement learning agent based on the initial optical path configuration strategy corresponding to the current time step according to the flow distribution data of the tidal flow network at the current time step, the following steps are also included:

[0025] Obtain the current flow matrix of the tidal flow network.

[0026] Another aspect of this application provides a meta-learning-based optical path planning device for tidal flow networks, comprising:

[0027] The reinforcement learning and meta-learning parallel module is used to retrain the reinforcement learning agent based on the traffic distribution data of the tidal flow network at the current time step, and to obtain the optimal optical path configuration strategy for the current time step; and, within the current time step, to update the initial optical path configuration strategy for the current time step based on the meta-learning algorithm to obtain the initial optical path configuration strategy for the next time step.

[0028] The optical path reconstruction module is used to reconstruct the optical network layer corresponding to the tidal traffic network based on the optimal optical path configuration strategy in order to update the topology of the IP layer corresponding to the tidal traffic network.

[0029] A third aspect of this application provides an SDON manager, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the meta-learning-based tidal flow network optical path planning method.

[0030] A fourth aspect of this application provides an SDON controller, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The SDON controller is communicatively connected to an SDON manager, which executes the meta-learning-based tidal flow network optical path planning method. The SDON manager asynchronously sends the optimal optical path configuration strategy corresponding to the current time step to the SDON controller within the current time step and receives global reconfiguration completion information sent by the SDON controller.

[0031] When the processor executes the computer program, it performs the following:

[0032] Receive the optimal optical path configuration strategy corresponding to the current time step;

[0033] If the optimal optical path configuration strategy corresponding to the previous time step is determined to be stored locally, then based on the comparison result between the optimal optical path configuration strategy corresponding to the previous time step and the optimal optical path configuration strategy corresponding to the current time step, the optical cross-connect device to be reconfigured is determined as the target device among the optical cross-connect devices in the optical network layer corresponding to the tidal flow network, and the optical cross-connect reconfiguration configuration data for each target device is extracted from the optimal optical path configuration strategy corresponding to the current time step.

[0034] The corresponding optical cross-connect reconfiguration configuration data is sent to each of the target devices, so that each of the target devices performs optical cross-connect reconfiguration according to the received optical cross-connect reconfiguration configuration data and returns the corresponding optical cross-connect reconfiguration completion information;

[0035] The system receives optical cross-connect reconfiguration completion information returned by each of the target devices. If it is determined that the optical cross-connect reconfiguration completion information corresponding to each of the target devices has been received, a global reconfiguration completion information is sent.

[0036] The fifth aspect of this application provides a meta-learning-based optical path planning system for tidal flow networks, the SDON manager mentioned in the third aspect above, and the SDON controller mentioned in the fourth aspect above;

[0037] The SDON manager is used to obtain the traffic distribution data of the tidal traffic network at the current time step from the traffic layer of the tidal traffic network;

[0038] The SDON controller communicates with each optical cross-connect device in the optical network layer corresponding to the tidal flow network.

[0039] The sixth aspect of this application provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the meta-learning-based optical path planning method for tidal flow networks.

[0040] The seventh aspect of this application provides a computer program product including a computer program that, when executed by a processor, implements the meta-learning-based optical path planning method for tidal flow networks as described.

[0041] The meta-learning-based optical path planning method for tidal traffic networks provided in this application retrains a reinforcement learning agent based on the traffic distribution data of the tidal traffic network at the current time step, thereby obtaining the optimal optical path configuration strategy for the current time step. Furthermore, within the current time step, the initial optical path configuration strategy is updated based on a meta-learning algorithm to obtain the initial optical path configuration strategy for the next time step. Based on the optimal optical path configuration strategy, the optical network layer corresponding to the tidal traffic network is reconstructed to update the topology of the IP layer corresponding to the tidal traffic network. This method reduces the time and energy consumption costs associated with agent retraining, thereby improving the optical network management level, providing an effective solution for traffic diversion in tidal services, and enhancing network service quality.

[0042] Additional advantages, objectives, and features of this application will be set forth in part in the description which follows, and will in part become apparent to those skilled in the art upon review of the following description, or may be learned by practice of the application. The objectives and other advantages of this application can be realized and obtained by means of the structures specifically pointed out in the specification and drawings.

[0043] Those skilled in the art will understand that the purposes and advantages that can be achieved with this application are not limited to those specifically described above, and that the above and other purposes that this application can achieve will be more clearly understood from the following detailed description. Attached Figure Description

[0044] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, do not constitute a limitation thereof. The components in the drawings are not drawn to scale but are merely for illustrating the principles of this application. For ease of illustration and description of certain parts of this application, corresponding portions in the drawings may be enlarged, i.e., may appear larger relative to other components in an exemplary device actually manufactured according to this application. In the drawings:

[0045] Figure 1 The diagram illustrates examples of existing greedy methods, linear programming-based methods, heuristic-based methods, and reinforcement learning-based methods.

[0046] Figure 2 This is a schematic diagram of the first process of the tidal flow network optical path planning method based on meta-learning in one embodiment of this application.

[0047] Figure 3 This is a schematic diagram of the second process of the meta-learning-based optical path planning method for tidal flow networks in one embodiment of this application.

[0048] Figure 4 This is a schematic diagram of the logical solution process of formulas (1) and (2) in one embodiment of this application.

[0049] Figure 5 This is a schematic diagram illustrating the parallel execution process of reinforcement learning and meta-learning in steps 110 and 120 of this application, as exemplified in one of the examples.

[0050] Figure 6 This is a schematic diagram of the process executed by the SDON controller in one embodiment of this application.

[0051] Figure 7 This is a schematic diagram of the SDON manager in one embodiment of this application.

[0052] Figure 8 This is a schematic diagram of the execution logic of the tidal flow network optical path planning system based on meta-learning in an application example of this application.

[0053] Figure 9 This is an interactive diagram illustrating the meta-learning-based tidal flow network optical path planning method executed by the meta-learning-based tidal flow network optical path planning system in an application example of this application. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and their descriptions are used to explain this application, but are not intended to limit it.

[0055] It should also be noted that, in order to avoid obscuring this application with unnecessary details, only the structures and / or processing steps closely related to the solution according to this application are shown in the accompanying drawings, while other details that are not closely related to this application are omitted.

[0056] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0057] It should also be noted that, unless otherwise specified, the term "connection" in this article can refer not only to a direct connection, but also to an indirect connection involving an intermediary.

[0058] In the following description, embodiments of the present application will be illustrated with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar parts, or the same or similar steps.

[0059] See Figure 1 Existing optical network path planning algorithms include greedy methods, linear programming-based methods, heuristic-based methods, and reinforcement learning-based methods.

[0060] Greedy algorithms are characterized by their simple rules and fast solution, making them a common solution for time-delay-sensitive communication systems. However, their main drawback is their low degree of optimization.

[0061] The greatest advantage of linear programming is its high degree of optimization; its solution is the optimal solution. However, its main disadvantages are the difficulty in modeling, the difficulty in transforming constraints, and the extremely slow solution speed under large-scale variables. Therefore, it is not a conventional solution.

[0062] Heuristic methods offer a balance between the optimization level and convergence speed of greedy algorithms and linear programming algorithms, and are often used as a compromise. However, they suffer from poor interpretability and unstable convergence speed.

[0063] Machine learning solutions are currently the most popular solutions in academia. A qualitative analysis of the performance and convergence speed of reinforcement learning-based methods is shown in Table 1, where the number of stars is positively correlated with performance. Compared to other algorithms, reinforcement learning-based methods are characterized by higher optimization levels and faster convergence speeds. A drawback is that the decision model needs to be retrained when business distribution changes.

[0064] Table 1 Qualitative analysis of the optimization performance and convergence speed of reinforcement learning-based methods.

[0065]

[0066] Therefore, in order to address the problems of high time and energy consumption associated with retraining agents under tidal traffic conditions in existing reinforcement learning-based network optical path planning methods, this application provides a meta-learning-based tidal traffic network optical path planning method, a meta-learning-based tidal traffic network optical path planning method apparatus for executing the meta-learning-based tidal traffic network optical path planning method, an SDON manager, an SDON controller, a meta-learning-based tidal traffic network optical path planning system, a computer-readable storage medium, and a computer program product, which can apply meta-learning to optical network optical path planning under tidal traffic conditions for traffic diversion.

[0067] The following examples will provide a detailed description.

[0068] Based on this, embodiments of this application provide a meta-learning-based optical path planning method for tidal traffic networks that can be implemented by an SDON manager. See [link to relevant documentation]. Figure 2 The meta-learning-based optical path planning method for tidal flow networks specifically includes the following:

[0069] Step 100: Based on the traffic distribution data of the tidal flow network at the current time step, retrain the reinforcement learning agent for the initial optical path configuration strategy corresponding to the current time step to obtain the optimal optical path configuration strategy corresponding to the current time step; and, within the current time step, update the initial optical path configuration strategy corresponding to the current time step based on the meta-learning algorithm to obtain the initial optical path configuration strategy corresponding to the next time step.

[0070] In one or more embodiments of this application, a time step t in time series typically refers to the time interval between one observation point and the next observation point in the sequence data, and can be represented by t. It defines the sampling frequency and time scale of the time series data. For example, if the data is collected once a day, then each time step represents one day; if it is recorded once an hour, then each time step is one hour. The total length of the time series is represented by T, which can be any time length or positive infinity, and can be set according to the actual application requirements.

[0071] It is understandable that a tidal flow network refers to a network in which the current flow exhibits a clear periodic fluctuation over time, similar to the periodic changes in ocean tides. This phenomenon is particularly evident in 5G networks, especially in areas such as universities, industrial parks, CBD business districts, and large residential areas.

[0072] In one or more embodiments of this application, both the initial optical path configuration strategy and the optimal optical path configuration strategy include configuration data corresponding to each optical cross-connect device in the optical network layer of the current tidal flow network. The configuration data packet contains the cross-connection relationship between each optical cross-connect device and other optical cross-connect devices.

[0073] The optical cross-connect (OXC) is an important network unit in the optical network. Its function is analogous to a switch in a time-division multiplexing network, primarily used to complete cross-connections between multi-wavelength ring networks. As a node in the mesh-like optical network, its purpose is to achieve automatic configuration, protection / recovery, and reconfiguration of the optical network. The various optical cross-connect devices in the tidal flow network can be abbreviated as OXCs.

[0074] In step 100, based on the traffic distribution data of the tidal flow network at the current time step, the reinforcement learning agent is retrained according to the initial optical path configuration strategy corresponding to the current time step. This can be performed using existing reinforcement learning agent retraining methods. For example, at the beginning of each time step, it is assumed that the service distribution at that time step is inconsistent with the previous step, i.e., distribution drift has occurred, and the reinforcement learning agent's decision has failed, requiring retraining. The SDON manager retrains the reinforcement learning DRL. Using the current traffic matrix as environmental features and maximizing throughput as the objective, a Markov decision process model is performed. After several rounds of training, the reinforcement learning agent can output the optimal optical path configuration strategy.

[0075] It is understandable that Software-Defined Optical Network (SDON) is a centralized control and separation scheduling architecture. The optical path refers to an end-to-end optical channel with a defined wavelength, configured by optical switches within the optical network. By reconfiguring the optical network layer corresponding to the tidal traffic network based on the optimal optical path configuration strategy to update the topology of the IP layer corresponding to the tidal traffic network, the replanning of the optical path is completed.

[0076] In step 100, within the current time step, the execution of updating the initial optical path configuration strategy corresponding to the current time step based on the meta-learning algorithm to obtain the initial optical path configuration strategy corresponding to the next time step, and the execution of optical path reconstruction of the optical network layer corresponding to the tidal flow network based on the optimal optical path configuration strategy to update the topology of the IP layer corresponding to the tidal flow network, can be performed synchronously or asynchronously. Specifically, the execution process of the meta-learning algorithm can be performed synchronously with the reinforcement learning process, without waiting to obtain the optimal optical path configuration strategy corresponding to the current time step before execution, thereby effectively improving the efficiency of agent retraining and reducing energy consumption costs.

[0077] Understandably, meta-learning, also known as "learn to learn," is an important algorithmic framework in the field of artificial intelligence. Traditional machine learning methods aim to learn the distribution model of data or a specific strategy, while meta-learning aims to learn the distribution model of data distribution or the ability to learn a certain strategy—it represents a higher-level learning model. Meta-learning has shown good performance in image and language domains. It is typically used in few-shot learning, leveraging the model's learning results in multi-shot domains to retain knowledge and achieve rapid adaptation in small-shot domains.

[0078] Step 200: Based on the optimal optical path configuration strategy, perform optical path reconstruction on the optical network layer corresponding to the tidal traffic network to update the topology of the IP layer corresponding to the tidal traffic network.

[0079] As can be seen from the above description, the tidal traffic network optical path planning method based on meta-learning provided in this application embodiment can reduce the time and energy consumption costs of agent retraining, improve the efficiency of agent retraining, thereby improving the optical network management and control level, providing an effective solution for traffic diversion of tidal services, and improving network service quality.

[0080] To further improve the efficiency and reliability of retraining reinforcement learning agents in the process of optical path planning for tidal flow networks based on meta-learning, in the optical path planning method for tidal flow networks based on meta-learning provided in this application embodiment, if the current time step is the first time step within a preset time period, i.e., t=0, then the initial optical path configuration strategy corresponding to the current time step is randomly generated in advance; if the current time step is not the first time step within a preset time period, then the initial optical path configuration strategy corresponding to the current time step is obtained in advance after updating the optimal optical path configuration strategy corresponding to the previous time step based on the meta-learning algorithm in the previous time step.

[0081] To further improve the effectiveness and reliability of updating the initial optical path configuration strategy corresponding to the current time step using a meta-learning algorithm in the process of optical path planning for tidal flow networks based on meta-learning, this application provides a meta-learning-based optical path planning method for tidal flow networks, see [link to relevant documentation]. Figure 3 Step 100 in the meta-learning-based tidal flow network optical path planning method specifically includes the following:

[0082] Step 110: Based on the traffic distribution data of the tidal flow network at the current time step, retrain the reinforcement learning agent for the initial optical path configuration strategy corresponding to the current time step to obtain the optimal optical path configuration strategy corresponding to the current time step.

[0083] And, step 120: within the current time step, according to the initial optical path configuration strategy corresponding to the current time step, the trajectory data set generated by reinforcement learning interaction during the retraining process of the reinforcement learning agent in the current time step, the preset reinforcement learning learning rate and the preset meta-learning learning rate, execute the preset meta-learning algorithm to obtain the updated result data of the initial optical path configuration strategy corresponding to the current time step, and use the updated result data as the initial optical path configuration strategy corresponding to the next time step.

[0084] In one example of step 120, the specific meta-learning algorithm is shown in formulas (1) and (2), wherein, after executing formula (1), it is necessary to [implement a function] in a preset window. The strategy parameters are updated for each fixed-length component. In this embodiment, the first update of the strategy parameters corresponding to formula (2) is used as an example. In actual applications, after executing formula (2), there may be more iterative update processes such as the second update of the strategy parameters.

[0085]

[0086] In the above formula, This represents a temporary variable used to characterize the impact of the current time step t flow on the first policy update;

[0087] This represents the preset reinforcement learning rate;

[0088] This indicates the preset meta-learning rate;

[0089] The strategy parameters represent the initial optical path configuration strategy corresponding to the current time step;

[0090] The policy parameters represent the initial optical path configuration policy update.

[0091] Indicates the optical path configuration strategy;

[0092] This indicates the initial optical path configuration strategy corresponding to the current time step;

[0093] This indicates the initial optical path configuration strategy update. w =1, then the first updated optical path configuration strategy is the updated result data of the initial optical path configuration strategy corresponding to the current time step, which is also the initial optical path configuration strategy for the next time step t+1; if w =2, then it is necessary to use As the current optical path configuration strategy to be updated, the subscripts of parameters with subscript 1 in formula (2) are all modified to the value of the current component w, i.e., 2, and the subscripts of parameters with subscript 0 are modified to w-1, i.e., 1. Then, the modified formula (2) is executed until the calculation of formula (2) corresponding to all components of the window is completed, and the final formula (2) is executed. The updated result data of the initial optical path configuration strategy corresponding to the current time step t is also the initial optical path configuration strategy for the next time step t+1.

[0094] This represents the first set of trajectory data generated through reinforcement learning interaction during the retraining process of the reinforcement learning agent using the initial optical path configuration strategy corresponding to the current time step. As a strategy, with M 0 A collection of trajectory data generated through reinforcement learning interactions in a traffic environment;

[0095] Indicated by As a strategy, with M 1 The trajectory data set generated during the second reinforcement learning interaction in a traffic environment;

[0096] Indicated by As a strategy, with Optimize the objective function for reinforcement learning of data;

[0097] Indicated by As a strategy, with Optimize the objective function for reinforcement learning of data;

[0098] M 0 This represents an example traffic environment;

[0099] M 1 This represents another example of a traffic environment;

[0100] Indicates to Find the partial derivatives;

[0101] This represents a temporary variable used to represent the accumulation of t.

[0102] in, Figure 4 At the current time step t=0 and t Taking =0 as an example, the logical solution process of the above formulas (1) and (2) is illustrated. In this case, the blue dashed arrow represents formula (1) and the black dashed arrow represents formula (2).

[0103] Based on this, the parallel execution process diagram of reinforcement learning and meta-learning in steps 110 and 120 provided in the embodiments of this application is as follows: Figure 5 As shown.

[0104] in, This represents the policy parameters of the initialization policy after optimization by the meta-learning algorithm at time t.

[0105] To further improve the effectiveness and reliability of meta-learning-based optical path planning for tidal flow networks, this application provides a meta-learning-based optical path planning method for tidal flow networks, see [link to relevant documentation]. Figure 3 Step 200 in the meta-learning-based tidal flow network optical path planning method specifically includes the following:

[0106] Step 210: In the current time step, the optimal optical path configuration strategy corresponding to the current time step is asynchronously sent to a controller, so that the controller can reconstruct the optical network layer corresponding to the tidal traffic network based on the optimal optical path configuration strategy corresponding to the current time step to update the topology of the IP layer corresponding to the tidal traffic network.

[0107] The controller described in this embodiment can also be called an SDON controller. To further improve the effectiveness and reliability of the execution of meta-learning-based tidal flow network optical path planning, see [link to relevant documentation]. Figure 6 The controller is used to perform the following steps:

[0108] Step 211: Receive the optimal optical path configuration strategy corresponding to the current time step.

[0109] Step 212: If the optimal optical path configuration strategy corresponding to the previous time step is determined to be stored locally, then based on the comparison result between the optimal optical path configuration strategy corresponding to the previous time step and the optimal optical path configuration strategy corresponding to the current time step, among the optical cross-connect devices in the optical network layer corresponding to the tidal flow network, the optical cross-connect device to be reconfigured is determined as the target device, and optical cross-connect reconfiguration configuration data for each target device is extracted from the optimal optical path configuration strategy corresponding to the current time step.

[0110] Step 213: Send the corresponding optical cross-connect reconfiguration configuration data to each of the target devices, so that each of the target devices performs optical cross-connect reconfiguration according to the received optical cross-connect reconfiguration configuration data and returns the corresponding optical cross-connect reconfiguration completion information.

[0111] Step 214: Receive the optical cross-connect reconfiguration completion information returned by each of the target devices respectively. If it is determined that the optical cross-connect reconfiguration completion information corresponding to each of the target devices has been received, then send out the global reconfiguration completion information.

[0112] To further improve the reliability of the optical path planning process for tidal flow networks based on meta-learning, this application provides a method for optical path planning in tidal flow networks based on meta-learning, see [link to relevant documentation]. Figure 3 The optical path planning method for tidal flow networks based on meta-learning, after step 200, specifically includes the following:

[0113] Step 300: If the global reconfiguration completion information sent by the controller is received and the initial optical path configuration strategy corresponding to the next time step has been obtained, then it is determined that the tidal flow network optical path planning corresponding to the current time step has been completed.

[0114] Furthermore, to further improve the accuracy and effectiveness of optical path planning in tidal flow networks, in a meta-learning-based optical path planning method for tidal flow networks provided in this application embodiment, the flow distribution data includes a flow matrix of the tidal flow network, wherein each element in the flow matrix represents the flow between two nodes in the tidal flow network, and the communication frequency corresponding to each element is represented by different colors or different numerical parameters of the same color, wherein the numerical parameters include at least one of hue value, saturation value, and brightness value; based on this, see... Figure 3 The optical path planning method for tidal flow networks based on meta-learning includes the following content before steps 110 and 120:

[0115] Step 010: Obtain the current flow matrix of the tidal flow network.

[0116] Specifically, an SDON traffic sensing component can be used to obtain the current traffic matrix of the tidal traffic network and transmit the current traffic matrix of the tidal traffic network to the SDON manager.

[0117] From a software perspective, this application also provides a meta-learning-based tidal flow network optical path planning apparatus for executing all or part of the aforementioned meta-learning-based tidal flow network optical path planning method, see [link to relevant documentation]. Figure 7 The meta-learning-based tidal flow network optical path planning device specifically includes the following components:

[0118] The reinforcement learning and meta-learning parallel module 10 is used to retrain the reinforcement learning agent based on the traffic distribution data of the tidal flow network at the current time step, for the initial optical path configuration strategy corresponding to the current time step, so as to obtain the optimal optical path configuration strategy corresponding to the current time step; and, within the current time step, update the initial optical path configuration strategy corresponding to the current time step based on the meta-learning algorithm to obtain the initial optical path configuration strategy corresponding to the next time step.

[0119] The optical path reconstruction module 20 is used to reconstruct the optical network layer corresponding to the tidal traffic network based on the optimal optical path configuration strategy in order to update the topology of the IP layer corresponding to the tidal traffic network.

[0120] The embodiments of the meta-learning-based tidal flow network optical path planning device provided in this application can be used to execute the processing flow of the embodiments of the meta-learning-based tidal flow network optical path planning method in the above embodiments. Its functions will not be repeated here, but can be referred to the detailed description of the embodiments of the meta-learning-based tidal flow network optical path planning method in the above embodiments.

[0121] The meta-learning-based tidal flow network optical path planning device can perform the meta-learning-based tidal flow network optical path planning part in either a server or a client device. The choice can be made based on the processing power of the client device and the limitations of the user's usage scenario. This application does not impose any limitations in this regard. If all operations are performed in the client device, the client device may further include a processor for the specific processing of the meta-learning-based tidal flow network optical path planning.

[0122] The aforementioned client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side; in other implementation scenarios, it may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a distributed server structure.

[0123] The server and the client device can communicate using any suitable network protocol, including those not yet developed as of the date of this application. Such network protocols may include, for example, TCP / IP, UDP / IP, HTTP, HTTPS, etc. Furthermore, such network protocols may also include RPC (Remote Procedure Call Protocol) and REST (Representational State Transfer Protocol) protocols used on top of the aforementioned protocols.

[0124] As can be seen from the above description, the tidal traffic network optical path planning device based on meta-learning provided in this application embodiment can reduce the time and energy consumption costs of agent retraining, thereby improving the optical network management and control level, providing an effective solution for traffic diversion of tidal services, and improving network service quality.

[0125] This application also provides an electronic device, which may include a processor, a memory, a receiver, and a transmitter. The processor is used to execute the meta-learning-based optical path planning method for tidal flow networks mentioned in the above embodiments. The processor and memory can be connected via a bus or other means, taking a bus connection as an example. The receiver can be connected to the processor and memory via wired or wireless means.

[0126] The processor can be a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips.

[0127] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the meta-learning-based tidal flow network optical path planning method in the embodiments of this application. The processor executes various functional applications and data processing by running the non-transitory software programs, instructions, and modules stored in the memory, thereby implementing the meta-learning-based tidal flow network optical path planning method in the above method embodiments.

[0128] The memory may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the processor, etc. Furthermore, the memory may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0129] The one or more modules are stored in the memory, and when executed by the processor, the meta-learning-based optical path planning method for tidal flow networks in the embodiment is executed.

[0130] In some embodiments of this application, the user equipment may include a processor, a memory, and a transceiver unit. The transceiver unit may include a receiver and a transmitter. The processor, memory, receiver, and transmitter may be connected via a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to send and receive signals.

[0131] As one implementation method, the functions of the receiver and transmitter in this application can be implemented by transceiver circuits or dedicated transceiver chips, and the processor can be implemented by dedicated processing chips, processing circuits or general-purpose chips.

[0132] As another implementation approach, the server provided in this application embodiment can be implemented using a general-purpose computer. That is, the program code implementing the processor, receiver, and transmitter functions is stored in memory, and the general-purpose processor implements the processor, receiver, and transmitter functions by executing the code in memory.

[0133] In one example, the electronic device may be an SDON manager, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the aforementioned meta-learning-based tidal flow network optical path planning method.

[0134] In addition, this application embodiment also provides an SDON controller, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The SDON controller is communicatively connected to an SDON manager. The SDON manager is used to execute the aforementioned meta-learning-based tidal flow network optical path planning method. The SDON manager is used to asynchronously send the optimal optical path configuration strategy corresponding to the current time step to the SDON controller within the current time step and receive global reconfiguration completion information sent by the SDON controller. When the processor executes the computer program, it implements the aforementioned steps 211 to 214.

[0135] Based on this, see Figure 8 This application also provides a meta-learning-based optical path planning system for tidal flow networks, which includes the aforementioned SDON manager and SDON controller.

[0136] The SDON manager is used to obtain the traffic distribution data of the tidal traffic network at the current time step from the traffic layer of the tidal traffic network; the SDON controller communicates with each optical cross-connect device in the optical network layer corresponding to the tidal traffic network.

[0137] like Figure 8 As shown, in cross-layer communication systems, with changes in service distribution, relying solely on IP layer regulation cannot solve the problem of unstable average hop count leading to poor global latency stability under a fixed IP virtual topology. The degree of adaptation between service distribution and topology determines the global average hop count. Therefore, this application starts with the SDON control system, aiming to simplify global cross-layer communication. It changes the IP layer topology through optical path reconfiguration at the optical network layer, ensuring good adaptation between service distribution and topology at each time point, thereby stabilizing the average hop count over the time scale and improving network service quality.

[0138] To further illustrate the above embodiments, this application also provides a specific application example of a meta-learning-based optical path planning method for tidal traffic networks, implemented using a meta-learning-based tidal traffic network optical path planning system. Under tidal traffic services, reinforcement learning agents fail when service distribution drifts. The algorithm proposed in this application example aims to reduce the cost (time, energy consumption, etc.) of agent retraining, thereby improving the optical network management level, providing an effective solution for traffic diversion in tidal services, and improving network service quality. See also... Figure 9 This application example specifically includes the following steps:

[0139] Step 1: As Figure 9As shown, at the beginning of each time step, it is assumed that the traffic distribution at that time step is inconsistent with that of the previous step, i.e., distribution drift has occurred, and the reinforcement learning agent's decision has failed, requiring retraining. The SDON manager retrains the reinforcement learning DRL. Using the current traffic matrix as environmental features and aiming to maximize throughput, a Markov decision process is modeled. After several rounds of training, the reinforcement learning agent can output the optimal optical path configuration strategy as the current optical path configuration strategy.

[0140] Step 2: After retraining through reinforcement learning, the SDON manager outputs the current optical path configuration policy (i.e., the global reconfiguration policy) and sends it to the SDON controller asynchronously. Upon receiving the policy, the SDON controller analyzes the optical cross-connect (OXC) reconfiguration scheme and sends reconfiguration commands to the OXCs involved in the reconfiguration. The reconfiguration policy for all OXCs in the network is performed synchronously to ensure reconfiguration efficiency. After receiving the reconfiguration command, each OXC executes the optical cross-connect reconfiguration scheme and sends a reconfiguration completion confirmation message to the SDON controller upon completion.

[0141] Step 3: After the SDON controller receives the reconfiguration confirmation information from all OXCs, it sends a reconfiguration completion confirmation message to the SDON manager, informing the SDON manager that the current global reconfiguration is complete.

[0142] Step 4: It can be executed in parallel with Step 1. It does not need to wait for the SDON manager to send the global optical path configuration policy to the SDON controller, nor does it need to wait for the asynchronous callback instruction of the instruction. The meta-learning algorithm as shown in Formula (1) and Formula (2) above can be executed in parallel.

[0143] Step 5: The SDON manager confirms the completion of the meta-learning algorithm and receives all OXC completion instructions, meaning the reconfiguration plan for the current time step is complete.

[0144] In summary, the meta-learning-based optical path planning method for tidal traffic networks provided in this application example applies meta-learning to optical network optical path planning under the background of tidal services for traffic diversion, and has the following beneficial effects:

[0145] 1. Reduce the average number of hops in the business process;

[0146] 2. Simplify cross-layer communication and reduce global service latency;

[0147] 3. Accelerate the network configuration strategy calculation process and improve network configuration efficiency;

[0148] 4. Improve network service quality and enhance user satisfaction.

[0149] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned meta-learning-based tidal flow network optical path planning method. The computer-readable storage medium can be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, floppy disks, hard disks, removable storage disks, CD-ROMs, or any other form of storage medium known in the art.

[0150] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the aforementioned meta-learning-based tidal flow network optical path planning method.

[0151] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave.

[0152] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0153] In this application, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.

[0154] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to the embodiments of this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for planning a lightpath of a tidal flow network based on meta-learning, characterized in that, Comprise: According to the tidal flow network in the current time step flow distribution data, for the current time step corresponding to the initial light path configuration strategy reinforcement learning agent retraining to obtain the current time step corresponding to the optimal light path configuration strategy; and, in the current time step, according to the current time step corresponding to the initial light path configuration strategy, and in the current time step in the process of reinforcement learning agent retraining generated by reinforcement learning interaction trajectory data set, the preset reinforcement learning learning rate and meta learning learning rate execute the preset meta learning algorithm, to obtain the updated result data of the initial light path configuration strategy corresponding to the current time step, and take the updated result data as the initial light path configuration strategy corresponding to the next time step; Based on the optimal light path configuration strategy, the optical network layer corresponding to the tidal flow network is reconstructed to update the topology structure of the IP layer corresponding to the tidal flow network; Wherein, the meta learning algorithm is shown in formula (1) and formula (2): denotes a temporary variable used to characterize the influence of the current time step t traffic on the first policy update; denotes a preset reinforcement learning learning rate; denotes a preset meta-learning learning rate; a policy parameter representing an initial light path configuration policy corresponding to the current time step; a policy parameter representing a light path configuration policy that is first updated; represents an optical path configuration strategy; represents the initial light path configuration policy corresponding to the current time step; represents the first updated optical path configuration strategy, if w =1, the first updated optical path configuration strategy is the updated result data of the initial optical path configuration strategy corresponding to the current time step, that is, the initial optical path configuration strategy of the next time step t+1; if w =2, take as the optical path configuration strategy to be updated currently, and modify the lower index of the parameter with subscript 1 in formula (2) to the value of the current component w, that is, 2, and modify the lower index of the parameter with subscript 0 to w-1, that is, 1, then execute the modified formula (2), until the calculation of formula (2) corresponding to all components of the window is completed, and take the final as the updated result data of the initial optical path configuration strategy corresponding to the current time step t, that is, the initial optical path configuration strategy of the next time step t+1; represents a trajectory data set generated for the first time in the retraining process of the reinforcement learning intelligent agent corresponding to the initial light path configuration strategy at the current time step, that is, the trajectory data set generated by the reinforcement learning interaction of the initial light path configuration strategy at the current time step is denoted as , and the trajectory data set generated by the reinforcement learning interaction of the policy is denoted as 0 , and the trajectory data set generated by the reinforcement learning interaction of the policy is denoted as representing that the policy is to M 1 a set of trajectory data generated from a second reinforcement learning interaction for a traffic environment representing that the policy is to optimize the objective function of reinforcement learning with data as the strategy data as the strategy representing that the policy is to optimize the objective function of reinforcement learning with data as the strategy data as the strategy M 0 and M 1 represent different traffic environments; denotes the derivative with respect to derivative represents a temporary variable used to characterize the accumulation of t; policy parameters representing the initialization strategy optimized by the neuron learning algorithm at time t.

2. The meta-learning based tidal flow network lightpath planning method of claim 1, wherein, If the current time step is the first time step in the preset time period, the initial light path configuration strategy corresponding to the current time step is randomly generated in advance; If the current time step is not the first time step in the preset time period, the initial light path configuration strategy corresponding to the current time step is obtained by updating the optimal light path configuration strategy corresponding to the last time step based on the meta learning algorithm in the last time step.

3. The meta-learning based tidal flow network lightpath planning method of claim 1, wherein, The optimal light path configuration strategy corresponding to the current time step is sent to a controller asynchronously in the current time step, so that the controller reconstructs the optical network layer corresponding to the tidal flow network based on the optimal light path configuration strategy corresponding to the current time step to update the topology structure of the IP layer corresponding to the tidal flow network. The controller is used to execute the following contents:

4. The meta-learning based tidal flow network lightpath planning method of claim 3, wherein, Receive the optimal light path configuration strategy corresponding to the current time step; If it is determined that the optimal light path configuration strategy corresponding to the last time step is stored locally, according to the comparison result between the optimal light path configuration strategy corresponding to the last time step and the optimal light path configuration strategy corresponding to the current time step, in each optical cross connection device in the optical network layer corresponding to the tidal flow network, determine the optical cross connection device to be reconstructed as the target device, and extract the optical cross connection reconstruction configuration data respectively for each target device from the optimal light path configuration strategy corresponding to the current time step; Send the corresponding optical cross connection reconstruction configuration data to each target device respectively, so that each target device respectively performs optical cross connection reconstruction according to the received optical cross connection reconstruction configuration data and returns the corresponding optical cross connection reconfiguration completion information; Receive the optical cross connection reconfiguration completion information returned by each target device respectively, and if it is determined that the optical cross connection reconfiguration completion information corresponding to all target devices is received, the global reconfiguration completion information is sent. Also include:

5. The meta-learning based tidal flow network lightpath planning method of claim 4, wherein, ​ If global reconfiguration completion information sent by the controller is received and the initial optical path configuration strategy corresponding to the next time step is obtained, it is determined that the optical path planning of the tidal traffic network corresponding to the current time step is completed.

6. The meta-learning based tidal flow network lightpath planning method of claim 1, wherein, The traffic distribution data includes a traffic matrix of the tidal traffic network, wherein each element in the traffic matrix represents the traffic between two nodes in the tidal traffic network, and the communication frequency corresponding to each element is represented by different colors or different numerical parameters of the same color, wherein the numerical parameters include at least one of hue value, saturation value and lightness value. Correspondingly, before the retraining of the reinforcement learning agent for the initial optical path configuration strategy corresponding to the current time step according to the traffic distribution data of the tidal traffic network at the current time step, the method further comprises: Obtaining the traffic matrix of the tidal traffic network at the current time step.

7. A SDON manager comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, The processor executes the computer program to implement the meta-learning-based tidal traffic network optical path planning method according to any one of claims 1 to 6.

8. A SDON controller comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, The SDON controller is in communication connection with an SDON manager, the SDON manager is configured to execute the meta-learning-based tidal traffic network optical path planning method according to any one of claims 1 to 6, and the SDON manager is configured to send the optimal optical path configuration strategy corresponding to the current time step to the SDON controller asynchronously and receive global reconfiguration completion information sent by the SDON controller at the current time step; The processor executes the computer program to implement the following contents: Receiving the optimal optical path configuration strategy corresponding to the current time step; If it is determined that the optimal optical path configuration strategy corresponding to the previous time step is stored locally, according to the comparison result between the optimal optical path configuration strategy corresponding to the previous time step and the optimal optical path configuration strategy corresponding to the current time step, in each optical cross-connect device in the optical network layer corresponding to the tidal traffic network, the optical cross-connect device to be reconfigured at the current time step is determined as a target device, and the optical cross-connect reconfiguration configuration data respectively for each target device is extracted from the optimal optical path configuration strategy corresponding to the current time step; Sending the optical cross-connect reconfiguration configuration data corresponding to each target device to each target device respectively, so that each target device performs optical cross-connect reconfiguration according to the received optical cross-connect reconfiguration configuration data and returns corresponding optical cross-connect reconfiguration completion information; Receiving the optical cross-connect reconfiguration completion information returned by each target device respectively, and if it is determined that the optical cross-connect reconfiguration completion information corresponding to each target device is received, global reconfiguration completion information is sent. 9.A meta-learning based tidal flow network optical path planning system, characterized in that, It comprises: The SDON manager according to claim 7 and the SDON controller according to claim 8; The SDON manager is configured to obtain the traffic distribution data of the tidal traffic network at the current time step from the traffic layer of the tidal traffic network; The processor executes the computer program to implement the meta-learning-based tidal traffic network optical path planning method according to any one of claims 1 to 6. The SDON controller is communicatively connected with each optical cross-connect device in the optical network layer corresponding to the tidal flow network, respectively.

Citation Information

Patent Citations

  • Network traffic engineering method and device based on meta-learning

    CN118194981A

  • Multi-objective reinforcement learning method and system for adaptive network flow optimization

    CN119071236A