Fault repair method for optical network, electronic device and program product
Through multi-agent decision model and path selection and wavelength allocation agent optimized by reinforcement learning, traditional fiber network failure recovery solutions solve the problems of high computing delay and insufficient adaptability in large-scale networks, and achieve fast and effective optical network failure repair.
Patent Information
- Application Number
- CN202510855995.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-08-29
AI Technical Summary
Traditional fiber network failure recovery solutions have high computational delays in large-scale network scenarios, which cannot meet the timeliness requirements, and lack adaptive decision-making capabilities in the face of sudden low-probability failures or unknown topological scenarios, and have low recovery efficiency.
The multi-agent decision model is adopted, and the path selection agent and the wavelength allocation agent are responsible for path selection and wavelength allocation respectively. Combined with reinforcement learning and optimization decision strategies, a fast and feasible recovery plan is generated, and adaptability is improved through online training and updates.
It improves the efficiency and effect of optical network failure recovery, can quickly deal with sudden low-probability failures and unknown topological scenarios, ensures the rationality and accuracy of recovery paths and wavelength allocations, and reduces transmission errors.
Smart Images

Figure CN120568232A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a fault repair method, electronic equipment, storage medium and program product for an optical network. Background Art
[0002] As the core carrier of modern communications infrastructure, the reliability of fiber optic networks directly impacts the continuity of internet services. Fiber failures (such as physical damage or natural disasters) can disrupt IP links and degrade network performance. Traditional recovery solutions have significant limitations in terms of dynamic resource allocation and real-time response. First, the decision space for fault recovery in large-scale network scenarios experiences a combinatorial explosion, leading to a surge in computational latency and an inability to meet timeliness requirements. Second, traditional solutions lack adaptive decision-making capabilities in the face of sudden, low-probability failures or unknown topology scenarios, significantly reducing recovery efficiency. Summary of the Invention
[0003] The present disclosure provides a fault repair method for an optical network, an electronic device, a storage medium, and a program product.
[0004] According to one aspect of the present disclosure, a fault repair method for an optical network is provided, comprising: in response to a fiber failure event in the optical network, determining a failed IP link in the optical network; using the failed IP link as input to a multi-agent decision model, and having the multi-agent decision model output a path recovery plan corresponding to the failed IP link, the path recovery plan comprising a recovery path determined for the failed IP link by a path selection agent in the multi-agent decision model, and a wavelength allocated to the recovery path by a wavelength allocation agent in the multi-agent decision model; updating the IP link capacity according to the path recovery plan; and redistributing traffic on the IP link according to the updated IP link capacity to achieve network fault repair.
[0005] According to one aspect of the technical solution, a failed IP link in an optical network is identified in response to a fiber failure event in the optical network. This failed IP link is used as input to a multi-agent decision model, which then outputs a path recovery plan corresponding to the failed IP link. The path recovery plan includes a recovery path determined for the failed IP link by a path selection agent in the multi-agent decision model and a wavelength assigned to the recovery path by a wavelength allocation agent in the multi-agent decision model. Subsequently, the IP link capacity is updated based on the path recovery plan, and traffic is redistributed across the IP link based on the updated IP link capacity to repair the network failure.
[0006] In this way, the path selection agent and wavelength allocation agent are responsible for path selection and wavelength allocation respectively, which reduces the decision-making complexity of each agent and allows them to focus on their respective tasks, thereby improving recovery efficiency and effectiveness.
[0007] Moreover, the intelligent agent can learn the optimal behavior strategy through continuous interaction with the environment, making it more adaptable to the environment and ensuring its adaptive decision-making ability even in the face of sudden low-probability failures or unknown topology scenarios.
[0008] In some embodiments of the present disclosure, the failed IP link is used as the input of a multi-agent decision model, and the multi-agent decision model outputs a path recovery plan corresponding to the failed IP link, including: determining the state space of a path selection agent based on the global wavelength occupancy information of the optical network, a set of candidate optical paths corresponding to the IP link, and a wavelength occupancy status corresponding to each candidate optical path; based on the state space of the path selection agent, the path selection agent selects one of the candidate optical paths corresponding to the failed IP link as a recovery path from the set of candidate optical paths corresponding to the failed IP link; and a wavelength allocation agent allocates a wavelength to the recovery path, and the state space of the wavelength allocation agent includes the wavelength occupancy status corresponding to the recovery path and the global wavelength occupancy information.
[0009] According to the technical solution of this embodiment, by constructing state spaces for the path selection agent and the wavelength assignment agent and making decisions based on these state spaces, it is possible to comprehensively consider global wavelength occupancy information and the wavelength occupancy status of candidate optical paths, ensuring the rationality of the selected restoration path and assigned wavelength. This helps to quickly generate effective path restoration plans and improve the efficiency of network fault recovery.
[0010] In some embodiments of the present disclosure, according to the state space of the path selection agent, the path selection agent selects one of the candidate optical paths corresponding to the failed IP link as a recovery path, including: determining a candidate optical path with an idle wavelength from the candidate optical path set corresponding to the failed IP link; according to the state space of the path selection agent, the path selection agent selects one of the candidate optical paths with an idle wavelength corresponding to the failed IP link as a recovery path corresponding to the failed IP link.
[0011] According to the technical solution of this embodiment, by pre-screening candidate optical paths with idle wavelengths, the selection range of the path selection agent can be narrowed, so that the path selection agent can find a suitable recovery path more quickly, thereby improving the efficiency of fault recovery.
[0012] In some embodiments of the present disclosure, allocating wavelengths to the restoration path by a wavelength allocation agent includes: determining an unavailable wavelength corresponding to the restoration path based on an occupancy status of the wavelength corresponding to the restoration path; and shielding the allocation action of the unavailable wavelength when the wavelength allocation agent allocates wavelengths to the restoration path.
[0013] According to the technical solution of this embodiment, by shielding the allocation actions of unavailable wavelengths, the action space of the wavelength allocation agent is narrowed, enabling it to find available wavelengths more quickly and improving the efficiency of wavelength allocation. Shielding the allocation actions of unavailable wavelengths can also avoid incorrect allocations, ensure the accuracy of wavelength allocation, and reduce transmission errors caused by wavelength conflicts.
[0014] In some embodiments of the present disclosure, the method further includes: generating a training data set based on the collected fault scenario data; training the path selection agent and wavelength allocation agent in the multi-agent decision model according to the training data set, and during the training process, giving a positive reward when an available wavelength is successfully allocated on the path, and giving a reward of zero if no available wavelength is found.
[0015] According to the technical solution of this embodiment, through training based on actual fault scenario data, the path selection agent and the wavelength allocation agent can learn effective recovery strategies to improve adaptability and reliability to various fault scenarios.
[0016] In some embodiments of the present disclosure, when training the path selection agent and the wavelength allocation agent in the multi-agent decision model, the wavelengths of the failed IP links are restored one by one based on a random order.
[0017] According to the technical solution of this embodiment, the randomized allocation order strategy can promote exploration and avoid falling into local optimality in the early stages of training. Randomizing the processing order helps avoid performance degradation caused by fixed patterns and improves the network's recovery robustness.
[0018] In some embodiments of the present disclosure, the failed IP link is used as the input of a multi-agent decision model, and the multi-agent decision model outputs a path recovery plan corresponding to the failed IP link, including: when there are multiple failed IP links, the multi-agent decision model restores the wavelength of each failed IP link one by one based on a random order, and outputs a corresponding path recovery plan.
[0019] According to the technical solution of this embodiment, when multiple wavelengths need to be restored, the wavelengths to be restored can be restored one by one in a random order, thereby reducing the computational overhead.
[0020] In some embodiments of the present disclosure, the method further includes: collecting newly added fault scenario data during the online operation phase of the multi-agent decision model; updating the training data set according to the newly added fault scenario data; and optimizing the parameters of the multi-agent decision model based on the updated training data set.
[0021] According to the technical solution of this embodiment, the adaptability and generalization ability of the path recovery solution generated by the multi-agent decision model can be improved.
[0022] In some embodiments of the present disclosure, IP link traffic is redistributed based on the updated IP link capacity, including: establishing a multi-commodity network flow optimization problem based on known traffic demand and a set of candidate paths between IP node pairs; solving the multi-commodity network flow optimization problem based on a solver to determine a traffic allocation strategy based on the updated IP link capacity.
[0023] According to the technical solution of this embodiment, by establishing and solving a multi-commodity network flow optimization problem, we can obtain the optimal traffic allocation strategy based on the updated IP link capacity. This ensures the rational distribution of network traffic, improves the utilization of link resources, reduces network congestion, and improves overall transmission efficiency.
[0024] According to another aspect of the present disclosure, a fault repair device for an optical network is provided, comprising: a determination module for determining a failed IP link in the optical network in response to a fiber failure event in the optical network; an inference module for taking the failed IP link as an input of a multi-agent decision model, and having the multi-agent decision model output a path recovery plan corresponding to the failed IP link, the path recovery plan comprising a recovery path determined for the failed IP link by a path selection agent in the multi-agent decision model, and a wavelength allocated to the recovery path by a wavelength allocation agent in the multi-agent decision model; an update module for updating the IP link capacity according to the path recovery plan; and a processing module for redistributing traffic on the IP link according to the updated IP link capacity to achieve network fault repair.
[0025] According to another aspect of the present disclosure, an electronic device is provided, comprising: a memory storing execution instructions; and a processor executing the execution instructions stored in the memory, so that the processor executes the fault repair method for an optical network according to any embodiment of the present disclosure.
[0026] According to another aspect of the present disclosure, a readable storage medium is provided, wherein the readable storage medium stores execution instructions, and when the execution instructions are executed by a processor, the method for repairing a fault in an optical network according to any embodiment of the present disclosure is implemented.
[0027] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the fault repair method for an optical network according to any embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings illustrate exemplary embodiments of the present disclosure and together with the description serve to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification.
[0029] Figure 1 A schematic diagram showing the mapping relationship between the optical layer topology and the IP layer topology is shown.
[0030] Figure 2 A schematic diagram showing the occurrence of an optical fiber failure event is shown.
[0031] Figure 3 A schematic flow chart of a fault repair method for an optical network according to an embodiment of the present disclosure is shown.
[0032] Figure 4 A flow chart of step S320 in a fault repair method for an optical network according to an embodiment of the present disclosure is shown.
[0033] Figure 5 A flow chart of step S322 in a fault repair method for an optical network according to an embodiment of the present disclosure is shown.
[0034] Figure 6 A flow chart of step S323 in a fault repair method for an optical network according to an embodiment of the present disclosure is shown.
[0035] Figure 7 A schematic diagram of a process for training a multi-agent decision model included in a fault repair method for an optical network according to an embodiment of the present disclosure is shown.
[0036] Figure 8 A schematic diagram of the process of online training of a multi-agent decision model included in a fault repair method for an optical network according to an embodiment of the present disclosure is shown.
[0037] Figure 9 A flow chart of step S340 in a fault repair method for an optical network according to an embodiment of the present disclosure is shown.
[0038] Figure 10 A schematic structural block diagram of an optical network restoration system based on multi-agent reinforcement learning according to an embodiment of the present disclosure is shown.
[0039] Figure 11 A network topology diagram used in a comparative experiment according to an embodiment of the present disclosure is shown.
[0040] Figure 12 A schematic diagram showing a comparison of maximum achievable throughputs of different network recovery methods in a comparative experiment according to an embodiment of the present disclosure is shown.
[0041] Figure 13 A schematic diagram showing a comparison of the computing time of different network recovery methods in a comparative experiment according to an embodiment of the present disclosure is shown.
[0042] Figure 14 A schematic diagram showing performance comparison of different network recovery methods in a comparative experiment according to an embodiment of the present disclosure is shown.
[0043] Figure 15 A schematic structural block diagram of a fault repair device for an optical network according to an embodiment of the present disclosure is shown.
[0044] Figure 16 A schematic structural block diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0045] The present disclosure is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are intended only to illustrate the relevant content and are not intended to limit the present disclosure. It should also be noted that, for ease of description, only the portions relevant to the present disclosure are shown in the accompanying drawings.
[0046] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present disclosure can be combined with each other. The technical solution of the present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0047] Fiber optic backbone networks form the infrastructure for long-distance communications in wide area networks (WANs), supporting high-capacity data transmission. Figure 1As shown, a high-capacity IP link maps data streams onto an optical path through a series of interconnected subsystems. For example, at node C, the router interface transmits electrical signals to the Optical Transport Network (OTN) switches, which aggregate traffic and perform sub-wavelength processing before forwarding the signal to the optical transceiver. The optical transceiver then converts the electrical signal into an optical carrier and assigns the wavelength to the spectral slot according to the grid spacing standard. These wavelengths are then combined through a multiplexer (MUX) and routed through a reconfigurable optical add / drop multiplexer (ROADM). The above process establishes the topological mapping between the IP layer and the optical layer, and the actual traffic distribution occurs at the IP link layer. Figure 1 Three IP links are shown: AB, AC, and AD.
[0048] Whenever a fiber failure occurs, two key steps are required to ensure optimal resource utilization. First, an alternative optical path must be restored for the affected IP link, allocating the appropriate wavelength. However, due to resource constraints, complete optical layer restoration is often not feasible, necessitating an optimized restoration strategy to maximize restored traffic. Consequently, traffic distribution at the IP layer must also be reconfigured to further improve overall restoration efficiency.
[0049] Early optical network recovery tools relied on resource redundancy strategies based on pre-set backup paths. While this approach ensured reliability, it also resulted in low resource utilization. Subsequently, fault-aware Traffic Engineering (TE) methods emerged, dynamically reallocating IP layer traffic but often overlooking underutilized wavelength resources in the optical layer. In contrast, optical restoration methods improve resource utilization by reconfiguring wavelengths, but because they fail to consider specific traffic demands, they are prone to suboptimal solutions.
[0050] A more attractive approach is to jointly consider wavelength allocation and traffic distribution by solving a joint modeling problem for the IP / optical layer. Currently, mixed integer linear programming (MILP) is often used to solve this problem. To improve the scalability of MILP, a two-stage traffic engineering model is proposed: multiple feasible optical layer repair solutions are generated in the offline phase, and the optimal solution is selected from these candidate solutions in the online phase to meet traffic demand. Although this method improves recovery efficiency through pre-computation, it still has several key limitations in practical applications: 1. Poor scalability Because this method relies on a pre-calculated probability distribution of fiber failures, the number of potential failure scenarios increases exponentially as the number of fibers in the network increases, making enumerating all scenarios computationally infeasible. Therefore, the aforementioned solution primarily targets high-probability failures. When encountering low-probability failures, a two-stage traffic engineering process must be performed after the failure. For example, in a 70-node optical network topology, this process can take up to 2.3 hours to calculate a recovery plan. This problem is further exacerbated without an accurate failure probability model.
[0051] 2. Unclear optical layer recovery configuration The above method uses the relaxed linear programming solution to round off candidate wavelength solutions, but does not propose an accurate optical layer recovery configuration. Figure 2 As shown in the figure, in the fault scenario where the AD fiber is broken, both IP links AC and AD are affected. Assuming that AC requires two wavelengths for recovery and AD also requires two wavelengths, then one candidate solution in the above method may allocate two wavelengths to AC and only one wavelength to AD. However, this solution does not clearly indicate which specific wavelengths are used for each candidate recovery path. λ 1 、 λ 2 、 λ 3 、 λ 4 For example, it does not specify whether AD uses wavelength via path ABCD. λ 2 recover.
[0052] 3. Poor feasibility of deployment of candidate solutions To obtain a deployable restoration strategy, an integer linear programming (ILP) model is typically constructed with the goal of maximizing the number of restored wavelengths. However, while the candidate restoration solutions generated by these methods may pass initial feasibility checks, they often fail to map to valid final configurations when solved using the ILP. This limitation means that a large number of seemingly feasible candidate solutions are actually undeployable, especially in large-scale networks, significantly limiting the practical application of these methods.
[0053] To this end, this disclosure proposes the following technical solution: a path selection agent and a wavelength assignment agent are responsible for path selection and wavelength assignment, respectively. This reduces the decision-making complexity of each agent, allowing them to focus on their respective tasks, thereby improving recovery efficiency and effectiveness. Furthermore, the agents can continuously learn optimal behavior strategies through interaction with the environment, giving them greater environmental adaptability. Even in the face of sudden, low-probability failures or unknown topology scenarios, their adaptive decision-making capabilities are maintained, while also ensuring the feasibility of the determined path recovery solution.
[0054] Figure 3 A schematic flow chart of a fault repair method for an optical network according to an embodiment of the present disclosure is shown.
[0055] like Figure 3 As shown, the fault repair method for an optical network includes at least steps S310 to S340, which are described in detail as follows.
[0056] In step S310, in response to a fiber failure event in an optical network, a failed IP link in the optical network is determined.
[0057] Among them, a fiber failure event can be a transmission interruption event caused by physical damage to the fiber link in the optical network, natural disasters, or aging of infrastructure.
[0058] A failed IP link can be one that is unable to transmit data due to a fiber failure. These links originally relied on the failed fiber to carry data traffic, but due to the fiber failure, the link loses its transmission capacity, requiring recovery through alternative paths.
[0059] In this embodiment, a fault detection mechanism, such as optical power detection or link status detection, can be deployed in the optical network to monitor the operating status of the optical fiber link in real time. When the optical power of the optical fiber link drops below a set threshold or the link status detection fails, the system triggers a fiber fault event.
[0060] When a fiber fault is detected, the specific location of the fault can be determined by analyzing the optical power monitoring data and / or link status information. For example, by detecting the point where the optical power drops, the location of the fiber break can be precisely located.
[0061] Then, using the mapping relationship between the optical layer topology and the IP layer topology of the current optical network, it is possible to determine which IP links are dependent on the failed fiber. It should be understood that this mapping relationship records the transmission path and wavelength resources used by each IP link in the optical layer. For example, if fiber AD fails, the above mapping relationship can be queried to identify IP link AD and other IP links that depend on this fiber. In this way, the failed IP links affected by the fiber failure event in the optical network can be accurately determined.
[0062] This allows for quick and accurate locating of the impact of fiber failures on the optical network. Once the failed IP link is precisely identified, targeted recovery strategies can be developed, improving optical network recovery efficiency and service quality.
[0063] In step S320, the failed IP link is used as the input of the multi-agent decision model, and the multi-agent decision model outputs a path recovery plan corresponding to the failed IP link. The path recovery plan includes a recovery path determined for the failed IP link by the path selection agent in the multi-agent decision model, and a wavelength allocated to the recovery path by the wavelength allocation agent in the multi-agent decision model.
[0064] The multi-agent decision model can be a pre-trained multi-agent reinforcement learning-based model that includes a path selection agent and a wavelength assignment agent. The model generates a corresponding path recovery solution for a failed IP link through collaborative decision-making between the agents.
[0065] A path recovery plan can be a recovery strategy generated by a multi-agent decision model for a failed IP link. It includes a corresponding recovery path and assigned wavelengths. The path selection agent determines the recovery path, while the wavelength assignment agent allocates available wavelengths to the recovery path.
[0066] In this embodiment, the determined failed IP link information can be used as input to the multi-agent decision model. In one example, the failed IP link information can include the failed IP link's identifier, traffic demand, or dependent fiber path. After receiving the failed IP link information, the path selection agent in the multi-agent decision model can determine a recovery path for the failed IP link, and the wavelength allocation agent can assign a wavelength to the recovery path based on the determined recovery path.
[0067] In this way, failed IP link information is used as input for the multi-agent decision-making model. It should be understood that the multi-agent reinforcement learning model is trained on a large amount of data in the offline phase, enabling rapid decision-making. In the online phase, it can quickly generate path restoration plans, improving fault recovery efficiency. Furthermore, the path selection agent and wavelength allocation agent work together to ensure the proper allocation of recovery paths and wavelength resources, avoiding resource conflicts and waste.
[0068] Furthermore, the multi-agent decision-making model continuously optimizes its decision-making strategy through reinforcement learning, improving the accuracy and reliability of path restoration solutions. Compared to traditional methods, reinforcement learning models can better adapt to network dynamics, generate more optimal restoration solutions, and reduce errors and failures during the restoration process. For example, in complex network topologies, reinforcement learning models can more effectively identify feasible restoration paths and wavelength combinations.
[0069] In step S330, the IP link capacity is updated according to the path recovery solution.
[0070] The IP link capacity may refer to the amount of data transmission that each IP link can carry in the IP layer network, which is affected by the number of wavelengths allocated at the optical layer and the transmission rate of each wavelength.
[0071] In this embodiment, in an optical network, the IP link capacity of unaffected IP links in the event of a fiber failure can be calculated based on network configuration or historical data. For each newly determined recovery path and its corresponding wavelength, the IP link capacity corresponding to each recovery path can be determined. For example, the IP link capacity corresponding to the new IP link (i.e., the recovery path) can be calculated based on the number of wavelengths corresponding to each recovery path and the transmission rate corresponding to each wavelength. The calculated new IP link capacity is then updated in the link state information at the IP layer.
[0072] In this way, updating IP link capacity based on the path restoration plan enables traffic engineering at the IP layer to accurately understand changes in optical layer resources. This helps fully utilize restored link resources during subsequent traffic allocation, optimize traffic distribution, and improve network resource utilization.
[0073] In step S340, traffic on the IP link is redistributed according to the updated IP link capacity to achieve network fault repair.
[0074] Traffic redistribution refers to readjusting the transmission path and distribution ratio of data traffic in the IP layer according to changes in network status, so as to optimize the utilization of network resources and improve transmission efficiency.
[0075] In this embodiment, after updating the IP link capacity, the transmission paths and allocation ratios of data traffic at the IP layer can be redefined based on the updated IP link capacity, thereby ensuring a reasonable distribution of network traffic, fully utilizing restored link resources, and improving overall resource utilization. This traffic allocation strategy is then applied to the optical network to achieve network fault recovery.
[0076] Therefore, based on Figure 3 The illustrated embodiment identifies a failed IP link in an optical network in response to a fiber failure event. This failed IP link is used as input to a multi-agent decision model, which then outputs a path recovery plan corresponding to the failed IP link. This path recovery plan includes a recovery path determined for the failed IP link by a path selection agent in the multi-agent decision model and a wavelength assigned to the recovery path by a wavelength assignment agent in the multi-agent decision model. Subsequently, the IP link capacity is updated based on the path recovery plan, and traffic is redistributed across the IP link based on the updated IP link capacity, thereby repairing the network failure.
[0077] In this way, the path selection agent and wavelength assignment agent are responsible for path selection and wavelength assignment, respectively. This reduces the decision-making complexity of each agent, allowing them to focus on their respective tasks, thereby improving recovery efficiency and effectiveness. Furthermore, the agents can learn optimal behavior strategies through continuous interaction with the environment, making them more adaptable to the environment. This ensures their adaptive decision-making capabilities even in the face of sudden, low-probability failures or unknown topology scenarios. Furthermore, the modular design facilitates future expansion. For example, the path selection agent can be combined with a neural network to dynamically generate more optimal routing strategies.
[0078] Figure 4 FIG. 1 shows a flow chart of step S320 in a method for repairing a fault in an optical network according to an embodiment of the present disclosure. Figure 4 As shown, step S320 at least includes steps S321 to S323, which are explained in detail as follows.
[0079] In step S321, the state space of the path selection agent is determined according to the global wavelength occupancy information of the optical network, the set of candidate optical paths corresponding to the IP link, and the wavelength occupancy status corresponding to each candidate optical path.
[0080] Global wavelength occupancy information refers to the usage of all wavelengths in the entire optical network, including the wavelength occupancy status of each optical fiber link. This enables intelligent agents to make globally aware decisions, reduce conflicts, and improve efficiency.
[0081] The candidate lightpath set can be a set of possible optical transmission paths pre-calculated for each IP link, used as a recovery path selection in the event of a failure. Each candidate lightpath is represented by the sequence of optical fibers it passes through, thus providing structured path options for the agent.
[0082] The wavelength occupancy status corresponding to the candidate optical path may be the usage status of each wavelength on the candidate optical path, which may indicate which wavelengths are available and which wavelengths are occupied, and thus may be used to evaluate the current availability of the candidate optical path.
[0083] In this embodiment, global wavelength occupancy information for the entire optical network, a set of candidate optical paths corresponding to IP links, and the wavelength occupancy status of each candidate optical path can be obtained to construct a state space for the path selection agent. For example, the state space of the path selection agent may include a matrix representing global wavelength occupancy, a list of all candidate optical paths, and a wavelength occupancy vector corresponding to each candidate optical path.
[0084] In step S322, according to the state space of the path selection agent, the path selection agent selects one of the candidate optical paths corresponding to the failed IP link as a recovery path.
[0085] In this embodiment, the path selection agent can understand the current state of the optical network through the currently constructed state space. Based on this state space, the path selection agent can use a preset algorithm to select the optimal recovery path from a set of candidate optical paths corresponding to the failed IP link. The path selection agent can base its selection on factors such as the number of available wavelengths and path length. The path selection agent can output the selected recovery path as a decision result for subsequent use by the wavelength allocation agent.
[0086] In one embodiment, the path selection agent may select the corresponding restoration path based on the K shortest path algorithm. It should be noted that in other embodiments, the path selection agent may also perform path selection based on other algorithms, which is not particularly limited.
[0087] In step S323, a wavelength allocation agent allocates a wavelength to the restoration path, and the state space of the wavelength allocation agent includes the wavelength occupancy state corresponding to the restoration path and the global wavelength occupancy information.
[0088] In this embodiment, the wavelength allocation agent's state space consists of the wavelength occupancy status of the restoration path and global wavelength occupancy information. This helps the agent understand which wavelengths are available on the selected restoration path. Based on its state space, the wavelength allocation agent can consider factors such as wavelength continuity and collision avoidance when selecting the appropriate wavelength for the restoration path.
[0089] In addition, the allocated wavelengths can be marked as occupied, and the global wavelength occupancy information and the wavelength occupancy status of the corresponding optical path can be updated so that they can reflect the latest resource usage and provide an accurate data basis for subsequent processing (such as the recovery of new failed IP paths, etc.).
[0090] By constructing the state spaces of the path selection agent and the wavelength assignment agent and making decisions based on these state spaces, we can comprehensively consider global wavelength occupancy information and the wavelength occupancy status of candidate optical paths, ensuring the rationality of the selected restoration paths and assigned wavelengths. This helps to quickly generate effective path restoration solutions and improve the efficiency of network fault recovery.
[0091] Figure 5 FIG. 1 shows a flow chart of step S322 in a method for repairing a fault in an optical network according to an embodiment of the present disclosure. Figure 5 As shown, step S322 at least includes steps S3221 to S3222, which are described in detail as follows.
[0092] In step S3221, a candidate optical path having an idle wavelength is determined from the set of candidate optical paths corresponding to the failed IP link.
[0093] The idle wavelength may be a wavelength that is not occupied on the candidate optical path and can be used for new data transmission.
[0094] In this embodiment, for the candidate optical path set corresponding to the failed IP link, the wavelength occupancy status of each candidate optical path can be obtained. Then, the wavelength occupancy status of each candidate optical path is checked to determine which candidate optical paths have idle wavelengths for subsequent path selection.
[0095] In step S3222, based on the state space of the path selection agent, the path selection agent selects one of the candidate optical paths with idle wavelengths corresponding to the failed IP link as the recovery path corresponding to the failed IP link.
[0096] In this embodiment, based on the state space of the path selection agent, the path selection agent can select a recovery path from the candidate optical paths with idle wavelengths in the set of candidate optical paths corresponding to the failed IP link. This pre-screening of candidate optical paths with idle wavelengths narrows the path selection agent's selection range, enabling the path selection agent to more quickly find a suitable recovery path and improve fault recovery efficiency.
[0097] Figure 6FIG. 1 shows a flow chart of step S323 in a method for repairing a fault in an optical network according to an embodiment of the present disclosure. Figure 6 As shown, step S323 at least includes steps S3231 to S3232, which are described in detail as follows.
[0098] In step S3231, the unavailable wavelength corresponding to the restoration path is determined according to the wavelength occupancy status corresponding to the restoration path.
[0099] In this embodiment, after the restoration path is determined by the path selection agent, the unavailable wavelength on the restoration path may be determined according to the wavelength occupancy status of the restoration path.
[0100] In step S3232, when the wavelength assignment agent assigns wavelengths to the restoration path, the assignment action of unavailable wavelengths is shielded.
[0101] In this embodiment, when the wavelength allocation agent allocates wavelengths to the recovery path based on its own state space, it can block allocation actions for unavailable wavelengths. In other words, the wavelength allocation agent can select an available wavelength from the filtered list for allocation. By blocking allocation actions for unavailable wavelengths, the wavelength allocation agent's action space is reduced, enabling it to find available wavelengths more quickly and improving wavelength allocation efficiency. Blocking allocation actions for unavailable wavelengths prevents incorrect allocations, ensures accurate wavelength allocation, and reduces transmission errors caused by wavelength conflicts.
[0102] Figure 7 FIG. 1 shows a flow chart of a multi-agent decision model training method included in a fault repair method for an optical network according to an embodiment of the present disclosure. Figure 7 As shown, training the multi-agent decision model includes at least steps S350 to S360, which are described in detail below.
[0103] In step S350, a training data set is generated based on the collected fault scenario data.
[0104] The fault scenario data may refer to relevant data when a fiber fault event occurs in an optical network, including information such as the time and location of the fault, the affected IP links, and the wavelength occupancy status.
[0105] The training data set may refer to organized and pre-processed fault scenario data, which is used to train the multi-agent decision model so that it can learn effective path selection and wavelength allocation strategies.
[0106] In this embodiment, historical fault scenario data can be collected from the optical network's SDN (Software Defined Network) controller and network management system, including information such as the time and location of fiber failure events, affected IP links, and wavelength occupancy status. This collected fault scenario data is then cleaned, normalized, and annotated to construct a training dataset. For example, each fault scenario can be converted into a state-action-reward sample that includes the network status, failed IP link, recovery path, and wavelength allocation.
[0107] In step S360, the path selection agent and the wavelength allocation agent in the multi-agent decision model are trained based on the training data set. During the training process, a positive reward is given when an available wavelength is successfully allocated on the path, and a reward is zero if no available wavelength is found.
[0108] In this embodiment, before model training is performed, the policy networks of the path selection agent and the wavelength assignment agent may be initialized. In one example, random initialization may be used.
[0109] During model training, the network state from the training dataset is fed into the path selection agent to obtain the recovery path it selects. The recovery path is then fed into the wavelength assignment agent to obtain the wavelength it assigns. A reward is calculated based on the wavelength assignment result: a positive reward is given if an available wavelength is successfully assigned; a zero reward is given if no available wavelength is found. The agent's policy network is then updated based on the reward signal to optimize its decision-making strategy. This training process is repeated until the agent's strategy converges, meaning its decision-making achieves stable and effective performance on the training dataset.
[0110] In this way, through training based on actual fault scenario data, the path selection agent and wavelength allocation agent can learn effective recovery strategies and improve adaptability and reliability to various fault scenarios.
[0111] Notably, in the aforementioned reinforcement learning framework, the primary goal is to maximize the number of successfully restored wavelengths without directly incorporating traffic matrix information. Unlike reinforcement learning, which introduces weights into actions, the disclosed solution treats all restored wavelengths as equivalent, reflecting its non-prioritized nature. While this strategy decouples optical and IP layer modeling during the recovery process, potentially resulting in suboptimal global recovery, it has been proven to be effective in practice.
[0112] based on Figure 7In some embodiments of the present disclosure, as shown in the embodiment, when training the path selection agent and the wavelength allocation agent in the multi-agent decision model, the wavelengths of the failed IP links are restored one by one based on a random order.
[0113] It should be noted that each failed IP link may have one or more wavelengths requiring restoration. When a failed IP link has multiple wavelengths requiring restoration, a random wavelength restoration order can be generated for each training sample. This random order determines which wavelength to restore first. This allows all wavelength restorations to be considered equivalent, reflecting the non-priority nature of the restoration process.
[0114] Furthermore, this recovery method avoids the local optimal solution that can be achieved by using a fixed recovery order, enhancing the adaptability and robustness of the agent in different failure scenarios. Furthermore, randomizing the order helps explore different recovery paths and wavelength allocation combinations, increasing the diversity of recovery strategies and helping to find a more optimal recovery solution.
[0115] In other implementations, a randomized restoration order may be generated based on the wavelengths to be restored corresponding to all failed IP paths, and the wavelengths may be restored one by one based on the randomized restoration order. Those skilled in the art may select a corresponding implementation method based on actual implementation needs, and this is not particularly limited.
[0116] This randomized allocation order strategy promotes exploration and avoids falling into local optima early in training. It should be understood that when each wavelength on an IP link is restored, the wavelength usage status on the optical fiber is updated. Previous allocations can significantly impact subsequent restoration. Improper processing order can lead to unavailable paths or wavelengths, resulting in restoration failure. Therefore, the allocation order is crucial for path and wavelength selection. Randomizing the processing order helps avoid performance degradation caused by fixed patterns and improves network recovery robustness.
[0117] Based on the above embodiment, in some implementations of the present disclosure, the failed IP link is used as an input to a multi-agent decision model, and the multi-agent decision model outputs a path recovery solution corresponding to the failed IP link, including: In the case that there are multiple failed IP links, the multi-agent decision model restores the wavelength of each failed IP link one by one based on a random order and outputs a corresponding path restoration solution.
[0118] In this embodiment, during the practical application of the multi-agent decision-making model, when a fiber failure occurs, one or more failed IP links may exist, and a failed IP link may have one or more wavelengths that need to be restored. Therefore, when multiple wavelengths need to be restored, the wavelengths to be restored can be restored one by one in a random order. Compared to restoring all wavelengths at once, this approach reduces computational overhead and eliminates the need to encode placeholder information for IP links with different failed wavelengths.
[0119] In some embodiments of the present disclosure, an action blocking mechanism can be introduced during multi-agent decision-making model training to improve training efficiency. It should be understood that during optical layer recovery, some paths or wavelengths may be unavailable due to insufficient resources. For example, if no wavelength is available on a path, it is blocked before path selection. If a path is selected but a wavelength is unavailable, this action is also blocked before wavelength allocation. By retaining only valid action options, this mechanism can effectively reduce the search space and avoid exploring invalid actions, significantly improving training efficiency.
[0120] Figure 8 FIG. 1 shows a flow chart of online training of a multi-agent decision model included in a fault repair method for an optical network according to an embodiment of the present disclosure. Figure 8 As shown, the online training of the multi-agent decision model includes at least steps S370 to S390, which are described in detail as follows.
[0121] In step S370, during the online operation phase of the multi-agent decision model, newly added fault scenario data is collected.
[0122] In this embodiment, during the online operation stage of the multi-agent decision model, the operating status of the optical network can be monitored in real time, thereby collecting new optical fiber fault events and related data (such as the event, location, affected IP link, wavelength occupancy status, etc. of the optical fiber fault), thereby obtaining corresponding fault scenario data.
[0123] In step S380, the training data set is updated according to the newly added fault scenario data.
[0124] In this embodiment, fault scenario data collected during the online operation phase can be compared with fault scenario data in the training dataset to identify newly emerging fault scenario data. For example, during the online operation phase of the multi-agent decision model, if a fault scenario data item is collected and the affected IP link does not appear in the training dataset, then this fault scenario data can be identified as newly emerging fault scenario data. In this case, the newly emerging fault scenario data can be added to the original training dataset, thereby increasing the diversity of the training dataset and ensuring its data quality.
[0125] In step S390, the parameters of the multi-agent decision model are optimized based on the updated training data set.
[0126] In this embodiment, when the amount of additional fault scenario data reaches a certain number (e.g., 500), or the multi-agent decision-making model has been running for a certain period of time, the multi-agent decision-making model can be fine-tuned or retrained based on the updated training dataset to achieve parameter optimization. This improves the adaptability and generalization of the path restoration solution generated by the multi-agent decision-making model.
[0127] Figure 9 FIG. 1 shows a flow chart of step S340 in a method for repairing a fault in an optical network according to an embodiment of the present disclosure. Figure 9 As shown, step S340 at least includes steps S341 to S342, which are described in detail as follows.
[0128] In step S341 , a multi-commodity network flow optimization problem is established based on known traffic demands and a set of candidate paths between IP node pairs.
[0129] The multi-commodity network flow optimization problem is a mathematical optimization model used to optimize resource allocation in a multi-commodity flow network, ensuring optimal overall network performance while satisfying various constraints. In this disclosure, this problem can be applied to optimize traffic distribution at the IP layer in optical networks, ensuring optimal traffic distribution over updated IP link capacity.
[0130] In this embodiment, in the IP layer network, the traffic between different source-destination node pairs can be defined as different commodity flows. Next, the optimization objectives are defined, such as minimizing network congestion and maximizing traffic transmission efficiency. Then, constraints are established, such as ensuring that the traffic allocation of each IP link does not exceed its updated capacity; ensuring that the traffic requirements of each commodity flow are met; ensuring that traffic is only transmitted through predefined candidate paths, etc. This establishes a multi-commodity network flow optimization problem corresponding to the current optical network. In step S342, the multi-commodity network flow optimization problem is solved based on the solver to determine a flow distribution strategy based on the updated IP link capacity.
[0131] In this embodiment, an efficient mathematical optimization solver, such as Gurobi or CPLEX, can be used to solve the multi-commodity network flow optimization problem established above. The updated IP link capacity, the flow requirements of the commodity flows, and the path set are input into the solver. The solver is run to calculate an optimal flow allocation strategy that satisfies all constraints based on the updated IP link capacity. This flow allocation strategy includes the flow allocation ratio or specific flow value for each commodity flow on each path.
[0132] Then, the traffic distribution strategy obtained is applied to the actual network configuration (for example, according to the traffic distribution strategy, the router's forwarding table is updated and the traffic forwarding path and ratio are adjusted), thereby realizing network fault repair.
[0133] By establishing and solving a multi-commodity network flow optimization problem, we can obtain the optimal traffic allocation strategy based on the updated IP link capacity. This ensures the proper distribution of network traffic, improves link resource utilization, reduces network congestion, and enhances overall transmission efficiency.
[0134] Based on the technical solutions of the above embodiments, a specific application scenario of the embodiments of the present application is introduced below: In some embodiments of the present disclosure, a multi-agent reinforcement learning-based optical network recovery system (LBOR) is provided, which designs a collaborative strategy for path selection and wavelength allocation and builds an end-to-end fault recovery framework for large-scale topologies, thereby solving the following three technical problems.
[0135] 1. By using a large amount of historical network status and fault scenario data for training in the offline phase, the model can quickly generate recovery strategies in the online phase without exhaustively enumerating all fault combinations, fundamentally solving the scalability problem of pre-calculation methods.
[0136] 2. Each agent directly outputs the specific recovery path and the wavelength resources used when generating the strategy, ensuring that the recovery plan has a clear and deployable optical layer configuration.
[0137] 3. The introduction of action masking and randomized recovery order strategies effectively trims the search space and avoids local optimality, allowing the generated candidate solutions to be deployed without the assistance of a complex ILP solver, improving the feasibility and practical operability of the recovery strategy.
[0138] like Figure 10As shown, LBOR consists of three main modules: a pre-training module 1010, an online inference module 1020, and an online training module 1030. These three modules work together to achieve fast optical network recovery. Each module is described below.
[0139] Pre-training module 1010: In this module, the MARL (Multi-Agent Reinforcement Learning) model is pre-trained using synthetic data (i.e., training dataset). The SDN controller 1040 first collects the optical layer topology, the IP layer topology, and the mapping between them, including the optical path and wavelength occupancy corresponding to each IP link. LBOR simulates various fiber failure scenarios to generate a training dataset, which is then used to pre-train the model, enabling it to make initial recovery decisions.
[0140] Online Inference Module 1020: LBOR uses a trained MARL model to generate an optical layer recovery plan (i.e., a path restoration plan) for recovering from a failure. It then uses a solver to perform traffic allocation. When a fiber failure occurs, the MARL model generates a recovery plan based on the current network state and transmits the updated IP link configuration (i.e., the restoration path and corresponding wavelength) to the SDN controller 1040. This result is then sent to the real-time TE (Traffic Engineering) controller 1050, which optimizes traffic allocation by solving a multi-commodity network flow optimization problem. When allocating traffic, the TE controller 1050 considers the capacity of each IP link, current traffic demand (which can be collected every five minutes by the SDN controller 1040), and a predefined set of paths. Using solvers such as Gurobi, the TE controller determines the optimal traffic split (i.e., the traffic allocation strategy).
[0141] Online Training Module 1030: LBOR regularly trains the MARL model. During actual recovery, new failure scenario data is added to the training dataset for model fine-tuning and redeployment. To avoid frequent retraining, the system regularly collects data and updates the model when the data volume reaches a certain scale, thereby improving the adaptability and generalization of the generated recovery strategy.
[0142] The aforementioned MARL model employs a dual-agent architecture based on distributed actions and a centralized value function to achieve efficient optical network restoration. These agents are the path selection agent and the wavelength allocation agent. This separation of responsibilities reduces the decision-making complexity of each agent, allowing each agent to focus on its own task, thereby improving restoration efficiency and effectiveness. Furthermore, this modular design facilitates future expansion. For example, the path selection agent can be combined with a neural network to dynamically generate optimal routing strategies. A set of candidate optical paths (i.e., the candidate optical path set) is pre-computed for each IP link, providing a clear search space for path selection.
[0143] The MARL model adopts a sequential decision-making paradigm. In LBOR, each failed IP link is restored in a random order, assigning one wavelength at a time. For each wavelength affected by a failure on an IP link, two agents collaborate to complete path selection and wavelength assignment.
[0144] Specifically, the state space of the path selection agent consists of the following three main parts.
[0145] Global wavelength occupancy information: used to reflect the usage of all wavelengths in the entire network, enabling intelligent agents to make globally aware decisions, reduce conflicts and improve efficiency.
[0146] Candidate path information: This refers to all candidate optical paths for each IP link. Each path is represented by the optical fiber sequence it passes through, providing structured path options.
[0147] Wavelength occupancy status of candidate paths: used to evaluate the current availability of the path.
[0148] The state of the wavelength assignment agent is dynamically updated according to the path selection results. It contains the wavelength occupancy status on the selected recovery path and the global wavelength occupancy information to ensure that its decision-making considers both local constraints and global resources.
[0149] For the action space of the agents, the two agents operate sequentially and make independent decisions but are associated with each other. That is, the path selection agent selects a recovery path from the set of candidate optical paths, and the wavelength allocation agent assigns an available wavelength to the recovery path.
[0150] For the agent's reward function, its main goal is to maximize the number of successfully recovered wavelengths. Whenever the agent successfully assigns an available wavelength on the path, it receives a positive reward; if no available wavelength is found, the reward is zero.
[0151] Whenever the wavelength of an IP link is successfully restored, the system will update the wavelength occupancy information of the optical fibers involved in the path, ensuring that subsequent decisions reflect changes in resource usage and avoid constraint conflicts.
[0152] Then, a fully connected neural network can be used as the policy function approximator of the model. This structure is simple but suitable for processing structured inputs. Its input is the state representation and its output is the action probabilities of the two agents.
[0153] Based on the LBOR constructed above, in order to verify its recovery quality and performance and demonstrate its advantage of high efficiency while ensuring high recovery quality, the following experiments can be used to illustrate.
[0154] 1. Experimental Setup A. Comparison with Baseline Methods LBOR is compared with the following baseline methods: Two-phase solver: This method first solves the maximum wavelength restoration problem on precomputed candidate optical paths to determine the optimal configuration at the optical layer. Based on this, it then reoptimizes the traffic distribution ratios along predefined IP layer paths to maximize overall network resilience.
[0155] ARROW: ARROW first constructs a linear programming (LP) model on candidate lightpaths to solve an initial restoration strategy. This LP solution then generates multiple candidate restoration solutions through random rounding. ARROW then uses the actual traffic matrix to calculate the optimal traffic allocation scheme under the current failure scenario to maximize restored traffic. Finally, this optimal allocation determines the traffic distribution ratio across available paths. The number of candidate solutions is set to 40.
[0156] B. Network Topology and Traffic Matrix Four network topologies (B4, IBM, Bellsouth, and Columbus) were selected from the Internet Topology Zoo as optical layer networks for evaluation. Figure 11 The IP layer networks for these topologies were derived from the wavelength distribution of each IP link in the ARROW data. Each optical layer topology corresponds to five different IP layer topologies. For each IP layer topology, 200 fault scenarios were randomly generated, totaling 1,000 datasets, 500 of which were used for training and 500 for testing. Furthermore, 30 different traffic demand matrices were generated for each topology using the YATES tool.
[0157] C. Channel Selection The k-shortest path algorithm is used to select candidate optical layer restoration paths and predefined IP layer paths. The number of candidate optical layer restoration paths is set to 3, and the number of predefined IP layer paths is set to 8.
[0158] 2. Large-scale simulation experiments The performance of LBOR on four network topologies is evaluated from two aspects: the maximum achievable throughput after recovery and the computation time.
[0159] A. Restore throughput performance Figure 12This shows the maximum traffic that can be satisfied by each dataset, given the total traffic demand. LBOR's performance is only slightly lower than Solver and ARROW, while ARROW does not demonstrate a significant advantage in improving traffic distribution. Further analysis reveals that LBOR outperforms ARROW in traffic distribution across 60 failure scenarios in the Columbus topology.
[0160] B. Computation Time Analysis Figure 13 The average solution time is shown. Directly solving integer linear programs (ILPs) is computationally expensive, and even for the smallest topology, the computation time for a single scenario can range from 0.1 seconds to 5 seconds, depending on the specific failure scenario. ARROW is computationally more burdensome because it requires multiple rounds of rounding and feasibility checks to generate 40 candidate solutions, and then solves the ILP again to obtain the final solution. This multi-step process causes ARROW to be 384 times slower than LBOR on the smallest topology and even 1000 times slower on the Columbus topology. In contrast, LBOR has significantly better scalability and maintains stable inference time under different topologies and failure scenarios. Regardless of the topology size or failure mode, LBOR's inference latency is kept within 0.82 seconds. Its main computational bottleneck lies in the LP solver in large topologies, rather than the learning model itself.
[0161] 3. Ablation Experiment To evaluate the effectiveness of the random order assignment and action masking techniques, we conducted ablation experiments to compare the performance of reinforcement learning models under different configurations. First, we constructed a baseline model without these two techniques and used its results as a normalization reference.
[0162] like Figure 14 As shown (where the horizontal axis represents the number of recovered wavelengths, that is, the number of recovered wavelengths normalized to the LBOR-naive method; the vertical axis is the cumulative distribution function (CDF), which represents the probability that the number of recovered wavelengths is less than or equal to a certain value), when only the "random assignment order" (random-only) is introduced, the number of recovered wavelengths is the same or greater in 77% of the fault scenarios; and when only the "action masking mechanism" (mask-only) is introduced, the performance is improved in 93% of the scenarios, and in some scenarios the number of recovered wavelengths is 3.38 times that of the baseline model. However, when these two technologies are used alone, there are also scenarios where the recovery rate is lower than 50% of the baseline model. In contrast, when the two technologies are used in combination, stable performance is achieved in all fault scenarios, indicating that they are complementary in improving recovery robustness.
[0163] The following describes an apparatus embodiment of the present disclosure, which can be used to perform the optical network fault repair method in the above-mentioned embodiment of the present disclosure. For details not disclosed in the apparatus embodiment of the present disclosure, please refer to the above-mentioned embodiment of the optical network fault repair method in the present disclosure.
[0164] Figure 15 FIG. 1 shows a schematic structural block diagram of a fault repair device for an optical network according to an embodiment of the present disclosure. Figure 15 As shown, the fault repair device 1500 includes a determination module 1510 , a reasoning module 1520 , an update module 1530 and a processing module 1540 .
[0165] Specifically, the determining module 1510 is configured to determine a failed IP link in the optical network in response to an optical fiber failure event in the optical network.
[0166] The reasoning module 1520 is used to take the failed IP link as the input of the multi-agent decision model, and the multi-agent decision model outputs a path recovery plan corresponding to the failed IP link. The path recovery plan includes the recovery path determined by the path selection agent in the multi-agent decision model for the failed IP link, and the wavelength allocated to the recovery path by the wavelength allocation agent in the multi-agent decision model.
[0167] The updating module 1530 is configured to update the IP link capacity according to the path recovery solution.
[0168] The processing module 1540 is used to redistribute traffic on the IP link according to the updated IP link capacity to achieve network fault repair.
[0169] In some embodiments of the present disclosure, the failed IP link is used as the input of a multi-agent decision model, and the multi-agent decision model outputs a path recovery plan corresponding to the failed IP link, including: determining the state space of a path selection agent based on the global wavelength occupancy information of the optical network, a set of candidate optical paths corresponding to the IP link, and a wavelength occupancy status corresponding to each candidate optical path; based on the state space of the path selection agent, the path selection agent selects one of the candidate optical paths corresponding to the failed IP link as a recovery path from the set of candidate optical paths corresponding to the failed IP link; and a wavelength allocation agent allocates a wavelength to the recovery path, and the state space of the wavelength allocation agent includes the wavelength occupancy status corresponding to the recovery path and the global wavelength occupancy information.
[0170] In some embodiments of the present disclosure, according to the state space of the path selection agent, the path selection agent selects one of the candidate optical paths corresponding to the failed IP link as a recovery path, including: determining a candidate optical path with an idle wavelength from the candidate optical path set corresponding to the failed IP link; according to the state space of the path selection agent, the path selection agent selects one of the candidate optical paths with an idle wavelength corresponding to the failed IP link as a recovery path corresponding to the failed IP link.
[0171] In some embodiments of the present disclosure, allocating wavelengths to the restoration path by a wavelength allocation agent includes: determining an unavailable wavelength corresponding to the restoration path based on an occupancy status of the wavelength corresponding to the restoration path; and shielding the allocation action of the unavailable wavelength when the wavelength allocation agent allocates wavelengths to the restoration path.
[0172] In some embodiments of the present disclosure, the processing module 1540 is further used to: generate a training data set based on the collected fault scenario data; train the path selection agent and wavelength allocation agent in the multi-agent decision model based on the training data set, and during the training process, give a positive reward when an available wavelength is successfully allocated on the path, and give a zero reward if no available wavelength is found.
[0173] In some embodiments of the present disclosure, when training the path selection agent and the wavelength allocation agent in the multi-agent decision model, the wavelengths of the failed IP links are restored one by one based on a random order.
[0174] In some embodiments of the present disclosure, the failed IP link is used as the input of a multi-agent decision model, and the multi-agent decision model outputs a path recovery plan corresponding to the failed IP link, including: when there are multiple failed IP links, the multi-agent decision model restores the wavelength of each failed IP link one by one based on a random order, and outputs a corresponding path recovery plan.
[0175] In some embodiments of the present disclosure, the processing module 1540 is also used to: collect new fault scenario data during the online operation phase of the multi-agent decision model; update the training data set based on the new fault scenario data; and optimize the parameters of the multi-agent decision model based on the updated training data set.
[0176] In some embodiments of the present disclosure, IP link traffic is redistributed based on the updated IP link capacity, including: establishing a multi-commodity network flow optimization problem based on known traffic demand and a set of candidate paths between IP node pairs; solving the multi-commodity network flow optimization problem based on a solver to determine a traffic allocation strategy based on the updated IP link capacity.
[0177] The present disclosure also provides an electronic device. Figure 16 A schematic diagram showing a hardware implementation using a processing system is shown.
[0178] like Figure 16 As shown, the hardware structure of electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnecting buses and bridges, depending on the specific application and overall design constraints of the hardware. Bus 1100 connects various circuits including one or more processors 1200, memory 1300, and / or hardware modules. Bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc. Bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component Architecture (EISA) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of illustration, the figure only uses a single connecting line, but this does not mean that there is only one bus or only one type of bus.
[0179] The present disclosure also provides a readable storage medium having a computer program stored therein, which is used to implement the above-mentioned method when the computer program is executed by a processor. "Readable storage medium" can be any device that can contain, store, communicate, propagate or transmit a program for use in an instruction execution system, device or equipment or in combination with these instruction execution systems, devices or equipment. More specific examples of readable storage media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.
[0180] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part using software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the process or function of the present disclosure is performed in whole or in part.
[0181] A computer program or instruction can be stored in a readable storage medium or transferred from one readable storage medium to another. For example, the computer program or instruction can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The readable storage medium can be any accessible medium or a data storage device such as a server or data center that integrates one or more accessible media. The accessible medium can be a magnetic medium such as a floppy disk, hard disk, or magnetic tape; an optical medium such as a digital video disk; or a semiconductor medium such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.
[0182] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0183] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, apparatuses, electronic devices, and computer program products according to the present disclosure. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0184] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0185] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0186] In the description of this specification, the description with reference to the terms "one embodiment / method", "some embodiments / methods", "example", "specific example", or "some examples" means that the specific features, structures, or characteristics described in conjunction with the embodiment / method or example are included in at least one embodiment / method or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / method or example. Moreover, the specific features, structures, or characteristics described may be combined in a suitable manner in any one or more embodiments / methods or examples. In addition, those skilled in the art may combine and combine different embodiments / methods or examples described in this specification and the features of different embodiments / methods or examples, unless they are contradictory.
[0187] Those skilled in the art will appreciate that the above embodiments are merely intended to clearly illustrate the present disclosure and are not intended to limit the scope of the present disclosure. Other changes or modifications may be made based on the above disclosure, and such changes or modifications are still within the scope of the present disclosure.
Claims
1. A fault repair method for an optical network, characterized in that: include: In response to a fiber failure event in an optical network, determining a failed IP link in the optical network; Taking the failed IP link as an input of a multi-agent decision model, the multi-agent decision model outputs a path recovery plan corresponding to the failed IP link, the path recovery plan including a recovery path determined for the failed IP link by a path selection agent in the multi-agent decision model and a wavelength assigned to the recovery path by a wavelength assignment agent in the multi-agent decision model; updating the IP link capacity according to the path restoration plan; as well as Based on the updated IP link capacity, traffic on the IP link is redistributed to repair network faults.
2. The method according to claim 1, wherein The failed IP link is used as an input of a multi-agent decision model, and the multi-agent decision model outputs a path restoration solution corresponding to the failed IP link, including: Determining a state space of a path selection agent based on global wavelength occupancy information of the optical network, a set of candidate optical paths corresponding to the IP link, and a wavelength occupancy state corresponding to each candidate optical path; According to the state space of the path selection agent, the path selection agent selects one of the candidate optical paths corresponding to the failed IP link as a recovery path; A wavelength allocation agent allocates a wavelength to the restoration path, and a state space of the wavelength allocation agent includes a wavelength occupancy state corresponding to the restoration path and the global wavelength occupancy information.
3. The method according to claim 2, wherein The path selection agent selects, according to the state space of the path selection agent, one of the candidate optical paths corresponding to the failed IP link as a recovery path, including: Determining a candidate optical path having an idle wavelength from a set of candidate optical paths corresponding to the failed IP link; According to the state space of the path selection agent, the path selection agent selects one of the candidate optical paths with idle wavelengths corresponding to the failed IP link as the restoration path corresponding to the failed IP link.
4. The method according to claim 2, wherein Allocating a wavelength to the restoration path by a wavelength allocation agent includes: determining an unavailable wavelength corresponding to the restoration path according to an occupation status of the wavelength corresponding to the restoration path; When the wavelength assignment agent assigns wavelengths to the restoration path, the assignment action of unavailable wavelengths is shielded.
5. The method according to claim 1, wherein The method further comprises: Generate a training dataset based on the collected fault scenario data; The path selection agent and the wavelength assignment agent in the multi-agent decision model are trained based on the training data set. During the training process, a positive reward is given when an available wavelength is successfully assigned on the path, and a reward is zero if no available wavelength is found.
6. The method according to claim 5, wherein When training the path selection agent and the wavelength allocation agent in the multi-agent decision model, the wavelengths of the failed IP links are restored one by one based on a random order.
7. The method according to claim 6, wherein The failed IP link is used as an input of a multi-agent decision model, and the multi-agent decision model outputs a path restoration solution corresponding to the failed IP link, including: In the case that there are multiple failed IP links, the multi-agent decision model restores the wavelength of each failed IP link one by one based on a random order and outputs a corresponding path restoration solution.
8. The method according to claim 1, wherein Redistribute IP link traffic based on the updated IP link capacity, including: Based on the known traffic demand and the set of candidate paths between IP node pairs, a multi-commodity network flow optimization problem is established; The multi-commodity network flow optimization problem is solved based on the solver to determine a flow allocation strategy based on the updated IP link capacity.
9. An electronic device, characterized in that: include: a memory storing execution instructions; as well as A processor, wherein the processor executes the execution instruction stored in the memory, so that the processor executes the fault repair method for an optical network according to any one of claims 1 to 8.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the fault repair method for an optical network according to any one of claims 1 to 8 is implemented.