Network fiber breakage resistance protection planning method based on offline learning and improved MCTS

By dynamically searching for the optimal backup path of the optical network based on offline learning and an improved MCTS multi-mode multi-objective optimization algorithm, the problems of poor timeliness and low service recovery rate in optical network restoration technology are solved, achieving fast and effective service recovery.

CN120602313APending Publication Date: 2025-09-05WUHAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510711142.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing optical network restoration technologies have poor timeliness and low service recovery rates, making it difficult to quickly and effectively restore service continuity when a fiber optic link fails.

Method used

A multi-mode multi-objective optimization algorithm based on offline learning and improved Monte Carlo Tree Search (MCTS) is adopted. The network anti-fiber break protection planning model is learned through the preset multi-mode multi-objective optimization algorithm, the optimal backup path is dynamically searched, and the improved MCTS algorithm is used for rapid recovery when a link failure occurs.

Benefits of technology

It improves the recovery efficiency and service recovery rate of the optical network when the optical fiber link fails, reduces data loss and interruption time, and improves the performance and efficiency of network anti-fiber break protection planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602313A_ABST
    Figure CN120602313A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent optimization, in particular to a network anti-fiber-breakage protection planning method based on offline learning and improved MCTS, and the method comprises the steps: solving a preset multi-target optimization model of network anti-fiber-breakage protection planning through employing a preset multi-mode multi-target optimization algorithm, obtaining at least one optimal standby path corresponding to each service; learning a topological structure and link weight information of each optimal standby path; and when a link fault exists, dynamically searching a target recovery path of the current damaged service according to the topological structure and link weight information of at least one optimal standby path corresponding to the current damaged service in all services based on a preset improved Monte Carlo tree search algorithm, and outputting the target recovery path. Therefore, the problems of poor timeliness and low service recovery rate in the existing optical network recovery technology are solved, and the performance and efficiency of network fiber breakage resistance protection planning are effectively improved by comprehensively optimizing the shortest recovery time and the maximum service recovery rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent optimization technology, and in particular to a network anti-fiber break protection planning method based on offline learning and improved MCTS (Monte Carlo Tree Search). Background Art

[0002] Network fiber-break protection planning is about finding a recovery path for damaged services as quickly as possible when a link failure occurs in the network, essentially maintaining network survivability. With the rapid development of internet services, elastic optical networks carry massive amounts of data. However, fiber link failures frequently threaten service continuity, making network fiber-break protection planning crucial for maintaining network survivability.

[0003] Elastic optical networks (ELNs) effectively overcome the shortcomings of traditional wavelength division multiplexing (WDM) networks, which often rely on coarse-grained bandwidth allocation and fixed modulation formats, and are widely considered to be a promising intelligent optical network. During real-time communication in ELNs, traffic volume fluctuates over time. With the rapid growth of ultra-high-definition video, mobile applications, big data, and cloud services, internet applications are experiencing tremendous growth, and networks are becoming increasingly large and complex, increasing the probability of network failures. Furthermore, because optical fibers carry vast amounts of traffic, failures in fiber links or nodes within the network can lead to massive data loss and service interruptions, inflicting immeasurable losses on network users and operators. The large number of nodes and links in ELNs, coupled with unique constraints (spectral continuity, spectral consistency, and spectral non-overlap), make this problem an NP-hard problem. Over the past decade, numerous researchers have conducted in-depth research on this problem, proposing a variety of precise and heuristic approaches to solve it.

[0004] While precise methods can theoretically find optimal solutions to problems, they are time-consuming and subject to many limitations in practical applications. Heuristic methods can obtain feasible solutions to problems within a limited time, but their effectiveness deteriorates as the task becomes larger and the constraints become more stringent.

[0005] Currently, there are two main solutions to the problem of fiber-break protection planning in elastic optical networks: optical network protection technology and optical network recovery technology. Optical network protection technology involves finding backup paths for services in advance. Its advantage is that when a failure occurs, the working path can be directly switched to the backup path, which is faster and almost instantaneous. However, resources for the backup path must be reserved, increasing network construction costs. Optical network recovery technology involves rerouting damaged services in the elastic optical network when a failure occurs. Its advantage is that it can maximize the utilization of network resources. However, the recovery performance when a failure occurs is strongly related to the number of services, and the current rerouting algorithm has poor performance, which urgently needs to be addressed. Summary of the Invention

[0006] This application provides a network anti-fiber break protection planning method based on offline learning and improved MCTS to solve the problems of poor timeliness and low service recovery rate in existing optical network restoration technologies. By comprehensively optimizing the shortest restoration time and maximum service recovery rate, the performance and efficiency of network anti-fiber break protection planning are effectively improved.

[0007] The first embodiment of the present application provides a network anti-fiber break protection planning method based on offline learning and improved MCTS, including the following steps:

[0008] Using a preset multi-mode multi-objective optimization algorithm to solve the preset multi-objective optimization model of the network anti-fiber break protection plan, at least one optimal backup path corresponding to each service is obtained;

[0009] Learn the topology and link weight information of each optimal backup path;

[0010] When there is a link failure, based on the preset improved Monte Carlo tree search algorithm, according to the topological structure and link weight information of at least one optimal backup path corresponding to the currently damaged business among all businesses, the target recovery path of the currently damaged business is dynamically searched and the target recovery path is output.

[0011] According to one embodiment of the present application, the preset multi-mode multi-objective optimization algorithm includes:

[0012] The population is divided into multiple sub-populations based on the preset affinity propagation clustering algorithm, where each sub-population evolves independently;

[0013] At least one non-dominated solution is screened out from all archives, and based on a preset Pareto solution set learning strategy, the at least one non-dominated solution is classified to obtain at least one potential search area, and Gaussian fitting is performed on the non-dominated solutions in each potential search area to obtain a Gaussian fitting result;

[0014] A target search area is determined based on the Gaussian fitting result, and location information of a Pareto optimal solution is learned according to the target search area to obtain a learning result, and a new solution is generated according to the learning result.

[0015] According to one embodiment of the present application, after generating a new solution according to the learning result, the method further includes:

[0016] The ratio of the local search population to the global search population is dynamically adjusted according to the evolutionary stage of the preset multi-mode multi-objective optimization algorithm and the state of the population.

[0017] According to one embodiment of the present application, the preset improved Monte Carlo tree search algorithm includes:

[0018] In the node selection phase, the weights of the edges and / or nodes in the preset path are adjusted based on a preset topological structure weight adjustment mechanism;

[0019] In the simulation phase, the weight value corresponding to each action of the current node is calculated, the target simulation action is determined according to the weight value corresponding to each action of the current node, and the target simulation action is simulated;

[0020] When the link failure occurs, at least one damaged service and local topology information corresponding to each damaged service are determined, and a target recovery path corresponding to each damaged service is searched based on the local topology information corresponding to each damaged service and at least one optimal backup path.

[0021] According to an embodiment of the present application, the objective function of the preset multi-objective optimization model of the network anti-fiber break protection plan includes at least one of a service recovery time minimization function and a service recovery rate maximization function.

[0022] According to the network anti-fiber break protection planning method based on offline learning and improved MCTS in the embodiment of the present application, a preset multi-mode multi-objective optimization algorithm is used to solve the preset multi-objective optimization model of the network anti-fiber break protection planning, obtain at least one optimal backup path corresponding to each service, and learn the topology structure and link weight information of each optimal backup path; when there is a link failure, based on the preset improved Monte Carlo tree search algorithm, according to the topology structure and link weight information of the optimal backup path corresponding to the currently damaged service among all services, the target recovery path of the currently damaged service is dynamically searched, and the target recovery path is output. In this way, the problems of poor timeliness and low service recovery rate in the existing optical network recovery technology are solved, and the performance and efficiency of the network anti-fiber break protection planning are effectively improved by comprehensively optimizing the shortest recovery time and the maximum service recovery rate.

[0023] A second embodiment of the present application provides a network anti-fiber break protection planning device based on offline learning and improved MCTS, including:

[0024] A solution module is used to solve a preset multi-objective optimization model of the network anti-fiber break protection plan using a preset multi-mode multi-objective optimization algorithm to obtain at least one optimal backup path corresponding to each service;

[0025] A learning module is used to learn the topology structure and link weight information of each optimal backup path;

[0026] The recovery module is used to dynamically search for a target recovery path for a currently damaged service based on a preset improved Monte Carlo tree search algorithm when a link failure occurs, according to the topological structure and link weight information of at least one optimal backup path corresponding to the currently damaged service among all services, and output the target recovery path.

[0027] According to one embodiment of the present application, the solution module is used to:

[0028] The population is divided into multiple sub-populations based on the preset affinity propagation clustering algorithm, where each sub-population evolves independently;

[0029] At least one non-dominated solution is screened out from all archives, and based on a preset Pareto solution set learning strategy, the at least one non-dominated solution is classified to obtain at least one potential search area, and Gaussian fitting is performed on the non-dominated solutions in each potential search area to obtain a Gaussian fitting result;

[0030] A target search area is determined based on the Gaussian fitting result, and location information of a Pareto optimal solution is learned according to the target search area to obtain a learning result, and a new solution is generated according to the learning result.

[0031] According to one embodiment of the present application, after generating a new solution based on the learning result, the solution module is further configured to:

[0032] The ratio of the local search population to the global search population is dynamically adjusted according to the evolutionary stage of the preset multi-mode multi-objective optimization algorithm and the state of the population.

[0033] According to one embodiment of the present application, the recovery module is configured to:

[0034] In the node selection phase, the weights of the edges and / or nodes in the preset path are adjusted based on a preset topological structure weight adjustment mechanism;

[0035] In the simulation phase, the weight value corresponding to each action of the current node is calculated, the target simulation action is determined according to the weight value corresponding to each action of the current node, and the target simulation action is simulated;

[0036] When the link failure occurs, at least one damaged service and local topology information corresponding to each damaged service are determined, and a target recovery path corresponding to each damaged service is searched based on the local topology information corresponding to each damaged service and at least one optimal backup path.

[0037] According to an embodiment of the present application, the objective function of the preset multi-objective optimization model of the network anti-fiber break protection plan includes at least one of a service recovery time minimization function and a service recovery rate maximization function.

[0038] According to the network anti-fiber break protection planning device based on offline learning and improved MCTS according to the embodiment of the present application, a preset multi-mode multi-objective optimization algorithm is used to solve the preset multi-objective optimization model of the network anti-fiber break protection planning, obtain at least one optimal backup path corresponding to each service, and learn the topology structure and link weight information of each optimal backup path; when there is a link failure, based on the preset improved Monte Carlo tree search algorithm, according to the topology structure and link weight information of the optimal backup path corresponding to the currently damaged service among all services, the target recovery path of the currently damaged service is dynamically searched, and the target recovery path is output. In this way, the problems of poor timeliness and low service recovery rate in the existing optical network recovery technology are solved, and the performance and efficiency of the network anti-fiber break protection planning are effectively improved by comprehensively optimizing the shortest recovery time and the maximum service recovery rate.

[0039] An embodiment of the third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and runnable on the processor, wherein the processor executes the program to implement the network anti-fiber break protection planning method based on offline learning and improved MCTS as described in the above embodiment.

[0040] The fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the network anti-fiber break protection planning method based on offline learning and improved MCTS as described in the above embodiment.

[0041] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0043] Figure 1 A flowchart of a network anti-fiber break protection planning method based on offline learning and improved MCTS according to an embodiment of the present application;

[0044] Figure 2 A flowchart of a multi-mode multi-objective algorithm based on a niche strategy and a Pareto solution set learning strategy according to one embodiment of the present application;

[0045] Figure 3 A flowchart of solving a network anti-fiber break protection planning problem according to one embodiment of the present application;

[0046] Figure 4 Schematic diagram of a network anti-fiber break protection planning device based on offline learning and improved MCTS according to an embodiment of the present application;

[0047] Figure 5 Schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0048] The following describes in detail embodiments of the present application. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0049] Those skilled in the art will understand that in classic network fiber-break protection planning models, the shortest recovery time is generally the optimization objective. However, in practical applications, multiple optimization objectives often exist, and simply considering the shortest recovery time alone cannot meet actual application needs. In different contexts, some scholars have proposed multiple optimization objectives, such as service recovery rate, resource allocation optimization, and minimum path cost. However, most research focuses on a single objective, and research results that combine and comprehensively consider multiple objectives are rare.

[0050] This invention utilizes an intelligent optimization algorithm, also known as a metaheuristic algorithm, which maintains excellent problem-solving capabilities and convergence efficiency even when faced with larger tasks and more stringent constraints. In particular, when solving NP-hard problems, this intelligent optimization algorithm can effectively address the "combinatorial explosion" phenomenon and obtain the optimal solution. This invention constructs a network fiber-break protection planning model with the shortest recovery time and maximum service recovery rate as optimization objectives.

[0051] Furthermore, the present invention considers multiple objectives simultaneously for optimization when solving the network anti-fiber break protection planning problem, and defines the network anti-fiber break protection planning problem as a multi-objective problem. For one service, there are multiple optimal paths in the network, which is consistent with the characteristics of the multi-mode multi-objective problem (one solution in the objective space corresponds to multiple solutions in the decision space). Therefore, the network anti-fiber break protection planning problem is a standard multi-mode multi-objective optimization problem. The present invention considers the most common single-fiber link failure scenario in elastic optical networks, and uses optical network recovery technology based on multi-mode multi-objective optimization algorithms to solve the network anti-fiber break protection planning model. In view of the shortcomings of existing optical network recovery technologies such as poor timeliness and low service recovery rate, a multi-mode multi-objective algorithm based on offline learning and improved Monte Carlo tree search is designed to solve the network anti-fiber break protection planning problem in elastic optical networks.

[0052] The following describes a network anti-fiber break protection planning method based on offline learning and improved MCTS according to an embodiment of the present application with reference to the accompanying drawings.

[0053] Specifically, Figure 1 A flowchart of a network anti-fiber break protection planning method based on offline learning and improved MCTS is provided in an embodiment of the present application.

[0054] like Figure 1 As shown, the network anti-fiber break protection planning method based on offline learning and improved MCTS includes the following steps:

[0055] In step S101, a preset multi-mode multi-objective optimization algorithm is used to solve a preset multi-objective optimization model of the network anti-fiber break protection plan to obtain at least one optimal backup path corresponding to each service.

[0056] In some embodiments, the objective function of the preset multi-objective optimization model of the network anti-fiber break protection planning includes at least one of a function of minimizing the service recovery time and a function of maximizing the service recovery rate.

[0057] Specifically, the network anti-fiber break protection planning algorithm model based on offline learning and improved Monte Carlo tree search in the embodiment of the present application is constructed as follows:

[0058] In network fiber-break protection planning models, the most common optimization objective is minimizing service recovery time. This is crucial in real-world scenarios: shorter service recovery time means faster communication restoration, improving efficiency. It also effectively reduces data loss during outages. When a network outage occurs, the goal is to find a recovery path for the damaged service in the shortest possible time, while satisfying model constraints. This minimizes service recovery time. The model formula for service recovery time is as follows:

[0059] Minf N ;

[0060] f i =start time -end time ;

[0061] f i <d j ;

[0062] Among them, f i is the recovery time of the i-th service; i = 1, 2, ..., N, where N is the total number of damaged services; start time The time when the i-th business starts to disconnect, end time is the time for the i-th service to be restored; d j is the duration of the jth outage; at any moment, the project recovery time cannot be greater than the duration of the outage d j Otherwise, it is considered that the business cannot be restored before the circuit breaker is restored.

[0063] The service recovery rate refers to the goal of being able to restore services before the circuit breaker is restored when a circuit breaker occurs in the elastic optical network. Failure to restore services will result in a large amount of resource loss and communication interruption.

[0064] In this embodiment of the application, the service recovery rate is used as an indicator of robustness, which is expressed as the ratio of recoverable services to the total damaged services in the project. The model formula based on the service recovery rate is as follows:

[0065] RR=∑ k∈S (service recover / service blocking );

[0066] Among them, service recover is the number of services that can be restored, service blocking The total number of businesses affected.

[0067] Therefore, the objective function of the network anti-fiber break protection planning problem model is:

[0068] Minf N ;

[0069] MaxRR;

[0070] Among them, Minf N is the function that minimizes the service recovery time, and MaxRR is the function that maximizes the service recovery rate.

[0071] Furthermore, in some embodiments, the preset multi-mode multi-objective optimization algorithm includes: dividing the population into multiple sub-populations based on a preset affinity propagation clustering algorithm, wherein each sub-population evolves independently; screening out at least one non-dominated solution from all archives, classifying the at least one non-dominated solution based on a preset Pareto solution set learning strategy to obtain at least one potential search area, and performing Gaussian fitting on the non-dominated solutions in each potential search area to obtain a Gaussian fitting result; determining the target search area based on the Gaussian fitting result, and learning the location information of the Pareto optimal solution according to the target search area to obtain a learning result, and generating a new solution based on the learning result.

[0072] Specifically, in order to solve the network anti-fiber break protection planning problem model, the embodiment of the present application proposes a multi-mode multi-objective optimization algorithm based on the niche strategy and the Pareto solution set learning strategy to find as many optimal paths as possible for each service.

[0073] Particle swarm optimization is a multi-objective evolutionary algorithm based on non-dominated sorting. It has been widely studied and applied in solving multi-objective problems. Its basic idea is to first randomly generate a parent population P with a population size of N. t , by learning the global optimum and individual historical optimum, the offspring population Q is generated t The two populations are merged and then non-dominated sorting is performed while calculating the crowding degree. According to the non-dominated relationship and crowding degree between individuals, suitable individuals are selected to form a new parent population P. t+1 Then update the global optimum and individual historical optimum, and repeat the cycle until the end condition is met. The basic process of the preset multi-mode multi-objective optimization algorithm proposed in this application is as follows Figure 2 As shown, in the present invention, the particle swarm algorithm is improved as follows:

[0074] In order to improve the diversity of multi-modal and multi-objective optimization algorithms, traditional algorithms usually adopt a small habitat strategy to improve the diversity of the population. The small habitat strategy divides the population into multiple regions, and each particle evolves only within the assigned region. This can reduce the speed of information transmission and avoid rapid convergence to the local optimum. At the same time, it can fully explore all areas of the decision space and improve the diversity of the generated solutions.

[0075] Traditional niche methods usually require preset parameters (such as the radius or number of niches), which makes the algorithm sensitive to parameters and difficult to adapt to the characteristics of different problems. In order to solve this problem, this application adopts affinity propagation clustering (APClustering) as an adaptive niche generation method. AP clustering is an adaptive clustering method that does not require preset parameters. It automatically determines the cluster center and the number of clusters through the message passing process. Its core idea is to determine the cluster center by calculating the similarity matrix between samples and iteratively updating Responsibility and Availability.

[0076] In each generation, AP clustering is used to divide the population into multiple subpopulations (niches). Each subpopulation evolves independently, ensuring that particles conduct local exploration within their respective niches, thereby enhancing the diversity of solutions. The distribution of niches can be dynamically adjusted based on the current state of the population to adapt to the search needs at different stages.

[0077] The adaptive niche strategy is the first stage of the algorithm. Its goal is to enhance solution diversity and identify all potential Pareto-optimal solution (PS) regions, providing learning samples for the Pareto solution set learning strategy. To identify all potential Pareto-optimal solution (PS) regions when solving MMOPs (Multi-Modal Multi-Objective Optimization Problems), it is crucial to have a population with strong search diversity. As mentioned above, affinity propagation clustering is a parameter-free clustering method. Through AP clustering, multiple subpopulations are adaptively created in each generation, effectively maintaining population diversity during the search process. Affinity propagation clustering determines clustering results based on a message passing process without specifying cluster centers or the number of clusters. Therefore, it avoids sensitivity to other parameters. At the same time, under APC clustering, individuals with similar characteristics are clustered in the same niche and are truly spatially adjacent.

[0078] Furthermore, when dealing with multimodal and multi-objective problems, few algorithms consider leveraging the positional information of existing Pareto solution sets to guide the generation of new solutions. Learning this positional information can improve algorithm performance. A key issue is how to learn the positional information of these PSs and increase the likelihood that other particles will explore the global optimal solution. This paper proposes a novel Pareto Optimal Set (PS) learning strategy that learns the positional information of the PSs during their evolution and, based on the learning results, enhances the convergence of new solutions.

[0079] The algorithm focuses on diverse searches in the early stages, thus employing an adaptive niche search strategy. The resulting Pareto solution sets provide learning samples for the later Pareto solution set learning strategy. To improve the convergence efficiency of the algorithm in the later stages of evolution, this application proposes a search strategy based on Pareto Solution Set Learning (PSL). This strategy analyzes the distribution of non-dominated solutions in the decision space, identifies potential search areas, and guides particles to conduct local searches.

[0080] First, non-dominated solutions are selected from all archives, and then these non-dominated solutions are classified and divided into multiple regions, each representing a potential search area, in order to improve the learning effect. Then, Gaussian fitting is performed on the non-dominated solutions in each region. Gaussian process is a kind of random process in probability theory and mathematical statistics. According to the mean μ(x) and covariance K(x j ,y j ) A Gaussian process can be defined, and the mean function and kernel function of the Gaussian process can be modified by observing samples. The expressions are as follows:

[0081]

[0082] in, is the probability density function of the Gaussian distribution, is a random vector, pi is the circumference of the circle, n is the dimension of the random vector, K is the covariance, is the mean of the random vector.

[0083] Using the Pareto solution set as a sample, Gaussian fitting is performed on each region to generate a probability distribution model. The Gaussian fitting results are intended to identify potential high-quality search areas, which is equivalent to performing segmented fitting on the Pareto solution set. The purpose of segmentation is to improve the performance of each segment, thereby improving the fitting accuracy of each potential region and providing better guidance for the subsequent particle learning process. For the learning samples, a special crowding degree calculation is first performed, and new solutions are generated near the particles with the highest crowding degree, that is, near the particles with the best diversity.

[0084] Furthermore, in some embodiments, after generating a new solution based on the learning results, it also includes: dynamically adjusting the ratio of the local search population and the global search population based on the evolutionary stage and the state of the population using a preset multi-mode multi-objective optimization algorithm.

[0085] Specifically, dual-population coevolution divides the population into two co-evolving components: a local search population and a global search population. The local search population uses Pareto solution set learning to perform local searches, focusing on developing potential search areas and improving the algorithm's convergence efficiency. The global search population performs a global search to ensure population diversity and prevent the algorithm from falling into local optima. Simultaneously, the number of particles in both populations is dynamically adjusted, and the ratio of the local search population to the global search population is adjusted based on the algorithm's evolutionary stage and population status.

[0086] In the early stage, the global search population dominates, enhancing diversity; in the later stage, the local search population dominates, improving convergence.

[0087] In step S102 , the topology structure and link weight information of each optimal backup path are learned.

[0088] In step S103, when there is a link failure, based on the preset improved Monte Carlo tree search algorithm, according to the topological structure and link weight information of at least one optimal backup path corresponding to the currently damaged business among all businesses, the target recovery path of the currently damaged business is dynamically searched and the target recovery path is output.

[0089] Furthermore, in some embodiments, the preset improved Monte Carlo tree search algorithm includes: in the node selection stage, adjusting the weights of the edges and / or the weights of the nodes in the preset path based on the preset topological structure weight adjustment mechanism; in the simulation stage, calculating the weight value corresponding to each action of the current node, determining the target simulation action based on the weight value corresponding to each action of the current node, and simulating the target simulation action; when there is a link failure, determining at least one damaged service and the local topology information corresponding to each damaged service, and searching for the target recovery path corresponding to each damaged service based on the local topology information corresponding to each damaged service and at least one optimal backup path.

[0090] Specifically, the embodiment of the present application learns the local topology and link weight information of the optimal paths obtained for each service, simulates link disconnection, and guides the Monte Carlo tree search for the recovery path of the damaged service. The core idea of ​​this strategy is to optimize the selection and simulation process of the Monte Carlo tree search through the local topology information obtained through preset path learning, thereby improving the search efficiency and solution quality, and mainly includes the following three improved parts.

[0091] The first part is weight adjustment based on pre-set path learning: In traditional Monte Carlo tree search, selection strategies (such as UCT (Upper Confidence Bound for Trees)) typically rely solely on node statistics (such as visit counts and reward values). To leverage the results of pre-set path learning, the improved strategy in this application introduces local topological weights during node selection. This gives higher weights to edges or nodes that frequently appear in the pre-set path, making them more likely to be selected during the search process.

[0092] The second part is the simulation process guided by local topology. During the simulation phase, traditional Monte Carlo tree search relies entirely on random sampling, which can lead to inefficient search. The improved strategy of this application introduces local topology guidance, prioritizing the simulation of edges or nodes that frequently appear in the preset path: first, for each possible action of the current node, its local topological weight is calculated, and then the actions are sorted according to the weight value, with actions with higher weights being prioritized for simulation. In this way, the simulation process can converge to a high-quality solution more quickly.

[0093] The third improvement addresses the potential for local changes in the network topology when a link failure occurs. To accommodate this dynamic nature, the proposed strategy introduces a mechanism for dynamically adjusting the search scope. When a link failure is detected, the algorithm first identifies the affected services and then locates the local topology information corresponding to the affected services. The algorithm then focuses its search resources, prioritizing repairs on the affected paths. For unaffected areas, the algorithm maintains its original search strategy, avoiding unnecessary computational overhead.

[0094] Therefore, the present invention aims at the multi-mode and multi-objective problem characteristics of the network anti-fiber break protection planning problem, and proposes a corresponding multi-mode and multi-objective optimization algorithm to solve the problem. Through the niche strategy and Pareto solution set learning strategy, as many optimal paths as possible are found for each service, and then the topological structure of these optimal paths and the weight information of the links are learned to guide the Monte Carlo tree search process, thereby improving the timeliness of the algorithm and improving the service recovery rate.

[0095] The following combination Figure 3 The network anti-fiber break protection planning method based on offline learning and improved MCTS proposed in this application is described in detail. Figure 3 As shown, the network anti-fiber break protection planning method based on offline learning and improved MCTS includes the following steps:

[0096] First, a simulated network model is constructed and a multi-mode, multi-objective algorithm is applied for offline learning to optimize the network's recovery strategy in the event of a fiber break. During the learning process, the algorithm iterates until the preset number of fiber breaks is reached. The algorithm then calculates the overall fiber break recovery rate to assess the network's recovery efficiency in the event of multiple fiber breaks. If the number of fiber breaks has not been reached, the simulation continues.

[0097] Secondly, after simulating a fiber break, we learn offline backup paths. These paths are pre-calculated before the fiber break occurs and are used to quickly restore network connectivity. We then use Monte Carlo search to perform network recovery.

[0098] Finally, the fiber break recovery time is calculated. This is an important indicator for measuring the performance of the algorithm and reflects the time required from the occurrence of fiber break to the complete recovery of the network.

[0099] Therefore, the multi-mode, multi-objective algorithm based on offline learning and improved Monte Carlo tree search proposed in the embodiments of the present application includes a multi-mode, multi-objective optimization algorithm based on a niche strategy and a Pareto solution set learning strategy, which is used to solve the optimal path for as many services as possible. A multi-mode, multi-objective optimization algorithm based on offline learning and improved Monte Carlo tree search is proposed. By learning the topology of the optimal path and link weight information, it guides the Monte Carlo tree search and uses optical network restoration technology to solve the network fiber break protection planning problem. To overcome the shortcomings of current optical network restoration technology in algorithm timeliness and service recovery rate, the network fiber break protection planning problem model is further enriched, and a network fiber break protection planning algorithm based on offline learning and improved Monte Carlo tree search is provided. During offline learning, a multi-mode, multi-objective algorithm based on PSO is used to solve as many learning samples as possible, improving the timeliness of real-time rerouting. While ensuring the accuracy of the algorithm solution, the present invention can significantly reduce the algorithm timeliness and improve the recovery rate of damaged services, improve efficiency, and reduce outage delay, thereby reducing the losses suffered by elastic optical networks during network outages.

[0100] In order to facilitate those skilled in the art to more clearly understand the network anti-fiber break protection planning method proposed in this application, it is described in detail below in conjunction with a specific embodiment. This embodiment adopts an optical network recovery technology based on a multi-mode multi-objective optimization algorithm. This embodiment is based on offline learning and an improved Monte Carlo tree search algorithm, provides a priority coding mechanism, and improves the particle swarm algorithm.

[0101] In this example, the experiment used a representative data set of actual services and optical network topology provided by a telecommunications company. This data set contained 1,984 network nodes and 8,732 bidirectional fiber links, providing a realistic scenario for testing the algorithm's fiber break resistance.

[0102] The test set primarily consists of three tables: node (network topology), oms (optical network path information), and service (service). The following describes the contents of each of these tables in detail.

[0103] The node table contains only two columns, node and nodeID, which represent the node and its ID.

[0104] OMS is a CSV file containing Optical Multiplex Section (OMS) information. It describes the logical paths and their attributes (cost, distance, spectrum, etc.) in the optical network, supporting path calculation, resource allocation, and network optimization (such as selecting low-cost or spectrum-specific paths). The following is an introduction to the key fields and contents of this table:

[0105] Table field description: OMS: Identifies the record type as an optical network path (fixed value). omsId: Unique identifier of the path. remoteOmsId: Identifier of the remote path (a symmetric record for a bidirectional path). src: The starting node ID of the path. snk: The ending node ID of the path. cost: The cost of the path. distance: The distance of the path. ots: The optical transmission segment identifier (fixed value 1). osnr: The optical signal-to-noise ratio (fixed value 0.000776). slice: The slice capacity (fixed value 6250). colors: Wavelength or spectrum allocation information (format: 0-960, some parts are empty).

[0106] Data characteristics: Each path has forward and reverse records (for example, omsId=4526 and remoteOmsId=4537 correspond to another omsId=4537 and remoteOmsId=4526).

[0107] Spectrum allocation: The colors field is mostly 0-960, indicating standard spectrum allocation. A few records contain other values ​​(such as 96-864, 0-964), indicating that special spectrum ranges or resources are occupied.

[0108] Association with other tables: node table: the src and snk fields are associated with the nodeId of the node table.

[0109] The service table primarily contains service information and includes nine columns: Index (service ID); src (source node ID); snk (destination node ID); sourceOtu (source OTU (Optical Transport Unit) ID); targetOtu (destination OTU ID); m_width (bandwidth, all values ​​are 24); bandType (band type); sourceDimColors (source node spectrum range); and targetDimColors (destination node spectrum range). This table describes over 1,800 services, their corresponding source and destination nodes, and their corresponding spectrum ranges.

[0110] Based on the above data, the specific implementation steps of the present invention are as follows:

[0111] The business recovery time and business recovery rate of the model are calculated according to the formulas described above. They will be used as the objective functions of the model in the following solution steps.

[0112] A multi-modal multi-objective optimization algorithm based on the niche strategy and Pareto solution set learning strategy is used to solve offline learning samples. The specific implementation steps are as follows:

[0113] (1) Adaptive microhabitat strategy

[0114] In each generation, the population in each niche evolves independently to avoid loss of population diversity. First, initialize a new population P new Used to store offspring. During the evolution process, the speed and position of the particles are updated through the speed and position update formula in PSO. The global optimal here is the global optimal subGbest of the subpopulation. k The individual optimality is the historical optimality in the individual archive, and by learning the global optimal subGbest k and individual optimality Generate new particles Will Add a new population P new and the archive of the i-th particle of the k-th subpopulation, and then update the individual optimal of the particle.

[0115] In the individual's historical archive In the process, the archived particles are sorted globally, and the individuals that are not global optimal or local optimal are removed. Then the special crowding degree is calculated and sorted in descending order according to the size of the special crowding degree. The first particle has the largest crowding degree and the best diversity, so it is marked as the updated individual optimal. After this, subGbest k and NDSet k Will be updated. Only when subGbest k The optimal individual updated by some When it is dominated, it will be Otherwise, it remains unchanged. This operation helps prevent frequent changes in direction during evolution and enhances convergence. Add to NDSet k In the process, the global PS or local PS is retained and the dominated individuals are discarded. The cycle will continue until all niches have completed evolution.

[0116] The historical positions collected during the iteration process are of great value to the subsequent evolution process. Therefore, two internal archives are created in the dynamic niche PSO, namely, the personal archive and Create a non-dominated solution of NDSet is used to store non-dominated individual positions. k All personal profiles of particles in the k-th niche are collected, and the non-dominated individuals in the k-th niche are saved, which are prepared and provided for the subsequent evolutionary process.

[0117] (2) Pareto solution learning strategy

[0118] When the evolutionary generation reaches the preset value T, the embodiment of the present application uses a dynamic dual population method to divide the population dynamics into two parts. The P1 population still uses the adaptive niche particle swarm optimization algorithm. The P2 population uses the Pareto set learning method. A is the archive of the global non-dominated solution. First, the archive is sorted based on the distance A The particles in the cluster are classified, with the minimum distance set to rmin, which is related to the size of the decision space. After classification, Gaussian fitting is applied to particles in each cluster to learn their position information. Gaussian fitting is chosen for its simplicity, speed, and ease of implementation. The particle with the highest crowding degree and its nearest neighbor are then selected. The positions between these two particles are explored, and a new solution is found using the learning results of the Pareto set learning method.

[0119] Furthermore, the optimal path topology and link weight information obtained by learning guide the Monte Carlo tree search process, which includes the following steps:

[0120] (1) Offline learning

[0121] During offline learning, the embodiment of the present application uses a particle swarm algorithm based on adaptive niche and Pareto solution set learning to search for backup paths for services. After finding as many backup paths as possible for each service, the local topology of these paths and the weight information of the links are learned.

[0122] The embodiment of the present application adopts a priority coding method to map the location information of the particle swarm with the path.

[0123] The following example illustrates how a chromosome based on priority encoding represents a path in a directed graph. Consider a chromosome [7, 3, 4, 6, 2, 5, 8, 10, 1, 9]. This chromosome is randomly generated using a genetic algorithm and is not an optimal solution. It's important to note that each value in the chromosome represents the priority of the corresponding node in the graph, not the order in which the nodes are visited. The length of the chromosome typically matches the number of nodes in the graph. For example, in this example, the chromosome length is 10, corresponding to the 10 nodes in the graph. If a starting point is specified in the path planning problem, the chromosome length can be adjusted appropriately.

[0124] To decode the chromosome and generate a path, we need to use a set nodes, which stores the endpoints of the directed edges of each node when it is the starting point, that is, the nodes that each node can reach next. For this problem, the set nodes is defined as: nodes = [[], [2, 3], [3, 4, 5], [5, 6], [7, 8], [4, 6], [7, 9], [8, 9], [9, 10],

[10] ]. Since the index of the list in Python starts at 0, and the node numbering in the graph starts at 1, the 0th element of nodes is set to an empty list for consistency. For example, the 1st element of nodes is [2, 3], which means that node 1 can reach nodes 2 and 3; the 2nd element is [3, 4, 5], which means that node 2 can reach nodes 3, 4, and 5, and so on.

[0125] The following is the decoding process of chromosome [7,3,4,6,2,5,8,10,1,9]:

[0126] Starting from node 1, search for the element [2,3] with index 1 in nodes, indicating that node 1 can reach node 2 or node 3. According to the chromosome, the priorities of nodes 2 and 3 are 3 and 4, respectively. Select node 3, which has a higher priority, as the next node to visit. Next, search for the element [5,6] with index 3 in nodes, indicating that node 3 can reach node 5 or node 6. According to the chromosome, the priorities of nodes 5 and 6 are 2 and 5, respectively. Select node 6, which has a higher priority, as the next node to visit. Repeat the above process, and the complete access path is finally obtained: 1→3→6→7→8→10.

[0127] After decoding the path, the fitness of the individual needs to be calculated. Fitness calculations are typically based on an optimization objective. In this problem, the optimization objective is the optimal path, so path length and routing hop count can be used as the optimization objective function values. Path length is defined as the cumulative sum of the weights of all directed edges along the path.

[0128] (2) Improved Monte Carlo Tree Search

[0129] The model simulates a link outage and, when a link outage occurs, uses an improved Monte Carlo tree search to find a recovery path for the damaged services. The core idea of ​​this strategy is to optimize the Monte Carlo tree search selection and simulation process using local topology information learned from pre-set paths, thereby improving search efficiency and solution quality. This strategy primarily includes the following three improvements.

[0130] The first part is weight adjustment based on pre-set path learning: In traditional MCTS, selection strategies (such as UCT) typically rely solely on node statistics (such as visit counts and reward values). To leverage the results of pre-set path learning, the improved strategy introduces local topological weights when selecting nodes. This gives higher weights to edges or nodes that frequently appear in the pre-set path, making them more likely to be selected during the search process.

[0131] The weight calculation formula is:

[0132] w(u,v)=α*freq(u,v)+(1-α)*reward(u,v);

[0133] Where w(u,v) is the weight of the edge between nodes u and v, u is a node, v is a node, freq(u,v) is the frequency of the edge in the local topology; reward(u,v) is the traditional reward value of the edge (u,v); α is the weight coefficient used to balance the information of the preset path and the traditional reward value.

[0134] The second part is the simulation process guided by local topology. During the simulation phase, traditional MCTS relies entirely on random sampling, which can lead to inefficient search. The improved strategy introduces local topology guidance to prioritize simulations of edges or nodes that frequently appear in the pre-set path. First, for each possible action at the current node, its local topological weight is calculated. Then, the actions are sorted by weight, with actions with higher weights being prioritized for simulation. This allows the simulation process to converge to a high-quality solution more quickly.

[0135] The third improvement addresses the potential for local changes in the network topology when a link failure occurs. To accommodate this dynamic nature, the improved strategy introduces a mechanism for dynamically adjusting the search scope. When a link failure is detected, the algorithm first identifies the affected services. It then identifies the local topology information corresponding to these services and focuses its search resources, prioritizing repairs on the affected paths. For unaffected areas, the algorithm maintains its original search strategy to avoid unnecessary computational overhead.

[0136] Based on the above operations, a multi-modal, multi-objective optimization algorithm based on offline learning and an improved Monte Carlo tree search was used to solve the problem. The algorithm parameters were set as follows: the initial population size was set to 100, the number of algorithm iterations was set to 50, the inertia weight w was set to 0.7298, and the coefficients C1 and C2 were set to 2.05. The experimental results, namely, the recovery time (s) as the traffic volume increases, are shown in Table 1, and the service recovery rate as the traffic volume increases, are shown in Table 2.

[0137] Table 1

[0138]

[0139] Table 2

[0140]

[0141]

[0142] As shown in Table 1, the OL-MCTS-RSA algorithm exhibits optimal recovery time performance under all traffic conditions, with an average recovery time of only 0.54 seconds, significantly outperforming other algorithms (p < 1). Especially under high traffic conditions, the OL-MCTS-RSA algorithm's recovery time is much shorter than that of other algorithms, demonstrating high real-time performance.

[0143] Table 2 shows that the OL-MCTS-RSA algorithm also performs best in terms of fault recovery rate, with an average recovery rate of 0.925. In particular, it achieves a 100% recovery rate at a low transaction volume (100) and maintains an 84.5% recovery rate at a high transaction volume (1836). While the recovery rates of all algorithms decrease with increasing transaction volume, the OL-MCTS-RSA algorithm experiences the smallest decrease (only 15.5 percentage points), demonstrating good robustness.

[0144] According to the network anti-fiber break protection planning method based on offline learning and improved MCTS in the embodiment of the present application, a preset multi-mode multi-objective optimization algorithm is used to solve the preset multi-objective optimization model of the network anti-fiber break protection planning, obtain at least one optimal backup path corresponding to each service, and learn the topology structure and link weight information of each optimal backup path; when there is a link failure, based on the preset improved Monte Carlo tree search algorithm, according to the topology structure and link weight information of the optimal backup path corresponding to the currently damaged service among all services, the target recovery path of the currently damaged service is dynamically searched, and the target recovery path is output. In this way, the problems of poor timeliness and low service recovery rate in the existing optical network recovery technology are solved, and the performance and efficiency of the network anti-fiber break protection planning are effectively improved by comprehensively optimizing the shortest recovery time and the maximum service recovery rate.

[0145] Next, a network anti-fiber break protection planning device based on offline learning and improved MCTS proposed in an embodiment of the present application will be described with reference to the accompanying drawings.

[0146] Figure 4 4 is a block diagram of a network anti-fiber break protection planning device based on offline learning and improved MCTS according to an embodiment of the present application.

[0147] like Figure 4 As shown, the network anti-fiber break protection planning device 10 based on offline learning and improved MCTS includes: a solution module 100, a learning module 200 and a recovery module 300.

[0148] Among them, the solution module 100 is used to use the preset multi-mode multi-objective optimization algorithm to solve the preset multi-objective optimization model of the network anti-fiber break protection plan to obtain at least one optimal backup path corresponding to each service; the learning module 200 is used to learn the topology structure and link weight information of each optimal backup path; the recovery module 300 is used to dynamically search for the target recovery path of the currently damaged service based on the preset improved Monte Carlo tree search algorithm when there is a link failure, according to the topology structure and link weight information of at least one optimal backup path corresponding to the currently damaged service among all services, and output the target recovery path.

[0149] Furthermore, in some embodiments, the solution module 100 is used to: divide the population into multiple sub-populations based on a preset affinity propagation clustering algorithm, wherein each sub-population evolves independently; screen out at least one non-dominated solution from all archives, classify the at least one non-dominated solution based on a preset Pareto solution set learning strategy to obtain at least one potential search area, and perform Gaussian fitting on the non-dominated solutions in each potential search area to obtain a Gaussian fitting result; determine the target search area based on the Gaussian fitting result, and learn the location information of the Pareto optimal solution according to the target search area to obtain a learning result, and generate a new solution based on the learning result.

[0150] Furthermore, in some embodiments, after generating a new solution based on the learning results, the solution module 100 is also used to dynamically adjust the ratio of the local search population and the global search population based on the evolutionary stage and population status of the preset multi-mode multi-objective optimization algorithm.

[0151] Furthermore, in some embodiments, the recovery module 300 is used to: in the node selection stage, adjust the weights of the edges and / or the weights of the nodes in the preset path based on a preset topology structure weight adjustment mechanism; in the simulation stage, calculate the weight value corresponding to each action of the current node, determine the target simulation action based on the weight value corresponding to each action of the current node, and simulate the target simulation action; when there is a link failure, determine at least one damaged service and the local topology information corresponding to each damaged service, and search for the target recovery path corresponding to each damaged service based on the local topology information corresponding to each damaged service and at least one optimal backup path.

[0152] Furthermore, in some embodiments, the objective function of the preset multi-objective optimization model of the network anti-fiber break protection plan includes at least one of a function of minimizing the service recovery time and a function of maximizing the service recovery rate.

[0153] It should be noted that the above explanation of the embodiment of the network anti-fiber break protection planning method is also applicable to the network anti-fiber break protection planning device of this embodiment, and will not be repeated here.

[0154] According to the network anti-fiber break protection planning device based on offline learning and improved MCTS according to the embodiment of the present application, a preset multi-mode multi-objective optimization algorithm is used to solve the preset multi-objective optimization model of the network anti-fiber break protection planning, obtain at least one optimal backup path corresponding to each service, and learn the topology structure and link weight information of each optimal backup path; when there is a link failure, based on the preset improved Monte Carlo tree search algorithm, according to the topology structure and link weight information of the optimal backup path corresponding to the currently damaged service among all services, the target recovery path of the currently damaged service is dynamically searched, and the target recovery path is output. In this way, the problems of poor timeliness and low service recovery rate in the existing optical network recovery technology are solved, and the performance and efficiency of the network anti-fiber break protection planning are effectively improved by comprehensively optimizing the shortest recovery time and the maximum service recovery rate.

[0155] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:

[0156] Memory 501 , processor 502 , and computer programs stored in the memory 501 and executable on the processor 502 .

[0157] When the processor 502 executes the program, the network anti-fiber break protection planning method based on offline learning and improved MCTS provided in the above embodiment is implemented.

[0158] Furthermore, the electronic device further includes:

[0159] The communication interface 503 is used for communication between the memory 501 and the processor 502 .

[0160] The memory 501 is used to store computer programs that can be run on the processor 502 .

[0161] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0162] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0163] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can communicate with each other through an internal interface.

[0164] The processor 502 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0165] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned network anti-fiber break protection planning method based on offline learning and improved MCTS.

[0166] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0167] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0168] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A network anti-fiber break protection planning method based on offline learning and improved MCTS, characterized by: The following steps are involved: Using a preset multi-mode multi-objective optimization algorithm to solve the preset multi-objective optimization model of the network anti-fiber break protection plan, at least one optimal backup path corresponding to each service is obtained; Learn the topology and link weight information of each optimal backup path; When there is a link failure, based on the preset improved Monte Carlo tree search algorithm, according to the topological structure and link weight information of at least one optimal backup path corresponding to the currently damaged business among all businesses, the target recovery path of the currently damaged business is dynamically searched and the target recovery path is output.

2. The method according to claim 1, characterized in that The preset multi-mode multi-objective optimization algorithm includes: The population is divided into multiple sub-populations based on the preset affinity propagation clustering algorithm, where each sub-population evolves independently; At least one non-dominated solution is screened out from all archives, and based on a preset Pareto solution set learning strategy, the at least one non-dominated solution is classified to obtain at least one potential search area, and Gaussian fitting is performed on the non-dominated solutions in each potential search area to obtain a Gaussian fitting result; A target search area is determined based on the Gaussian fitting result, and location information of a Pareto optimal solution is learned according to the target search area to obtain a learning result, and a new solution is generated according to the learning result.

3. The method according to claim 2, characterized in that After generating a new solution according to the learning result, the method further includes: The ratio of the local search population to the global search population is dynamically adjusted according to the evolutionary stage of the preset multi-mode multi-objective optimization algorithm and the state of the population.

4. The method according to claim 1, wherein The preset improved Monte Carlo tree search algorithm includes: In the node selection phase, the weights of the edges and / or nodes in the preset path are adjusted based on a preset topological structure weight adjustment mechanism; In the simulation phase, the weight value corresponding to each action of the current node is calculated, the target simulation action is determined according to the weight value corresponding to each action of the current node, and the target simulation action is simulated; When the link failure occurs, at least one damaged service and local topology information corresponding to each damaged service are determined, and a target recovery path corresponding to each damaged service is searched based on the local topology information corresponding to each damaged service and at least one optimal backup path.

5. The method according to claim 1, wherein The objective function of the preset multi-objective optimization model of the network anti-fiber break protection planning includes at least one of a service recovery time minimization function and a service recovery rate maximization function.

6. A network anti-fiber break protection planning device based on offline learning and improved MCTS, characterized in that: include: A solution module is used to solve a preset multi-objective optimization model of the network anti-fiber break protection plan using a preset multi-mode multi-objective optimization algorithm to obtain at least one optimal backup path corresponding to each service; A learning module is used to learn the topology structure and link weight information of each optimal backup path; The recovery module is used to dynamically search for a target recovery path for a currently damaged service based on a preset improved Monte Carlo tree search algorithm when a link failure occurs, according to the topological structure and link weight information of at least one optimal backup path corresponding to the currently damaged service among all services, and output the target recovery path.

7. The device according to claim 6, characterized in that The solution module is used to: The population is divided into multiple sub-populations based on the preset affinity propagation clustering algorithm, where each sub-population evolves independently; At least one non-dominated solution is screened out from all archives, and based on a preset Pareto solution set learning strategy, the at least one non-dominated solution is classified to obtain at least one potential search area, and Gaussian fitting is performed on the non-dominated solutions in each potential search area to obtain a Gaussian fitting result; A target search area is determined based on the Gaussian fitting result, and location information of a Pareto optimal solution is learned according to the target search area to obtain a learning result, and a new solution is generated according to the learning result.

8. The device according to claim 7, characterized in that After generating a new solution according to the learning result, the solution module is further configured to: The ratio of the local search population to the global search population is dynamically adjusted according to the evolutionary stage of the preset multi-mode multi-objective optimization algorithm and the state of the population.

9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the network anti-fiber break protection planning method based on offline learning and improved MCTS as described in any one of claims 1 to 5.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program is executed by a processor to implement the network anti-fiber break protection planning method based on offline learning and improved MCTS as described in any one of claims 1 to 5.

Citation Information

Cited By

  • Heuristic-based cell PCI (Peripheral Component Interconnect) planning method and system

    CN122160781A

  • Intelligent substation secondary circuit fault arc online monitoring method and device

    CN122193838A