N-1 maintenance planning method for communication network based on dynamic risk map
By using a dynamic risk mapping method, high-risk areas are identified and maintenance batches are optimized, which solves the network overload problem caused by uncoordinated resource scheduling and traffic redistribution in existing technologies, thereby improving the stability and resource efficiency of communication networks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANXI ELECTRIC POWER CO POWER COMM CENT
- Filing Date
- 2026-04-08
- Publication Date
- 2026-07-21
AI Technical Summary
The existing N-1 maintenance plan scheduling method for communication networks fails to effectively coordinate resource allocation, resulting in the decentralized execution of maintenance tasks. It is impossible to predict and avoid the risk of overload in surrounding areas caused by traffic redistribution, which affects network redundancy and overall operational stability.
Based on the dynamic risk map, a risk level distribution map is generated by acquiring risk assessment data. Geographically adjacent nodes are clustered to form a high-risk maintenance priority cluster. The shortest path algorithm is used to optimize the resource scheduling path, and the traffic load changes during maintenance are simulated to dynamically adjust the maintenance execution order to avoid network overload. Global stability iterative evaluation is adopted to ensure the stability of the final maintenance plan.
It achieves efficient resource utilization, reduces maintenance costs, avoids network overload, ensures stable operation and service continuity of the communication network in the N-1 maintenance mode, and meets the requirements of optimal resource utilization and stable network operation.
Smart Images

Figure CN122437761A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication network maintenance technology, and in particular to a method for scheduling N-1 maintenance plans for communication networks based on dynamic risk maps. Background Technology
[0002] In the field of modern communication network operation and maintenance, ensuring network stability and reliability has always been a core issue. As the infrastructure for information transmission in modern society, the operational quality of communication networks directly affects the normal operation of the social economy and the continuous supply of public services. With the increasing scale of networks, the diversification of service types, and the growing complexity of network structures, how to efficiently complete the maintenance of network equipment and lines while ensuring uninterrupted service has become a common challenge for operators and equipment manufacturers. Against this backdrop, the N-1 maintenance plan has been widely adopted as an important operation and maintenance strategy. N-1 maintenance refers to ensuring that when one component in a network (such as nodes or lines) is taken out of service for maintenance, the remaining N-1 components can continue to carry all service traffic without causing network overload or service interruption. The core of this maintenance model lies in verifying the network's redundancy and resilience, and it is an important means of ensuring the high availability of communication networks.
[0003] However, existing technologies have significant limitations when implementing the N-1 maintenance plan. Early approaches relied primarily on manual experience and fixed maintenance cycles, scheduling maintenance tasks based on equipment importance or fault history. This static scheduling method often struggles to adapt to dynamically changing network loads. With the development of data analytics, some methods have introduced risk assessment mechanisms, predicting equipment and line failure probabilities and prioritizing higher-risk areas. However, these methods often treat maintenance tasks as independent units, focusing on individual nodes or lines without considering the overall network operation. Regarding resource allocation, existing solutions typically use proximity to dispatch maintenance personnel and equipment. However, this locally optimal approach often ignores the interrelationships between multiple maintenance tasks, leading to frequent long-distance deployments and resource waste. More importantly, existing technologies fail to adequately consider the redistribution of network traffic during maintenance when arranging the maintenance sequence. When a certain area is under concentrated maintenance, surrounding equipment may experience new performance bottlenecks due to excessive load. This causes the redundancy capability that the N-1 maintenance was originally intended to verify to fail due to unreasonable scheduling. This phenomenon of neglecting one aspect for another is difficult to avoid effectively in existing solutions. Summary of the Invention
[0004] Therefore, the technical problem to be solved by the present invention is to overcome the shortcomings of the prior art in the N-1 maintenance plan arrangement of communication networks, which only focuses on the maintenance arrangement of a single device or line, resulting in a lack of overall planning of resource scheduling, decentralized execution of maintenance tasks, and inability to effectively predict and avoid the risk of overload in the surrounding area caused by traffic redistribution during centralized maintenance, making it difficult to guarantee the network's redundancy and overall operational stability. The present invention provides a communication network N-1 maintenance plan arrangement method based on dynamic risk map, which can reduce resource scheduling costs by clustering and integrating geographically adjacent high-risk nodes, avoid local overload by simulating traffic carrying capacity changes during maintenance and dynamically adjusting the execution order, and ensure that the final maintenance plan meets the requirements of optimal resource allocation and stable network operation through global stability iterative evaluation.
[0005] To address the aforementioned technical problems, this invention provides a method for scheduling N-1 maintenance plans for communication networks based on dynamic risk maps, comprising the following steps: Acquire risk assessment data, including the failure probability and impact range of nodes and lines in the communication network, and generate a risk level distribution map; Based on the risk level distribution map, nodes that are geographically close are clustered to identify at least one high-risk maintenance priority cluster. Analyze the connection relationships between adjacent nodes in the high-risk maintenance priority cluster. When the connection relationship meets the set connection strength condition, the corresponding nodes are assigned to the same maintenance batch, forming a preliminary batch division set. Based on the initial batch partitioning set, the resource scheduling path between each batch is calculated using the shortest path algorithm, and the resource allocation scheme is determined with the goal of minimizing scheduling cost. Based on the resource allocation plan, the impact of each batch of maintenance on network traffic capacity is simulated. If the simulation results show that a certain batch of maintenance will cause the network load in the surrounding area to exceed the capacity, the execution order of the batches will be adjusted to distribute the high load period and generate an optimized maintenance execution order. Based on the optimized maintenance execution sequence, its impact on the stability of the entire communication network is assessed. If the assessment results do not meet the preset stability indicators, clustering, resource allocation, and simulation optimization are re-executed until a final maintenance plan that meets the stability indicators is obtained.
[0006] In one embodiment of the present invention, the risk assessment data includes: counting the number of failures of each node and line in the historical operating cycle, calculating the failure frequency per unit time as the failure probability; and when a node or line fails, counting the number and distribution range of other nodes and lines affected by it and causing service interruption, as the scope of influence.
[0007] In one embodiment of the present invention, generating a risk level distribution map includes: Obtain the maximum failure probability value P among all nodes and lines in the entire network. max and the minimum failure probability value P min And the maximum range of influence value R max and the minimum influence range value R min ; For any node or line, according to the formula: P norm =(PP min ) / (P max -P min ); Calculate its normalized failure probability value P norm , where P is the original fault probability value of the node or line; According to the formula: R norm =(RR min ) / (R max -R min ); Calculate its normalized influence range value R norm , where R is the original influence range value of the node or line; P norm and R norm Multiply by the product to obtain the risk coefficient S of the node or line; Using the network topology map as the base map, each node and line is colored with a grayscale value or RGB color value corresponding to the risk coefficient S of the node or line, forming a spatial distribution map in which the degree of risk is represented by the color intensity.
[0008] In one embodiment of the present invention, nodes that are geographically close are clustered to identify at least one high-risk maintenance priority cluster, including: Set the clustering distance threshold D and the risk coefficient threshold S. th ; Starting with any node not belonging to any cluster, search for other nodes whose geographical distance is less than the clustering distance threshold D. If the risk coefficient S of the searched node is greater than the risk coefficient threshold S... th If so, the node and the starting node will be grouped into the same temporary cluster; Repeat the above search process until there are no more risk coefficients greater than the risk coefficient threshold S within the range of the temporary cluster's perimeter that are less than the clustering distance threshold D. th The unassigned nodes will be identified as a high-risk, priority maintenance cluster for the temporary cluster. Iterate through all nodes in the network, repeating the above steps until all risk coefficients S are greater than the risk coefficient threshold S. th All nodes were assigned to the corresponding high-risk maintenance priority cluster.
[0009] In one embodiment of the present invention, determining that the connection relationship meets the set connection strength conditions includes: obtaining the link type, link bandwidth and physical distance between adjacent nodes; when the link type is the same, the link bandwidth can carry the service traffic transferred during maintenance, and the physical distance supports the round-trip scheduling of maintenance personnel in a single working period, the connection relationship is determined to meet the set connection strength conditions.
[0010] In one embodiment of the present invention, the resource scheduling path between batches is calculated using a shortest path algorithm, including: Determine the center coordinates of each maintenance batch, and use the center position as the path node; Obtain the location coordinates of the resource storage point, as well as the road traffic distances between any two path nodes and between the resource storage point and each path node; Starting from the resource storage point, select the path node that is closest to the current node in terms of road traffic distance from the unvisited path node in turn as the next node to be visited, and mark the node as visited; Repeat the above steps until all path nodes have been visited once. Returning to the resource storage point from the last visited path node, the distances of all the road segments traversed in the above process are added together to obtain the total scheduling distance, and the path corresponding to the total scheduling distance is determined as the resource scheduling path.
[0011] In one embodiment of the present invention, simulating the impact of performing each batch of maintenance on network traffic capacity includes: removing the nodes and lines of the batch to be maintained from the current network topology, simulating the redistribution of traffic on each link in the remaining network, calculating the ratio of the simulated traffic value of each link to the rated capacity of the link as the load rate, and determining that the batch maintenance will cause the network load in the surrounding area to exceed the carrying capacity if the load rate of any link exceeds a preset overload threshold.
[0012] In one embodiment of the present invention, adjusting the execution order of batches to distribute high-load periods includes: acquiring multiple maintenance batches that cause overload on the same link, arranging these batches at intervals on the time axis, inserting at least one other maintenance batch that will not cause link overload between two adjacent maintenance batches that cause link overload, and ensuring that at most one adjacent batch is under maintenance during any maintenance period.
[0013] In one embodiment of the present invention, evaluating its impact on the stability of the entire communication network includes: Following the optimized maintenance execution sequence, the load conditions of each link during each batch of maintenance are simulated sequentially; If, during any maintenance batch, the load on any link exceeds its rated capacity, it is determined that the maintenance execution sequence will cause network overload and fail to meet stability requirements. If the load on all links does not exceed their rated capacity during the execution of all maintenance batches, then the maintenance execution sequence is deemed to meet the stability requirements.
[0014] In one embodiment of the present invention, re-performing clustering, resource allocation, and simulation optimization includes: In the maintenance execution order generated in the previous iteration, obtain the number of link overloads that caused overload in the surrounding area when each maintenance batch was executed, and use the number of overloads as the feedback coefficient F of the area where the batch is located. According to the formula: D new =D old ×(1-αF) adjusts the clustering distance threshold, where D old The clustering distance threshold used in the previous iteration is α, which is the adjustment step size factor. With the adjusted clustering distance threshold D new Repeat the steps of clustering geographically proximate nodes and subsequent steps.
[0015] The technical solution of the present invention has the following advantages compared with the prior art: The present invention discloses an N-1 maintenance plan scheduling method for communication networks based on dynamic risk maps. By introducing dynamic risk maps, it achieves accurate identification and priority processing of high-risk areas in the communication network. By clustering geographically adjacent nodes and optimizing the division of maintenance batches, it effectively avoids the problems of independent maintenance tasks and lack of overall planning in traditional solutions. The calculation of resource scheduling paths and simulation of traffic carrying capacity significantly reduce resource waste and avoid network overload that may occur during maintenance, thereby ensuring the stable operation of the network and the continuity of services under the N-1 maintenance mode. Attached Figure Description
[0016] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings, wherein: Figure 1 This is a flowchart of the steps of the communication network N-1 maintenance plan scheduling method based on dynamic risk map of the present invention; Figure 2 This is a flowchart of the steps in generating a visualized risk level distribution map according to the present invention; Figure 3 This is a flowchart of the cluster analysis steps of the present invention; Figure 4 This is a flowchart of the steps in this invention to calculate the resource scheduling path between batches; Figure 5This is a flowchart illustrating the steps of this invention to simulate the impact of each batch of maintenance on network traffic capacity. Figure 6 This is a flowchart illustrating the steps of clustering, resource allocation, and simulation optimization in this invention. Detailed Implementation
[0017] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.
[0018] Reference Figure 1 As shown, this invention provides a method for scheduling N-1 maintenance plans for communication networks based on dynamic risk maps. First, it starts with network operation data to obtain risk assessment data, including the probability and impact range of node and line failures. Based on this data, a visualized risk distribution map is generated, thus comprehensively understanding the risk situation in different areas of the network. On this basis, cluster analysis is performed on geographically adjacent nodes according to the risk distribution map to divide them into high-risk maintenance priority clusters. This approach integrates dispersed high-risk points according to spatial proximity, avoiding the waste of resources from frequent cross-regional relocation. Subsequently, the connection relationships between adjacent nodes in the high-risk maintenance priority clusters are further analyzed. When the connection relationship meets the set connection strength conditions, the corresponding nodes are grouped into the same maintenance batch, forming a preliminary batch division set. This step, by considering the coupling degree between equipment, ensures that maintenance operations within the same batch will not generate additional risks due to mutual interference. After the batch division is completed, the method calculates the resource scheduling path between each batch based on the preliminary batch division set using a shortest path algorithm. The resource allocation scheme is determined with the goal of minimizing scheduling costs, thereby achieving optimal spatial allocation of maintenance personnel and equipment. To mitigate the impact of maintenance on network operation, the method simulates the impact of each batch of maintenance on network traffic capacity based on a resource allocation scheme. If the simulation results show that a certain batch of maintenance causes the network load in the surrounding area to exceed its capacity, the execution order of this batch is automatically adjusted to distribute high-load periods, generating an optimized maintenance execution order. Finally, the impact of the optimized maintenance execution order on the stability of the entire communication network is evaluated. If the evaluation results do not meet the preset stability indicators, the iterative process of clustering, resource allocation, and simulation optimization is repeated until a final maintenance plan that meets the stability indicators is obtained.
[0019] Compared with existing technologies, this invention achieves several beneficial effects through the above-mentioned technical solutions, including: First, in terms of resource utilization efficiency, by clustering high-risk nodes in geographical proximity, dispersed maintenance tasks are integrated into regionalized work clusters, significantly reducing the frequent movement of maintenance personnel and equipment over long distances, lowering scheduling costs and time consumption. Simultaneously, resource scheduling path planning based on the shortest path algorithm further optimizes the flow of materials between batches, making the overall resource allocation scheme cost-optimal. Second, in terms of network stability assurance, the method identifies potential overload risks in advance by simulating traffic load changes during maintenance, and smooths out peaks and valleys by adjusting the execution order of maintenance batches, avoiding network congestion in surrounding areas caused by localized concentrated maintenance. This ensures that even in an N-1 scenario where any component exits service, the remaining network can still smoothly carry all traffic, truly leveraging the protective role of redundancy design. More importantly, an iterative evaluation mechanism for global stability is introduced, feeding back the results of each simulation optimization to the initial clustering stage. Through cyclic verification, it ensures that the final generated maintenance plan not only meets local resource optimization but also conforms to the overall network operation stability requirements. This design approach of dynamic adjustment and closed-loop optimization transforms the scheduling of the N-1 maintenance plan from a static and isolated arrangement to a dynamic and systematic collaborative scheduling, providing a reliable guarantee for the safe and stable operation of the communication network.
[0020] Specifically, this application further proposes specific methods for obtaining risk assessment data, including counting the number of failures of each node and line in the historical operating cycle, calculating the failure frequency per unit time as the failure probability; and when a node or line fails, counting the number and distribution range of other nodes and lines affected by it that cause business interruption, as the scope of impact.
[0021] The calculation of failure probability specifically refers to calculating the frequency of failures per unit time by statistically analyzing the number of failures of each node and line in the communication network within a historical operating period. For example, data sources such as network management systems, alarm logs, and maintenance records can be used to extract the specific time points and frequency of failures for each node (such as routers, switches, servers, etc.) and each line (such as fiber optic links, microwave links, etc.) within the past year, two years, or longer (i.e., the historical operating period). Then, dividing the statistically analyzed number of failures by the length of the historical operating period yields the frequency of failures for that node or line per unit time, which is used as its failure probability. This statistical method based on historical data can objectively reflect the reliability level of network components.
[0022] The statistical analysis of the impact range specifically refers to the process of analyzing or simulating the impact of a failure in a node or line within a communication network on network services and topology. This involves determining the number and distribution of other nodes and lines affected by the failure, leading to service interruptions. For example, using network topology information and service routing configurations, when a node or line is marked as faulty, path calculation or service interruption analysis tools can identify all service flows dependent on the faulty component, affected user terminals, and other upstream or downstream nodes and lines that are consequently unable to function properly. The impact range can be quantified as the number of affected nodes, affected lines, affected users, and even the geographical area of the affected region. This method comprehensively assesses the potential cascading effects and service losses caused by a single failure event.
[0023] In this embodiment, a problem that needs to be solved is how to transform abstract risk assessment data into an intuitive and quantifiable spatial distribution map to facilitate subsequent risk identification and maintenance plan scheduling. Simply obtaining the probability of failure and the scope of impact cannot directly reflect the severity of risk and its geographical distribution characteristics in each area of the network, which may lead to inaccurate identification of high-risk areas, thereby affecting the effectiveness of maintenance plans.
[0024] In this regard, refer to Figure 2 As shown, this application further proposes a method for generating a risk level distribution map, which specifically includes the following steps: First, obtain the maximum failure probability value P among all nodes and lines in the entire network. max and the minimum failure probability value P min And the maximum range of influence value R max and the minimum influence range value R min This step aims to establish a benchmark for risk assessment data by obtaining the extreme values of failure probability and impact range across the entire network, providing a reference range for subsequent normalization processing. This ensures that risk data from different nodes and lines are compared on a uniform scale, avoiding assessment biases caused by differences in the units of the original data. Specifically, the system will traverse the historical operational data of all communication nodes and lines, statistically analyze their failure probability and impact range, and then filter out the maximum and minimum values.
[0025] Next, for any node or line, according to the formula: P norm =(PP min ) / (P max -P min ); Calculate its normalized fault probability value P norm , where P is the original fault probability value of the node or line. Normalized fault probability value P normThe calculation maps the original failure probability P to the interval [0, 1]. This process eliminates the influence of the original failure probability's dimensions, allowing for unbiased combination with other normalized indicators. This enables comparisons of failure probabilities across different network sizes or device types within a unified framework.
[0026] At the same time, according to the formula: R norm =(RR min ) / (R max -R min ); Calculate its normalized influence range value R norm Where R is the original influence range value of the node or line. Normalized influence range value R norm The calculation of the impact range R is similar to the normalization of the failure probability, aiming to map the original impact range R to the interval [0, 1]. The impact range is typically measured in terms of the amount of traffic affected, the number of users, or the number of links interrupted. Normalization eliminates the differences in units of measurement for the impact range, making it consistent with the normalized failure probability P. norm This provides a basis for subsequent comprehensive risk assessment, as the two methods are comparable.
[0027] Then, P norm and R norm Multiplying these yields the risk coefficient S of the node or line; the risk coefficient S is calculated by multiplying the normalized failure probability P... norm and the normalization influence range R norm Performing a product operation. This multiplicative combination method reflects the essence of risk: risk depends not only on the probability of an event occurring (failure probability) but also on the consequences of the event (scope of impact). Multiplication allows for a more sensitive identification of truly high-risk points that are both highly probable and have a large impact, because the product S only increases significantly when both are high. For example, a node with a high failure probability but a small impact might have a lower S value than a node with a moderate failure probability but a large impact, thus more accurately reflecting the actual risk priority.
[0028] Finally, using the network topology map as a base map, each node and line is colored with a grayscale value or RGB color value corresponding to its risk coefficient S, forming a spatial distribution map where color intensity represents the degree of risk. This step visualizes the abstract risk coefficient S. The network topology map provides geographical location information for communication network devices and links. By mapping the calculated risk coefficient S to visually distinguishable grayscale values (e.g., the larger the S value, the darker the color) or RGB color values (e.g., from green to yellow to red, representing low to high risk), the risk distribution of the entire network can be visually represented in geospatial space. This visualization method allows maintenance personnel to easily identify high-risk areas; for example, areas displayed in dark red or dark gray on the map are priority maintenance areas. This greatly improves the efficiency and accuracy of risk identification.
[0029] Reference Figure 3 As shown, based on the risk level distribution map, this application further proposes a method for clustering geographically adjacent nodes to identify at least one high-risk maintenance priority cluster. The specific steps include: First, set the clustering distance threshold D and the risk coefficient threshold S. th The clustering distance threshold D limits the geographical proximity between nodes; only nodes whose geographical distance is less than D can be considered to belong to the same cluster. Risk coefficient threshold S th This is used to filter out high-risk nodes that truly require priority maintenance; only when the risk coefficient S of a node is greater than this threshold S... th Only when certain conditions are met will a cluster be eligible for inclusion in the high-risk maintenance priority cluster. The setting of these two thresholds can be based on a comprehensive consideration of factors such as historical data analysis, expert experience, network size, and desired maintenance granularity to ensure the rationality and effectiveness of the clustering results.
[0030] Next, starting with any node not belonging to any cluster, search for other nodes whose geographical distance is less than the clustering distance threshold D. If the risk coefficient S of the searched node is greater than the risk coefficient threshold S... th If a node is found to be a seed node (starting node), it will be grouped into the same temporary cluster as the starting node. This step describes a cluster growth mechanism based on the seed node (starting node). First, the system selects a node that is not yet included in any cluster as the starting point. Then, it searches for its geographically neighboring regions centered on the starting node. During the search, other nodes are only included in the currently being built temporary cluster if they meet two conditions: first, their geographical distance from the starting node is less than a preset clustering distance threshold D, ensuring the geographical compactness of the cluster; second, their own risk coefficient S is greater than a preset risk coefficient threshold S0. thThis ensures the high-risk nature of the cluster. This dual screening mechanism ensures that the members of the temporary cluster are geographically proximate and high-risk nodes.
[0031] Subsequently, the search process is repeated until there are no more risk coefficients greater than the risk coefficient threshold S within the range of the temporary cluster's perimeter that are less than the clustering distance threshold D. th For any nodes not yet assigned to a cluster, the temporary cluster is identified as a high-risk, priority maintenance cluster. This search process continues, adding eligible nodes to the temporary cluster until no new nodes not assigned to any cluster and with a risk coefficient S higher than the risk coefficient threshold S are found within the geographical boundaries of the temporary cluster (i.e., within a radius equal to the clustering distance threshold D centered on any node in the cluster). th The nodes. At this point, the temporary cluster is regarded as a complete and independent high-risk maintenance priority cluster, whose internal nodes are geographically closely connected and all have a high risk level.
[0032] Finally, iterate through all nodes in the network, repeating the above steps until all risk coefficients S are greater than the risk coefficient threshold S. th All nodes are assigned to the corresponding high-risk maintenance priority cluster. To ensure that all nodes meeting the high-risk criteria are effectively identified and managed, the above cluster construction process is performed throughout the entire communication network. The system continuously selects new nodes that are not yet assigned to any cluster and whose risk coefficient S is greater than the risk coefficient threshold S. th Starting from a specific node, the process of cluster growth and determination is repeated. This traversal mechanism ensures that all high-risk nodes are included in the corresponding priority maintenance clusters, avoiding omissions and providing comprehensive and accurate basic data for subsequent maintenance planning.
[0033] In some embodiments described above, this application proposes clustering geographically proximate nodes based on a risk level distribution map and analyzing the connectivity between adjacent nodes in a high-risk maintenance priority cluster to group nodes that meet set connectivity strength conditions into the same maintenance batch. However, in practice, accurately and effectively determining whether the connectivity between adjacent nodes meets the connectivity strength conditions required for maintenance batch division, to ensure the rationality and operability of maintenance batches, is a problem that requires careful consideration. If the connectivity strength conditions are not clearly defined or are inaccurately determined, it may lead to unreasonable maintenance batch division, affecting maintenance efficiency and network stability.
[0034] In response, this application further proposes a method for determining whether a connection relationship meets the set connection strength conditions. This method includes obtaining the link type, link bandwidth, and physical distance between adjacent nodes. When the link type is the same, the link bandwidth is sufficient to carry the service traffic transferred during maintenance, and the physical distance supports the round-trip scheduling of maintenance personnel in a single working period, the connection relationship is determined to meet the set connection strength conditions.
[0035] Specifically, when analyzing the connectivity between adjacent nodes in a high-risk maintenance priority cluster, it is necessary to obtain the detailed attributes of the links between these adjacent nodes. First, the link type between adjacent nodes needs to be determined, such as whether it is a fiber optic link, cable link, or microwave link. Different link types may correspond to different maintenance processes, required tools, and professional skills. Therefore, grouping nodes with the same link type into the same batch helps to unify the allocation of maintenance resources and improve efficiency. Second, the link bandwidth between adjacent nodes needs to be obtained, that is, the maximum data traffic that the link can carry. During N-1 maintenance, the service traffic on the node or line being maintained needs to be transferred to other links. Therefore, determining whether the link bandwidth can handle the transferred service traffic during maintenance is crucial. This requires that even if some network resources are occupied during maintenance, the bandwidth of the remaining links is sufficient to handle all service traffic, avoiding network congestion or service interruption. Third, the physical distance between adjacent nodes needs to be obtained, that is, the actual geographical distance between the two nodes. Physical distance is a key factor in evaluating the efficiency of maintenance personnel scheduling, as it directly affects the time required for maintenance personnel to travel from one node to another. When physical distance allows maintenance personnel to complete round-trip scheduling within a single work period, it means that maintenance tasks within that batch can be completed efficiently within a single workday, avoiding the additional costs and management complexities associated with cross-day operations.
[0036] After forming the initial batch division set, resource scheduling needs to be performed for each maintenance batch, referring to... Figure 4 As shown, this application further proposes a method for calculating resource scheduling paths between batches using a shortest path algorithm. This method includes the following steps: First, determine the coordinates of the center location of each maintenance batch, and use these center locations as path nodes. In practice, a maintenance batch may contain multiple geographically close nodes or routes. To simplify path calculation and improve efficiency, each maintenance batch can be considered as a whole, and its geographical center location can be used to represent the batch. The coordinates of this center location can be obtained by calculating the average of the geographical coordinates (such as latitude and longitude) of all nodes or routes within the batch, or by using a weighted average (e.g., weighted according to the importance or size of the nodes). Using these center locations as path nodes effectively abstracts the maintenance batches, facilitating subsequent path planning algorithm processing.
[0037] Secondly, the location coordinates of resource storage points, as well as the road distances between any two path nodes and between resource storage points and each path node, are obtained. Resource storage points typically refer to warehouses or bases that store maintenance equipment, spare parts, and personnel. Their location coordinates can be obtained through Global Positioning System (GPS) data or Geographic Information System (GIS) databases. Similarly, the center location coordinates of each maintenance batch are obtained in a similar manner. The road distances between any two path nodes and between resource storage points and each path node can be obtained by calling map service APIs (such as Amap, Baidu Maps, Google Maps, etc.) to obtain real-time or historical road distance data, or by performing path analysis calculations based on GIS data. This distance data forms the basis for subsequent shortest path algorithm decisions.
[0038] Next, starting from the resource storage point, the algorithm sequentially selects the node with the closest road traffic distance to the current node from among the unvisited path nodes as the next node to be visited, and marks this node as visited. This step describes a greedy strategy: starting from the current position, it always selects the closest unvisited maintenance batch for scheduling. In practice, a set of visited nodes and a set of unvisited nodes can be maintained. In each step, the algorithm traverses all nodes in the unvisited node set, calculates their road traffic distances to the current node, and selects the node with the smallest distance as the next target. Once selected, this node is moved from the unvisited set to the visited set. This process is repeated until all path nodes have been visited once. This loop continues until the unvisited node set is empty, meaning all maintenance batches have been scheduled on the path.
[0039] Finally, returning to the resource storage point from the last visited path node, the total scheduling distance is obtained by summing the distances of all road segments traversed in sequence during the process. The path corresponding to this total scheduling distance is then determined as the resource scheduling path. After all maintenance batches have been visited, the resource scheduling vehicle or personnel need to return to the resource storage point. Therefore, the last segment of the path is the return trip from the last visited maintenance batch to the resource storage point. The total distance of this resource scheduling is obtained by summing the road traffic distances of all road segments traversed in sequence throughout the entire scheduling process. This path with the shortest total distance is determined as the resource scheduling path, representing the complete journey of the resource from the storage point, sequentially visiting all maintenance batches, and finally returning to the storage point.
[0040] During the N-1 maintenance plan preparation process for the communication network, it is necessary to accurately assess the impact of each maintenance batch on the network traffic carrying capacity to avoid network overload during the maintenance period. If the simulation assessment is inaccurate, it may cause local or global network outages during the actual execution of the maintenance plan, thereby affecting the stable operation of the communication network.
[0041] Reference Figure 5 As shown, this application further proposes a method for simulating the impact of performing each batch of maintenance on network traffic capacity, which includes: removing the nodes and lines of the batch to be maintained from the current network topology, simulating the redistribution of traffic on each link in the remaining network, calculating the ratio of the simulated traffic value of each link to the rated capacity of the link as the load rate, and determining that the batch maintenance will cause the network load in the surrounding area to exceed the carrying capacity if the load rate of any link exceeds a preset overload threshold.
[0042] Specifically, when simulating the impact of various maintenance batches on network traffic capacity, the nodes and lines of the batches to be maintained first need to be logically removed from the current network topology. This operation aims to simulate these network components being offline or unavailable during maintenance, thereby reflecting the changes in network topology caused by maintenance. In practice, these nodes and lines can be marked as unavailable in the network topology data model, or the connection weights associated with these components can be set to infinity during traffic calculation, preventing them from being selected in routing calculations.
[0043] Subsequently, the redistribution of traffic across the remaining network links is simulated. When some nodes and lines are removed, the original traffic paths may be interrupted, and service traffic in the network will seek new available paths for transmission. This step involves running traffic engineering algorithms, such as shortest path-based algorithms or multi-path routing algorithms, to recalculate the traffic paths for each source-destination traffic demand pair based on the remaining network topology and link capacity, and allocating traffic to these paths, thereby obtaining simulated traffic values for each link during the maintenance period.
[0044] Next, the load factor is calculated as the ratio of the simulated traffic value to the rated capacity of each link. The rated capacity of a link refers to the maximum traffic that the link is designed to carry. By dividing the simulated traffic value by the rated capacity, the load pressure on each link during maintenance can be quantitatively assessed, intuitively reflecting its utilization rate. For example, if a link has a rated capacity of 10Gbps and a simulated traffic value of 8Gbps, then its load factor is 0.8.
[0045] Finally, if the load rate of any link exceeds the preset overload threshold, it is determined that the batch maintenance will cause the network load in the surrounding area to exceed its carrying capacity. The preset overload threshold is a configurable percentage value, such as 80%, 90%, or 100%, and its setting is based on the network's redundancy design, service importance, and historical operating experience. Once the load rate of any link is found to exceed this threshold, it indicates that the maintenance batch may cause local or overall network overload risk.
[0046] Specifically, by simulating the impact of each batch of maintenance on network traffic capacity, it is possible to identify which maintenance batches cause the network load in the surrounding area to exceed its capacity. To address this issue, this application further proposes a method to adjust the execution order of batches to distribute high-load periods. Specifically, after simulating the impact of each batch of maintenance on network traffic capacity, the system identifies which maintenance batches cause the load rate of a specific communication link to exceed a preset overload threshold. At this point, it is necessary to obtain all maintenance batches that cause overload on the same link. This can be achieved by establishing a data structure, such as a list or hash table, where the key is the link identifier that caused the overload, and the value is the set of all maintenance batches that caused the link to overload. When the simulation results show that a link is overloaded, the system adds the corresponding maintenance batch to the set for that link, thereby clearly identifying all potential conflicting batches.
[0047] Subsequently, to distribute network load, these overloaded batches need to be arranged at intervals along the timeline. This can be achieved using heuristic or optimization algorithms. For example, for a set of batches causing overload on the same link, they can be removed from the original time series. Then, when re-inserting them, priority should be given to time periods with lower current network load and sufficient time intervals between them and the scheduled overloaded batches. This strategy aims to avoid overloaded batches being too concentrated in time, thereby effectively distributing network load.
[0048] To further ensure network stability, this application proposes inserting at least one other maintenance batch that will not cause link overload between two adjacent maintenance batches that would lead to link overload. When rescheduling overload batches, the system selects one or more batches as buffer batches from the set of other maintenance batches that will not cause link overload. These buffer batches can be maintenance batches that have a minor impact on the specific link or that affect other unrelated links. The scheduling algorithm prioritizes inserting these buffer batches between two adjacent overload batches, thereby effectively isolating overload batches in time and reducing network risk.
[0049] The ultimate goal is to ensure that at most one adjacent batch is under maintenance on a link during any given maintenance period. During scheduling, the system continuously monitors how many adjacent batches are under maintenance on each link within each time period. If, within a certain time period, the number of adjacent batches on a link exceeds one, the scheduling algorithm will trigger an adjustment, rescheduling one or more batches to other time periods until the condition of at most one adjacent batch being under maintenance is met. Here, adjacent batches refer to those identified in the simulation as potentially affecting the link. This constraint aims to minimize the stress on a single link within a short period, avoiding the superposition of multiple failures or overloads.
[0050] After initial optimization, this application further proposes to assess its impact on the stability of the entire communication network, including: simulating the load of each link during each batch of maintenance according to the optimized maintenance execution sequence; if the load of any link exceeds its rated capacity during the execution of any maintenance batch, it is determined that the maintenance execution sequence will cause network overload and fail to meet the stability requirements; if the load of all links does not exceed its rated capacity during the execution of all maintenance batches, it is determined that the maintenance execution sequence meets the stability requirements.
[0051] Specifically, assessing the impact on the stability of the entire communication network refers to a comprehensive and in-depth verification of the optimized maintenance execution sequence generated by the above methods. This ensures that the maintenance plan will not lead to service interruptions, performance degradation, or localized overloads during actual implementation. This assessment aims to grasp the robustness of the maintenance plan as a whole, ensuring it conforms to the N-1 principle, meaning that while one batch is being maintained, the remaining network components can still normally support services.
[0052] To achieve the above assessment, the load on each link during each maintenance batch needs to be simulated sequentially according to the optimized maintenance execution order. This means simulating the previously determined maintenance batches, which have been time-distributed and resource-optimized, one by one in the simulation environment, according to their predetermined execution order. During the simulation of each maintenance batch, the nodes and lines included in that batch will be treated as offline, and network traffic will be rerouted according to a preset routing policy. At this time, it is necessary to accurately calculate the traffic load on all remaining active links in the network and record this load data. This simulation process can be implemented using professional network simulation tools or custom algorithms based on network topology and traffic matrices to ensure the accuracy of the simulation results.
[0053] During the simulation, if any link experiences a load exceeding its rated capacity during any maintenance batch, the maintenance sequence is deemed to have caused network overload and failed to meet stability requirements. The rated capacity of a link is its maximum capacity to handle, determined during its design. If the simulation results show that even a single link's actual load exceeds its rated capacity during a maintenance batch, it indicates a flaw in the maintenance sequence, potentially leading to interruption or severe performance degradation of that link and related services. This stringent judgment ensures high network stability and avoids potential risks.
[0054] Conversely, if the load on all links does not exceed their rated capacity during the execution of all maintenance batches, the maintenance execution sequence is deemed to meet the stability requirements. This means that throughout the entire maintenance plan, regardless of which batch is under maintenance, all links in the communication network can normally handle the rerouted service traffic without overload. Only when the maintenance execution sequence passes this rigorous stability test can it be considered reliable and feasible.
[0055] In some of the embodiments described above in this application, when network stability is found to be substandard after simulating the maintenance plan, clustering, resource allocation, and simulation optimization need to be re-executed. However, simply repeating these steps without specifically adjusting key parameters may lead to inefficient iteration processes and even difficulty in converging to a final maintenance plan that meets stability requirements. In particular, if the initial clustering parameters are inappropriate, it may be impossible to effectively identify and optimize the maintenance batch division of high-risk areas.
[0056] Reference Figure 6 As shown, this application further proposes steps for re-executing clustering, resource allocation, and simulation optimization, including: First, the system obtains the number of link overloads that caused overload in the surrounding area during the execution of each maintenance batch in the maintenance execution sequence generated in the previous iteration, and uses this number of overloads as the feedback coefficient F for the area where the batch is located. This feedback coefficient F aims to quantify the negative impact of the maintenance plan on the network in the previous iteration. That is, if the execution of a certain maintenance batch causes link overload in the surrounding area, then the area where the batch is located is considered to have a high risk or room for optimization. By counting the number of link overloads, a specific value can be obtained, which intuitively reflects the degree of instability or the degree of improvement needed for the area under the current maintenance plan. For example, it can record which links were overloaded and the frequency of overload during each simulated maintenance. The feedback coefficient F can be simply defined as the number of links that caused network overload during a certain batch of maintenance, or more precisely, the sum of the overload levels of all overloaded links.
[0057] Based on this, the system will follow the formula: D new =D old ×(1-αF) adjusts the clustering distance threshold, where D old The clustering distance threshold D is the one used in the previous iteration, and α is the adjustment step size factor. The clustering distance threshold D is a key parameter that determines the size and density of high-risk maintenance priority clusters. When the feedback coefficient F (representing the degree of overload) is high, it means that the current clustering method may be too coarse or unreasonable, resulting in excessively large or concentrated maintenance batches, thus causing network overload. Through the formula, D... newIt will adaptively adjust based on the value of F. If the value of F is large (severe overload), then (1-αF) will become smaller, causing D to... new Less than D old This means reducing the clustering distance threshold. Reducing the clustering distance threshold means that clustering becomes more refined, dividing high-risk nodes into smaller, more dispersed clusters, which may result in smaller, more manageable maintenance batches, reducing the impact of a single maintenance on the network. Conversely, if the F value is small (overload is not severe or there is no overload), then (1-αF) is close to 1, and D new Approaching D old The clustering distance threshold adjustment range is small. The adjustment step size factor α is used to control the sensitivity of the adjustment, and its value is usually between 0 and 1. It can be set empirically or determined through pre-experiments based on the actual network characteristics and optimization objectives.
[0058] Subsequently, the adjusted clustering distance threshold D was used. new The steps for clustering geographically nearest nodes and subsequent steps are re-executed. This step ensures the effectiveness of the iterative process. At the clustering distance threshold D... new After being intelligently adjusted, the entire maintenance plan scheduling process (including clustering, batch partitioning, resource scheduling, simulation evaluation, etc.) will start from scratch, but this time based on a more optimized clustering parameter. In this way, the system can adaptively adjust its strategy for identifying and partitioning high-risk areas based on the experience of previous failures, and thus hopefully find a maintenance plan that meets stability requirements in subsequent iterations.
[0059] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A method for scheduling N-1 maintenance plans for a communication network based on a dynamic risk map, characterized in that, Includes the following steps: Acquire risk assessment data, including the failure probability and impact range of nodes and lines in the communication network, and generate a risk level distribution map; Based on the risk level distribution map, nodes that are geographically close are clustered to identify at least one high-risk maintenance priority cluster. Analyze the connection relationships between adjacent nodes in the high-risk maintenance priority cluster. When the connection relationship meets the set connection strength condition, the corresponding nodes are assigned to the same maintenance batch, forming a preliminary batch division set. Based on the initial batch partitioning set, the resource scheduling path between each batch is calculated using the shortest path algorithm, and the resource allocation scheme is determined with the goal of minimizing scheduling cost. Based on the resource allocation plan, the impact of each batch of maintenance on network traffic capacity is simulated. If the simulation results show that a certain batch of maintenance will cause the network load in the surrounding area to exceed the capacity, the execution order of the batches will be adjusted to distribute the high load period and generate an optimized maintenance execution order. Based on the optimized maintenance execution sequence, its impact on the stability of the entire communication network is assessed. If the assessment results do not meet the preset stability indicators, clustering, resource allocation, and simulation optimization are re-executed until a final maintenance plan that meets the stability indicators is obtained.
2. The method for scheduling N-1 maintenance plans for communication networks based on dynamic risk maps according to claim 1, characterized in that: Risk assessment data includes: statistics on the number of failures of each node and line during historical operating cycles, and calculation of the failure frequency per unit time as the failure probability; and statistics on the number and distribution range of other nodes and lines affected by a failure that cause service interruption, as the scope of impact.
3. The method for scheduling N-1 maintenance plans for communication networks based on dynamic risk maps according to claim 1, characterized in that: Generate a risk level distribution map, including: Obtain the maximum failure probability value P among all nodes and lines in the entire network. max and the minimum failure probability value P min And the maximum range of influence value R max and the minimum influence range value R min ; For any node or line, according to the formula: P norm =(P-P min ) / (P max -P min ); Calculate its normalized failure probability value P norm , where P is the original fault probability value of the node or line; According to the formula: R norm =(R-R min ) / (R max -R min ); Calculate its normalized influence range value R norm , where R is the original influence range value of the node or line; P norm and R norm Multiply by the product to obtain the risk coefficient S of the node or line; Using the network topology map as the base map, each node and line is colored with a grayscale value or RGB color value corresponding to the risk coefficient S of the node or line, forming a spatial distribution map in which the degree of risk is represented by the color intensity.
4. The method for scheduling N-1 maintenance plans for communication networks based on dynamic risk maps according to claim 3, characterized in that: Cluster geographically proximate nodes to identify at least one high-risk maintenance priority cluster, including: Set the clustering distance threshold D and the risk coefficient threshold S. th ; Starting with any node not belonging to any cluster, search for other nodes whose geographical distance is less than the clustering distance threshold D. If the risk coefficient S of the searched node is greater than the risk coefficient threshold S... th If so, the node and the starting node will be grouped into the same temporary cluster; Repeat the above search process until there are no more risk coefficients greater than the risk coefficient threshold S within the range of the temporary cluster's perimeter that are less than the clustering distance threshold D. th The unassigned nodes will be identified as a high-risk, priority maintenance cluster for the temporary cluster. Iterate through all nodes in the network, repeating the above steps until all risk coefficients S are greater than the risk coefficient threshold S. th All nodes were assigned to the corresponding high-risk maintenance priority cluster.
5. The method for scheduling N-1 maintenance plans for communication networks based on dynamic risk maps according to claim 1, characterized in that: Determining whether a connection meets the set connection strength conditions includes: obtaining the link type, link bandwidth, and physical distance between adjacent nodes. When the link type is the same, the link bandwidth can carry the service traffic transferred during maintenance, and the physical distance supports the round-trip scheduling of maintenance personnel in a single working period, the connection is determined to meet the set connection strength conditions.
6. The method for scheduling N-1 maintenance plans for communication networks based on dynamic risk maps according to claim 1, characterized in that: The resource scheduling paths between batches are calculated using the shortest path algorithm, including: Determine the center coordinates of each maintenance batch, and use the center position as the path node; Obtain the location coordinates of the resource storage point, as well as the road traffic distances between any two path nodes and between the resource storage point and each path node; Starting from the resource storage point, select the path node that is closest to the current node in terms of road traffic distance from the unvisited path node in turn as the next node to be visited, and mark the node as visited; Repeat the above steps until all path nodes have been visited once. Returning to the resource storage point from the last visited path node, the distances of all the road segments traversed in the above process are added together to obtain the total scheduling distance, and the path corresponding to the total scheduling distance is determined as the resource scheduling path.
7. The method for scheduling N-1 maintenance plans for communication networks based on dynamic risk maps according to claim 1, characterized in that: The simulation of the impact of each batch of maintenance on network traffic capacity includes: removing the nodes and lines of the batch to be maintained from the current network topology, simulating the redistribution of traffic on each link in the remaining network, calculating the ratio of the simulated traffic value of each link to the rated capacity of the link as the load rate, and determining that the batch maintenance will cause the network load in the surrounding area to exceed the carrying capacity if the load rate of any link exceeds the preset overload threshold.
8. The method for scheduling N-1 maintenance plans for communication networks based on dynamic risk maps according to claim 7, characterized in that: Adjusting the execution order of batches to distribute high-load periods includes: obtaining multiple maintenance batches that cause overload on the same link, arranging these batches at intervals on the timeline, inserting at least one other maintenance batch that will not cause link overload between two adjacent maintenance batches that cause link overload, and ensuring that at most one adjacent batch is under maintenance during any maintenance period.
9. The method for scheduling N-1 maintenance plans for communication networks based on dynamic risk maps according to claim 1, characterized in that: Assess its impact on the stability of the entire communication network, including: Following the optimized maintenance execution sequence, the load conditions of each link during each batch of maintenance are simulated sequentially; If, during any maintenance batch, the load on any link exceeds its rated capacity, it is determined that the maintenance execution sequence will cause network overload and fail to meet stability requirements. If the load on all links does not exceed their rated capacity during the execution of all maintenance batches, then the maintenance execution sequence is deemed to meet the stability requirements.
10. The method for scheduling N-1 maintenance plans for communication networks based on dynamic risk maps according to claim 1, characterized in that: Re-execute clustering, resource allocation, and simulation optimization, including: In the maintenance execution order generated in the previous iteration, obtain the number of link overloads that caused overload in the surrounding area when each maintenance batch was executed, and use the number of overloads as the feedback coefficient F of the area where the batch is located. According to the formula: D new =D old ×(1-αF) adjusts the clustering distance threshold, where D old The clustering distance threshold used in the previous iteration is α, which is the adjustment step size factor. With the adjusted clustering distance threshold D new Repeat the steps of clustering geographically proximate nodes and subsequent steps.