Method, system, medium and electronic device for automatic recovery of network equipment after disconnection
By acquiring network topology maps, the system automatically identifies and isolates disconnected nodes, solving the problem of slow response caused by device disconnection in large-scale networks. This enables rapid fault recovery and data transmission, improving network management efficiency and reliability.
Patent Information
- Application Number
- CN202411498697.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-25
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2044-10-25
AI Technical Summary
In the event of a large-scale network or multiple devices losing connection simultaneously, existing technologies rely on manual intervention, resulting in slow response times and low efficiency. This makes it difficult to quickly locate problems and take recovery measures, affecting the stability and reliability of network services.
By acquiring the network topology map, the system automatically identifies lost nodes and applies scheduling commands to isolate them and transmit the target data to backup nodes, thus achieving automated recovery.
It improves fault response speed, reduces human error, ensures data integrity and availability, lowers maintenance costs, and adapts to situations where multiple devices are out of service in complex network environments.
Smart Images

Figure CN119520233B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network management technology, specifically to an automatic recovery method, system, medium, and electronic device after a network device loses connection. Background Technology
[0002] As networks continue to expand in scale and complexity, the stability and reliability of network devices have become critical issues in network management. In large network environments, device disconnection is a common failure scenario, potentially leading to network service interruptions, data transmission delays or losses. Therefore, quickly identifying and restoring disconnected devices is essential for maintaining normal network operation.
[0003] Currently, common methods for handling network device outages primarily rely on manual intervention. After network administrators discover a device is out of service through monitoring systems, they need to manually analyze the cause of the failure, formulate recovery strategies, and execute corresponding operations remotely or on-site. This method may be adequate for handling single device outages, but it is often slow and inefficient when dealing with large-scale networks or multiple devices failing simultaneously. Summary of the Invention
[0004] This application provides an automatic recovery method, system, medium, and electronic device after a network device loses connection. In the case of large-scale network or multiple devices losing connection at the same time, it can quickly locate the problem and take corresponding measures, thereby improving the fault response speed.
[0005] In a first aspect, this application provides an automatic recovery method for a network device after it loses connection, comprising:
[0006] Obtain the network topology map corresponding to the target area, wherein the nodes of the network topology map correspond to the devices in the target area, and the edges of the network topology map correspond to the connection relationships of the devices;
[0007] If it is determined that there are disconnected nodes in the network topology graph, then a scheduling instruction is applied to the disconnected nodes and the disconnected nodes are repaired.
[0008] The scheduling instruction is used to isolate the target data currently being transmitted by the disconnected node, and to schedule the target node to transmit the target data to a backup node in the network topology.
[0009] A second aspect of this application provides an automatic recovery system after a network device loses connection, comprising:
[0010] A network topology map acquisition module is used to acquire a network topology map corresponding to a target area, wherein the nodes of the network topology map correspond to the devices in the target area, and the edges of the network topology map correspond to the connection relationships of the devices;
[0011] The scheduling instruction application module is used to apply scheduling instructions to the lost node and repair the lost node if it is determined that there is a lost node in the network topology graph; wherein, the scheduling instructions are used to isolate the target data currently being transmitted by the lost node, and to schedule the target node to transmit the target data to a backup node in the network topology.
[0012] A third aspect of this application provides a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the method steps described above.
[0013] A fourth aspect of this application provides an electronic device, comprising: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the above-described method steps.
[0014] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:
[0015] By acquiring the network topology map corresponding to the target area, a comprehensive understanding of the connectivity and operational status of network devices can be obtained. When a lost node is detected in the topology map, the system automatically applies scheduling commands and performs repairs without manual intervention. This method significantly improves fault response speed, especially when dealing with large-scale networks or multiple devices losing connection simultaneously, enabling rapid problem location and appropriate measures. The application of scheduling commands achieves automatic isolation of lost nodes and intelligent scheduling of data transmission, effectively preventing fault propagation and ensuring the continuity of network services. By transmitting the target data to a backup node, this method ensures data integrity and availability, minimizing the risk of data loss due to device loss.
[0016] Compared to existing methods that rely on manual intervention, the automatic recovery mechanism of this invention significantly improves network management efficiency, reduces human error, and lowers maintenance costs. Furthermore, its automated nature enables it to quickly respond to situations where multiple devices simultaneously lose connection in complex network environments, overcoming the slow response and low efficiency of traditional methods in such scenarios. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating an automatic recovery method for a network device after it loses connection, as provided in an embodiment of this application.
[0018] Figure 2 This is a schematic diagram of the structure of an automatic recovery system for a network device after it loses connection, provided in an embodiment of this application.
[0019] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0020] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0021] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.
[0022] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0023] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an automatic recovery method for network devices after connection loss, provided by an embodiment of the present invention. This method can be implemented using a computer program, a microcontroller, or run on an automatic recovery system for network devices based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application. Specifically, the method may include the following steps:
[0024] Step 101: Obtain the network topology map corresponding to the target area, where the nodes of the network topology map correspond to the devices in the target area, and the edges of the network topology map correspond to the connection relationships of the devices.
[0025] The target area refers to a specific network region requiring automatic recovery management of network device connectivity issues. It can be understood as a logically or physically independent network subsystem containing a set of interconnected network devices and connections. This area may be geographically defined, such as the network of a city or region; or it may be functionally or service-type defined, such as a company's core business network or a dedicated network for a specific service. In this embodiment, the target area may include various types of network devices, such as routers, switches, servers, etc., and various communication links connecting these devices. It also includes a network management center responsible for managing and controlling these devices. These devices and links together constitute a relatively complete network subsystem capable of independently performing specific network functions or supporting specific business requirements.
[0026] Correspondingly, a network topology diagram refers to a mathematical model and graphical representation used to describe the structure and relationships of network devices and their interconnections within a target area. It can be understood as a complex graph composed of nodes and edges, where nodes represent various devices in the network, and edges represent the physical or logical connections between these devices. In this embodiment, the network topology diagram is not merely a static network structure diagram, but a dynamic, information-rich data model. It contains detailed attribute information for each device, such as device type, function, importance, and current status. For example, network devices may be divided into a management control plane and a service forwarding plane, with the network management center responsible for transmitting control commands and service data to these devices. Furthermore, the edges in the topology diagram also contain connection attributes, such as bandwidth, latency, and connection type.
[0027] Based on the above embodiments, as an optional embodiment, step 101, which involves obtaining the network topology map corresponding to the target area, may further include the following steps:
[0028] Step 201: Obtain the device types of the devices in the target area and the connection relationships of each device. The device types include the network management center and network devices. The network devices include the management control plane and the service forwarding plane. The network management center transmits control commands to the network devices through the management control plane and transmits service data to the network devices through the service forwarding plane. The connection relationships of the devices include: at least one network device establishes a communication connection with the network management center, and at least one network device exchanges information through the network center.
[0029] Device type refers to the functional classification of various devices in the network. It can be understood as a system for classifying and identifying network devices to distinguish the specific responsibilities and functions that different devices undertake in the network. This classification not only includes the hardware attributes of the devices, but more importantly, it reflects the functional positioning and data processing capabilities of the devices in the overall network architecture.
[0030] Furthermore, the device types in this application embodiment are mainly divided into two categories: network management center and network devices. The network management center is the core of control and management of the entire network, responsible for the overall network monitoring, configuration, and decision-making. Network devices are various devices that constitute the network infrastructure, such as routers, switches, and firewalls. More specifically, network devices can be divided into two functional layers: the management control plane and the service forwarding plane. The management control plane is mainly responsible for receiving and executing control commands from the network management center, handling tasks such as network configuration and route calculation; while the service forwarding plane is responsible for the actual packet forwarding and processing, executing specific network services.
[0031] Defining device types helps the system adopt targeted strategies during fault diagnosis and recovery. When a device is detected to be out of service, the system can quickly assess the scope and severity of the impact based on the type of device. For example, if the network management center is out of service, the backup system may need to be activated immediately; while if a problem occurs in the service forwarding plane of a network device, the system may attempt to reroute data flows or activate a backup device.
[0032] Furthermore, device type information is used to optimize network resource allocation and performance management. The system can dynamically adjust network configuration based on the characteristics and load of different types of devices to achieve more balanced resource utilization and more efficient data processing.
[0033] In this context, connectivity refers to the physical and logical relationships between devices in a network, reflecting the data flow path and dependencies between devices. This connectivity includes not only direct physical connections between devices but also logical connections established through network protocols and virtualization technologies.
[0034] Furthermore, the connectivity in this embodiment is mainly manifested in two aspects. One is the connection between network devices and the network management center; the other is the information exchange between network devices through the network management center. The network management center transmits control commands to network devices through the management control plane, and simultaneously transmits service data to network devices through the service forwarding plane. This dual connectivity mechanism ensures effective network management and efficient data transmission.
[0035] Defining connectivity relationships is primarily used to construct an accurate network topology diagram. By understanding the connection methods between devices in detail, the system can generate a complete network structure model, which is crucial for understanding the overall network architecture and operating mechanisms. For example, the system can clearly identify which devices are directly connected and which devices need to communicate through intermediate nodes.
[0036] Secondly, connectivity information is crucial for fault diagnosis and recovery. When a device becomes unreachable, the system can quickly trace other potentially affected devices and services based on connectivity relationships. For example, if a critical network node becomes unreachable, the system can immediately assess which downstream devices might be affected, thereby developing a more targeted recovery strategy.
[0037] Step 202: Determine the nodes and edges according to the device type and connection relationship of each device.
[0038] In this context, a node refers to a basic unit in the network topology diagram, representing various devices or functional entities within the network. In this embodiment, a node can be understood as a specific device in the network, such as a router, switch, or server, as well as more abstract functional units, such as a network management center, management control plane, and service forwarding plane. Each node carries rich attribute information, including but not limited to device type, functional role, processing capacity, and current status. These attributes enable the system to accurately assess the importance and potential impact of each node in the network.
[0039] Correspondingly, an edge refers to a line connecting nodes in a network topology diagram, representing various connection relationships between nodes. In this embodiment, an edge can be understood as a physical connection between devices, such as a network cable or fiber optic cable; or a logical connection, such as a logical link in a virtual private network or software-defined network. Each edge also carries important attribute information, such as connection type, bandwidth capacity, current load, and latency characteristics. These attributes enable the system to accurately simulate the data flow and information exchange process in the network.
[0040] A precise network topology graph can be constructed by defining nodes and edges. By abstracting complex network structures into combinations of nodes and edges, the system can create a mathematically rigorous and intuitive network representation. This representation makes complex network structures easier to understand and analyze.
[0041] For example, when a node goes offline or an edge breaks, the system can immediately assess the severity of the failure and the potential scope of its impact by analyzing the properties of the relevant nodes and edges. For instance, the loss of a highly connected node may have a greater impact on the entire network, thus requiring more urgent handling.
[0042] Secondly, based on the analysis of node and edge attributes, the system can quickly determine the optimal recovery path. For example, among multiple possible recovery paths, the system will prioritize those paths composed of high-bandwidth, low-latency edges to ensure the efficiency of the recovery process.
[0043] Furthermore, the definitions of nodes and edges are also used for network optimization and resource management. By continuously monitoring the load status of nodes and the traffic of edges, the system can dynamically adjust the network configuration, achieve load balancing, and improve overall network performance.
[0044] Finally, when new devices are added or the network structure changes, the network topology can be easily updated simply by adding new nodes or edges and defining the corresponding attributes.
[0045] Step 203: Generate a network topology graph based on the nodes and edges.
[0046] Based on the above embodiments, as an optional embodiment, this application also introduces the concept of node hierarchy. By classifying nodes into critical nodes and ordinary nodes, a more refined strategy foundation is provided for network management and fault recovery. Specifically, step 202: determining nodes and edges according to the device type and connection relationship of each device may further include the following steps:
[0047] Step 301: Based on the equipment type and connection relationship of each device, determine the critical nodes and ordinary nodes. The critical nodes are more important than the ordinary nodes.
[0048] In this context, a critical node refers to a device or functional entity that plays a core role in the network and has a significant impact on the overall network functionality and performance. In this embodiment, a critical node can be understood as a network device that meets one or more of the following conditions: it undertakes critical network functions, such as a network management center or core router; it has high connectivity and is directly connected to a large number of other devices; it is located on a critical path in the network and is responsible for forwarding or processing a large amount of data; it runs important services or business operations and has a significant impact on overall network performance. These nodes typically have a high overall score, exceeding a preset importance threshold.
[0049] Correspondingly, ordinary nodes refer to devices or functional entities that play a secondary or auxiliary role in the network compared to critical nodes. In the embodiments of this application, ordinary nodes can be understood as devices that, while necessary for the normal operation of the network, have a relatively small impact on the overall network, such as terminal access switches and ordinary workstations. The overall score of these nodes is usually lower than a preset importance threshold.
[0050] Furthermore, identifying critical nodes and ordinary nodes offers the following advantages for managing network devices:
[0051] (1) Optimize fault detection and recovery strategies. The system will perform more frequent and in-depth monitoring of critical nodes and prioritize resource allocation for recovery when a fault occurs. This ensures that the core functions of the network can be maintained and restored as quickly as possible in the event of a fault.
[0052] (2) Achieve intelligent resource allocation. The system can allocate more network resources to key nodes, such as higher bandwidth priority, stricter security measures, and more frequent data backups, thereby improving the overall network performance and reliability.
[0053] (3) Used for accurate fault impact assessment. When a node has a problem, the system can quickly assess the possible scope and severity of the impact based on whether it is a critical node or a normal node, and thus formulate corresponding response strategies.
[0054] (4) It provides important guidance for network optimization and upgrades. When carrying out network transformation or expansion, the system can prioritize enhancing the performance of key nodes or increasing their redundancy, thereby more effectively improving the quality of the entire network.
[0055] Based on the above embodiments, in the automatic recovery process after network device disconnection, accurately assessing the importance of each device is crucial for formulating an effective recovery strategy. Therefore, this application introduces a device importance calculation method based on a first preset formula. Specifically, step 301 may further include the following steps:
[0056] Step 401: Based on the first preset formula, calculate the importance of each device by considering its device type and connection relationship.
[0057] The first preset formula is:
[0058]
[0059] In the formula, R i T represents the importance of device i. i n represents the device type coefficient of device i. i m represents the number of devices directly connected to device i. i K represents the number of devices indirectly connected to device i. ij K represents the connection type coefficient where device i is directly connected to device j. ik The connection type coefficient, C, for the indirect connection between device i and device k. j I represents the number of devices directly connected to device j. k λ represents the number of devices indirectly connected to device k, and μ represents the corresponding weight coefficients. The number of hops for the indirectly connected devices is less than or equal to 3.
[0060] Specifically, in real-world networks, the functions and roles of devices vary greatly. To reflect the inherent importance of different types of devices within the network, a device type coefficient is needed to differentiate their fundamental importance. For example, the network management center, as the control core of the entire network, undertakes critical tasks such as network configuration, monitoring, and management; its device type coefficient might be set to 2 or higher. In contrast, ordinary network devices such as switches or routers, while also important, play a relatively minor role in overall network management, and therefore their device type coefficient might be set to 1.
[0061] Furthermore, the main body of the formula considers the connection relationships between devices, divided into two main terms: direct connection impact and indirect connection impact. Directly connected devices typically have closer functional relationships and more direct impacts, while indirect connections represent the potential influence of devices over a wider range.
[0062] In the section on the impact of direct connections, the formula considers all other devices directly connected to the device. For each directly connected device, the formula calculates the product of the importance of its connection type and the number of connections that device has, and then sums these products for all directly connected devices. The importance of a connection type is reflected by a connection type coefficient; for example, a connection carrying critical control commands might be assigned a higher connection type coefficient (e.g., 1.5), while a connection carrying ordinary data transmission might have a lower coefficient (e.g., 1.0). This distinction allows the system to more accurately assess the contribution of each connection to the device's importance.
[0063] Correspondingly, the structure of the indirect connection impact section is similar to that of the direct connection impact section, but it considers the case of indirect connections through other devices. For each indirectly connected device, the formula also calculates the product of the importance of its connection type and the number of connections that device has, and then sums these products for all indirect connections. The introduction of this section allows the formula to assess the impact of a device across a wider network.
[0064] Furthermore, the setting that the number of hops for indirect connections is less than or equal to 3 is based on the practical needs of network management and performance optimization. For large network topologies, considering all possible indirect connections would not only lead to a sharp increase in computational complexity, but also potentially cause the number of indirect connections for most nodes to become similar, making it difficult to effectively distinguish the importance of nodes. Therefore, this embodiment adopts a more refined and practical method, namely, considering only indirect connections with hop counts of 1, 2, and 3.
[0065] This approach first focuses on directly connected devices (hop count 1), which typically have the most direct impact on node performance and functionality. Next, it considers devices with one-hop indirect connections (hop count 2), which, although not directly connected, are still within the node's neighborhood and significantly influence local network characteristics. Finally, it extends the scope to devices with two-hop indirect connections (hop count 3), which, while relatively distant, remain within a reasonable range of influence and reflect the node's position and importance within the larger network structure.
[0066] By employing the step-by-step approach described above, the system can obtain sufficiently detailed network topology information while maintaining computational efficiency. This balance enables the system to quickly assess the importance of nodes, identify the potential impact range of failures, and perform effective load balancing or backup node selection when necessary. For example, when selecting a backup for a critical node, the system prioritizes devices directly connected to it or separated by only one node, as these devices can typically take over the functionality of the lost node more quickly and effectively. Simultaneously, considering devices with a hop count of 3 provides the system with a broader perspective, helping to identify potential alternative paths or resources in large-scale networks.
[0067] To balance the contributions of direct and indirect connections to device importance, the formula introduces direct connection weights and indirect connection weights. Typically, the direct connection weight is greater than the indirect connection weight; for example, the direct connection weight might be set to 1.5, while the indirect connection weight is set to 1.0. This setting reflects a common understanding that direct connections generally have a greater impact on device importance than indirect connections. However, these weights can be adjusted based on the specific characteristics of the network and management needs. For example, in a highly distributed network, it might be necessary to increase the indirect connection weight to emphasize the importance of indirect connections.
[0068] Furthermore, in network management, different types of connections have varying degrees of impact on network operation and stability. The first pre-defined formula quantifies this difference by introducing a connection type coefficient, particularly emphasizing two key connection functions between the network management center and network devices: control command transmission and service data transmission. This distinction is crucial for accurately assessing device importance and optimizing fault recovery strategies.
[0069] Control command transmission connections are primarily used by the network management center to send management commands, configuration updates, and monitoring requests to network devices. These connections directly impact the overall management and control capabilities of the network. For example, through such connections, the network management center can adjust routing policies for network devices in real time, update security policies, or perform critical operations such as software upgrades. Because these operations directly affect network stability and security, control command transmission connections are typically given high importance.
[0070] In contrast, business data transmission connections are primarily used to carry regular data traffic within the network, such as users browsing the web, transferring files, or streaming video. While these connections are equally important for the daily operation of the network, in most cases, brief interruptions or delays do not immediately threaten the overall operational status of the network.
[0071] Based on the above, the system can assign a higher connection type coefficient value, such as 1.5, to control command transmission connections. This higher value reflects the crucial role of control connections in maintaining normal network operation. Correspondingly, service data transmission connections may be assigned a relatively lower value, such as 1.0. This differentiated assignment reflects that in network management, maintaining effective control over the network is more urgent and important than maintaining specific service data flows.
[0072] The above technical solution provides practical application value for network operation and maintenance by calculating the device importance value based on the first preset formula.
[0073] Regarding fault recovery prioritization, the system sorts out disconnected devices based on calculated importance values. When multiple devices fail simultaneously, the system prioritizes allocating resources to devices with higher importance. For example, suppose three devices A, B, and C fail simultaneously, with importance values of 8.5, 6.2, and 7.8, respectively. The system will first allocate resources to restore device A, followed by C and B. This recovery order based on quantitative indicators ensures that core network functions and critical services can be restored as quickly as possible, thereby minimizing the impact of network outages on overall business operations.
[0074] Secondly, regarding resource allocation strategies, the network management system can dynamically adjust resource allocation based on the importance of devices. For devices with higher importance, the system will allocate more network bandwidth, more frequent monitoring and checks, and higher-level security protection measures. For example, if the importance value of a router exceeds a preset threshold, the system may automatically increase its data backup frequency from once a day to once an hour and allocate a dedicated backup power supply to ensure its continuous and stable operation.
[0075] In terms of network optimization, importance values provide clear optimization directions. The system can periodically analyze the distribution of importance values for all devices in the network, identifying areas where importance is either too concentrated or too dispersed. For example, if multiple devices in a certain area have very high importance values (e.g., all exceeding 8.5), this may indicate a single point of failure risk in that area. The system can then add redundant devices or redesign the network topology in that area to mitigate the risk and improve the overall reliability of the network.
[0076] Step 402: Identify devices with an importance level greater than or equal to the threshold as critical nodes, and identify devices with an importance level less than the threshold as ordinary nodes.
[0077] Step 302: Determine the backup nodes based on the critical nodes and ordinary nodes.
[0078] In this context, a backup node refers to a device or functional entity in the network topology that can take over the functions of a critical node when necessary. In this embodiment, a backup node can be understood as an ordinary node that has been evaluated and selected by the system, possesses sufficient processing power and network connectivity, and can quickly take over the functions of a critical node when it fails.
[0079] Specifically, backup nodes serve as a core component of network redundancy. When critical nodes fail or become unavailable, they can quickly take over their functions, maintaining basic network operation and the continuity of critical services. Secondly, backup nodes also play a load balancing role in daily network operations, distributing some network traffic when needed to prevent overload of critical nodes.
[0080] In one feasible implementation, the selection and configuration of backup nodes is a dynamic process. The system continuously adjusts and optimizes the settings of backup nodes based on changes in network topology, evolution of business needs, and real-time performance of each node. This dynamism enables the network to maintain an optimal redundancy strategy at all times, ensuring high reliability while avoiding excessive waste of resources.
[0081] Step 303: Treat the key node, the ordinary node, and the backup node as nodes, and the connection relationship as edges.
[0082] Based on the above embodiments, as an optional embodiment, step 302: determining the backup node based on the critical node and the ordinary node may further include the following steps:
[0083] Step 501: Obtain the propagation path of the key nodes.
[0084] In this context, the propagation path refers to the main channels and directions through which data, control commands, and network status information flow and spread within the network. In this embodiment, the propagation path can be understood as a set of routes and connections starting from a key node, passing through a series of intermediate nodes and links, and ultimately reaching the network edge or a specific target node. These paths include not only physical network connections but also logical data flow and control information transmission processes.
[0085] Specifically, the system first extracts all nodes directly connected to key nodes from the network topology, forming the first layer of the propagation path. Then, the system expands outward layer by layer, analyzing the connections between these directly connected nodes and other nodes, until it covers the network edge or reaches a preset hop count limit. In this process, the system considers not only physical connections but also logical connections, such as logical links in Virtual Private Networks (VPNs) or Software-Defined Networking (SDN).
[0086] Simultaneously, the system also analyzes the main data flows originating from key nodes by combining real-time network traffic data. This includes the transmission paths of control commands, the flow of business data, and the interaction paths of various network protocols. For example, for a core router, the system pays particular attention to the propagation path of its routing table updates, as this directly affects the routing decisions of the entire network.
[0087] Furthermore, the system considers network redundancy and load balancing mechanisms. For example, if multiple parallel data transmission paths exist, the system analyzes their usage frequency and load to determine which are the primary propagation paths and which are backup paths. This analysis helps identify critical links and potential bottlenecks in the network.
[0088] Once the propagation path of a critical node is obtained, the system can more accurately assess the potential impact of a critical node failure. This provides an important basis for subsequent selection of backup nodes. For example, if it is found that the propagation path of a critical node covers a large area of the network, then when selecting its backup node, candidate nodes that can cover the same area need to be considered.
[0089] Step 502: Select ordinary nodes that can replace key nodes in the propagation path as target nodes.
[0090] Specifically, the system first analyzes the characteristics of each propagation path based on the previously acquired information about the propagation paths of key nodes, including path length, bandwidth capacity, and latency characteristics. Then, the system searches the network topology for ordinary nodes that can cover the same or similar propagation paths. This process involves complex path comparison and matching algorithms.
[0091] In one feasible implementation, the system can use a shortest path algorithm or a multi-constraint path algorithm to find ordinary nodes that can replace critical nodes at the lowest cost.
[0092] For example, the system can apply shortest path algorithms, such as Dijkstra's algorithm or the Bellman-Ford algorithm, to calculate the shortest paths from the critical node to all other nodes in the network. These paths constitute the propagation path set of the critical node. The system records the characteristics of these paths, including path length, total delay, minimum bandwidth, etc.
[0093] The system then iterates through all ordinary nodes in the network, repeating the above process for each node to calculate the shortest path from that node to all other nodes in the network. The system compares these paths with the propagation path set of key nodes, evaluating their similarity. The similarity calculation may consider factors such as the overlap of nodes covered by the path, differences in path length, and the degree of matching of key performance indicators.
[0094] In some cases, a single shortest path may not adequately reflect the complexity of the network. In such situations, the system will switch to using multi-constraint path algorithms. These algorithms, such as SAMCRA (Self-Adaptive Multiple ConstraintsRouting Algorithm) or TAMCRA (Tunable Accuracy Multiple ConstraintsRouting Algorithm), can simultaneously consider multiple network constraints, such as latency, bandwidth, hop count, and reliability. Using these algorithms, the system can identify ordinary nodes that are close to critical nodes in multiple performance metrics.
[0095] Specifically, multi-constraint path algorithms define a multi-dimensional cost function. For example, a three-dimensional vector (latency, bandwidth, reliability) can be used to represent the path cost. The algorithm searches for non-dominated solutions in this multi-dimensional space—paths that are superior to other solutions in at least one dimension. This allows the algorithm to identify ordinary nodes that best approximate the performance characteristics of the critical node after considering multiple factors.
[0096] Furthermore, the system considers not only physical topology but also logical topology. For example, in a software-defined networking (SDN) environment, the system analyzes the logical connections of the control plane to ensure that the selected target node can not only replace critical nodes in data forwarding but also play a corresponding role in network control and management. Simultaneously, the system also considers factors such as the node's hardware capabilities and software version compatibility to ensure that the target node meets the technical specifications required to replace critical nodes.
[0097] In addition, the system will evaluate the impact of each potential target node on network performance. This includes analyzing the potential changes to network latency, throughput, load balancing, etc., if the node is used as a backup. The system will prioritize ordinary nodes that can best maintain the original network performance characteristics after replacing critical nodes as target nodes.
[0098] In practice, the system may select multiple target nodes for each critical node, forming a candidate pool. These target nodes are ranked according to their matching degree of replacement capability, providing flexibility for the final determination of subsequent backup nodes. For example, if the most matching target node cannot serve as a backup node for some reason, the system can quickly select the next target node from the candidate pool.
[0099] Step 503: Obtain the redundancy of the target node.
[0100] Specifically, after identifying the target nodes that can replace the critical node propagation path, the system can further determine the redundancy of these target nodes. Redundancy reflects the remaining resources of a node in terms of processing power, storage space, network bandwidth, etc., and is a key indicator for selecting suitable backup nodes.
[0101] Specifically, the system performs a comprehensive resource assessment for each target node. This assessment process involves multiple dimensions, including but not limited to CPU utilization, memory usage, storage space utilization, and network interface bandwidth utilization. The system collects this data in real time and analyzes it in conjunction with historical data to obtain more accurate and stable redundancy assessment results. For example, for CPU utilization, the system considers not only the current instantaneous value but also the average and peak values over a period of time to ensure that the selected backup nodes can operate stably under various load conditions.
[0102] During the evaluation process, the system assigns weights to each resource metric and then calculates a comprehensive redundancy score through weighted averages. This score reflects the overall backup capability of the node. For example, for critical nodes that need to process large amounts of data, the system may assign higher weights to CPU and memory utilization; while for nodes mainly responsible for data forwarding, network bandwidth utilization may be given higher weights.
[0103] In addition, the system also considers the location and connectivity characteristics of the target node. For example, a node with high redundancy but a poor network location may not be as suitable as a backup as a node with slightly lower redundancy but a better location. Therefore, the system combines the node's topological location information with its resource redundancy to form a more comprehensive evaluation metric.
[0104] It's important to note that the process of obtaining target node redundancy is dynamic and continuous. The system updates this data periodically to adapt to changes in network load and fluctuations in device status. This dynamic evaluation mechanism ensures that when a backup node needs to be selected, the system can always make a decision based on the latest and most accurate information.
[0105] Step 504: If there are target nodes that are integrated into a target node with a redundancy greater than the threshold, then the target node with a redundancy greater than the threshold is identified as a backup node.
[0106] In this context, target nodes with redundancy exceeding a threshold refer to those nodes that have sufficient margin in resource utilization and performance reserves, enabling them to handle additional network load without affecting their normal functions. In this embodiment, target nodes with redundancy exceeding a threshold can be understood as network devices or functional entities whose overall resource utilization is lower than a preset security limit and whose remaining capacity exceeds a specified threshold, as assessed by the system. These include, but are not limited to, CPU utilization, memory usage, storage space, and network bandwidth utilization.
[0107] Step 505: If there is no target node with redundancy greater than the threshold, then perform load balancing on the critical node and the target node so that the redundancy of the target node is greater than the threshold.
[0108] Specifically, in order to solve the problem of uneven distribution of network resources and ensure that a suitable backup node can be found under any circumstances, during the automatic recovery process after network equipment is disconnected, when the system determines that there is no target node with redundancy greater than the threshold, load balancing is performed on critical nodes and target nodes to make the redundancy of the target node greater than the threshold.
[0109] Specifically, when the system detects that none of the potential target nodes meet the redundancy threshold requirements, it initiates a dynamic load balancing process. This process first performs a comprehensive analysis of the current network load distribution to identify which critical nodes and target nodes can have their loads redistributed. The system considers multiple factors, including but not limited to the processing capacity of each node, its current load level, network topology, and service priorities.
[0110] The system may attempt to migrate some non-critical services from high-load target nodes to less-loaded nodes, or distribute auxiliary functions of some critical nodes to other nodes to alleviate the pressure on critical nodes. During load balancing, the system strictly controls the impact on network performance. It sets a series of constraints, such as ensuring that the service quality of critical services is not affected, controlling the increase in network latency, and maintaining load balance. The system will perform multiple iterations, fine-tuning the load distribution each time, and evaluating the adjusted network status in real time until a solution that meets all constraints and ensures that the redundancy of the target node exceeds a threshold is found. In addition, load balancing also needs to consider the dynamic characteristics of the network. The system will predict short-term load change trends to ensure that the load distribution after the shift remains stable for a certain period of time.
[0111] Based on the above embodiments, the methods for determining device disconnection mainly rely on the status monitoring of the network device's management control plane and service forwarding plane. Specifically, this may include one or more of the following methods:
[0112] (1) Management and control plane detection:
[0113] Unresponsive Control Commands: The network management center issues control commands (such as service configuration, alarm query, software upgrade, etc.) to network devices through the management control plane. If the management control plane of a network device fails, these commands issued by the network management center will not receive a response. At this time, the management connection between the device and the network management center is lost, indicating that the device is disconnected.
[0114] Disconnected status: When the network management center is unable to communicate with the device through the control plane, the device is considered to be in a "disconnected" state, i.e., out of contact.
[0115] (2) Service forwarding plane detection
[0116] Service forwarding plane is normal but control plane is malfunctioning: Even if the control plane is disconnected, if the device's service forwarding plane is still functioning normally, the network management center can send specific maintenance commands (such as restart commands) through the service forwarding plane. If the service forwarding plane can accept and process these commands, it indicates that the device's forwarding plane is still normal, and the cause of the disconnection is mainly concentrated in the control plane.
[0117] Service connection validity check: By checking the service connection status between the network management center and the device, it is confirmed whether the service forwarding plane is working properly. If the service connection is normal but there is no response to control commands, the device is considered to be partially disconnected (control plane disconnection).
[0118] (3) Periodic polling mechanism
[0119] Physical port status check: Network devices periodically poll the operational status of all physical ports. If the physical ports are functioning normally, but the routing protocol status in the control plane and the OAM protocol status in the forwarding plane are both in a failed state, it can be inferred that the device's control plane or forwarding plane may have failed.
[0120] Protocol status check: By checking the routing protocol status of the control plane and the protocol status of the forwarding plane, it is possible to further confirm whether the device is in a disconnected state. If the protocol status shows as invalid multiple times consecutively, the system assumes that the device's control plane and forwarding plane may have failed simultaneously.
[0121] (4) Automated loss of connection determination
[0122] Multiple unresponsiveness determination: The system will automatically determine that the device has lost connection if it is detected as unresponsive multiple times through continuous polling and command sending mechanism.
[0123] Reboot mechanism: Once the device is confirmed to be disconnected, the system will trigger an automatic recovery mechanism to attempt to restore the device to normal operation by rebooting the device's control plane and / or forwarding plane.
[0124] (5) Graph Theory-Based Network Fault Impact Assessment
[0125] Topology Analysis: In more complex scenarios, the system uses graph theory to model the topology of the entire network. By analyzing the connections between devices, the system can assess which devices may be indirectly affected by the missing device and predict the fault propagation path, thereby quickly identifying the missing device.
[0126] Step 102: If it is determined that there are disconnected nodes in the network topology diagram, then apply scheduling instructions to the disconnected nodes and repair the disconnected nodes. The scheduling instructions are used to isolate the target data currently being transmitted by the disconnected nodes and to schedule the target data transmission to the backup nodes in the network topology.
[0127] Specifically, the system applies corresponding scheduling instructions based on the characteristics of the lost node. The main purpose of the scheduling instructions is to isolate the target data currently being transmitted by the lost node and reroute this data to a pre-determined backup node.
[0128] On the one hand, scheduling instructions primarily prevent data loss or interruption due to node disconnection, ensuring business continuity. On the other hand, isolating disconnected nodes can prevent the further spread of potential network failures, maintaining the stability of the entire network. The system first identifies all data streams currently being processed by the disconnected node, and then formulates detailed scheduling strategies based on the priority and characteristics of these data streams. For high-priority or time-sensitive data, the system will immediately redirect it to the most suitable backup node; for some non-critical data that can tolerate short-term delays, the system may temporarily cache it, waiting for better network conditions before transmission.
[0129] During the scheduling process, the system fully utilizes the previously constructed network topology map and node classification information. For example, for a lost device identified as a critical node, the system prioritizes backup nodes with redundancy exceeding a threshold as the target for data redirection. These backup nodes, due to their ample resource reserves and similar network locations to the critical nodes, can quickly take over the functions of the lost node, minimizing service interruption. Simultaneously, the system also considers the overall network load balancing to avoid overloading certain nodes during the redirection process.
[0130] The execution of scheduling instructions is a dynamic and continuous process. The system monitors the effectiveness of data redirection in real time, including performance metrics such as latency and packet loss rate on the new path. If a backup node is found to be unable to effectively handle the redirected data stream, the system quickly adjusts its strategy and selects the next most suitable backup node. This dynamic adjustment mechanism ensures that data transmission efficiency and reliability are maintained even in complex and ever-changing network environments.
[0131] Simultaneously, the system will initiate a process to repair the lost node. The specific repair strategy depends on the cause of the loss of connection and the type of node. For losses due to software failure, the system may attempt to restart the device's management control plane by sending specific maintenance commands through the service forwarding plane to trigger the restart. If the loss of connection is caused by hardware problems, the system may initiate a remote diagnostic program to collect detailed fault information and attempt automatic repair according to a preset repair procedure. In some cases, if automatic repair fails to resolve the issue, the system will generate a detailed fault report and notify the network administrator for manual intervention.
[0132] During the repair process, the system continuously evaluates the progress and effectiveness of the repair. If a certain repair strategy is found to be ineffective, it will quickly switch to an alternative. For example, if restarting the management control plane fails to resolve the issue, the system may attempt to update the device firmware or reset it to the previous stable configuration version. This flexible repair strategy ensures that the system can handle a variety of complex failure scenarios.
[0133] Based on the above embodiments, as an optional embodiment, if the lost node is not a critical node, and there are at least two backup nodes corresponding to the lost node, then the tightness of each backup node is calculated based on at least two of the path length, bandwidth, and link load between each backup node and the lost node, and the backup node with the highest tightness is taken as the target backup node.
[0134] Specifically, during the automatic recovery process of network devices after a connection failure, when the system detects that the lost node is not a critical node and multiple backup nodes are available, an efficient method is needed to select the most suitable backup node. To this end, this application introduces a backup node selection method based on density. This method comprehensively considers key factors such as path length, bandwidth, and link load, and quantifies the suitability of backup nodes by calculating density, thereby achieving the selection of the optimal backup node.
[0135] Specifically, the system first identifies all backup nodes corresponding to the lost node. These backup nodes were pre-determined in the previous network planning and possess the basic capability to take over the lost node. Then, the system collects network status data between these backup nodes and the lost node, including path length, bandwidth, and link load. After collecting this data, the system applies a preset formula for calculating the tightness of connection:
[0136]
[0137] In the formula, C represents the degree of tightness, L represents the path length between the target node and the lost node, B represents the bandwidth of the path between the target node and the lost node, U represents the link load of the path between the target node and the lost node, and w1, w2, and w3 represent the corresponding weight coefficients.
[0138] The path length between the backup node and the lost node refers to the number of hops or physical distance on the shortest path from the backup node to the lost node in the network topology. In this embodiment, it can be understood as the number of intermediate devices or the actual physical distance that a data packet needs to pass through to travel from the backup node to the location of the lost node. Path length is used to assess the network distance between the backup node and the lost node, reflecting the potential latency and complexity of data transmission. A shorter path length generally means faster data transmission speeds and lower network overhead, which is beneficial for quickly restoring network functionality.
[0139] The bandwidth between the backup node and the lost node refers to the maximum available data transmission capacity on the network path connecting these two nodes. In this embodiment, it can be understood as the minimum bandwidth value among all links on the path from the backup node to the lost node, typically expressed in bits per second. Bandwidth measures the data processing capability of the backup node when it takes over the functions of the lost node, reflecting the throughput potential of the network path. A higher bandwidth value means that the backup node can handle the data traffic of the lost node more effectively, helping to maintain network performance and quality of service.
[0140] The link load between the backup node and the lost node refers to the current resource occupancy on the network path connecting these two nodes. In this embodiment, it can be understood as the average or maximum utilization of all links on the path from the backup node to the lost node, typically expressed as a percentage. Link load is used to assess the network congestion risk that the backup node may face when taking over the functions of the lost node, reflecting the current busyness of the network path. A lower link load value indicates that the path has more available resources, making it more suitable for data migration and service switching, which is beneficial for balancing network load and improving recovery efficiency.
[0141] Furthermore, the system applies this formula to each backup node to calculate their respective density. After calculation, the system compares the density values of all backup nodes and selects the node with the highest density as the target backup node. This selection process ensures that the selected backup node is closest to the ideal state in terms of network location, transmission capacity, and current load.
[0142] This method of selecting backup nodes improves the accuracy of backup plans. By comprehensively considering multiple network factors, the system can more accurately assess the suitability of each backup node, avoiding the bias that may arise from a single metric. Secondly, this method enhances the efficiency of network recovery. Selecting the node with the highest density of backups means that data migration and service switching can be completed in the shortest possible time, thereby minimizing network downtime. Furthermore, this method helps optimize the utilization of network resources. By considering link load, the system can avoid adding extra burden to already heavily loaded paths, thus maintaining overall network balance.
[0143] It is worth noting that the effectiveness of this method largely depends on the appropriate setting of the weighting coefficients. Network administrators can adjust the values of w1, w2, and w3 according to the specific network environment and business needs. For example, in latency-sensitive application scenarios, the weight of w1 can be increased; while in networks primarily handling large data transmissions, a higher value for w2 may be required. This flexibility allows the method to adapt to various network environments and operational requirements.
[0144] Based on the above embodiments, as an optional embodiment, if the lost node is a critical node, the minimum cut algorithm is used to determine the backup node.
[0145] Specifically, during the automatic recovery process after a network device loses connection, when the system determines that the lost node is a critical node, it uses the minimum cut algorithm to determine backup nodes. The minimum cut algorithm is an algorithm used for network flow graphs. Its purpose is to find a set of edges in the network such that cutting these edges decouples the source and sink nodes. The goal of the minimum cut is to find the set of edges with the smallest total capacity, thus achieving an effective "partition." In network device recovery scenarios, the minimum cut algorithm can help determine which nodes or links can serve as backup nodes, so that network connectivity can be maintained even when the critical node fails.
[0146] In practice, the system acquires the latest network topology map and identifies the critical nodes that have lost connectivity. Let's assume critical node A is out of service. The system then transforms the network topology map into a traffic model. In this model, the out-of-service critical node A is designated as the source node, while other nodes in the network, such as B, C, and D, are considered potential destination or sink nodes. This abstraction allows the system to better analyze the distribution of network traffic and identify potential bottlenecks.
[0147] After the transformation, the system applies a minimum cut algorithm, such as the Edmonds-Karp algorithm, to this traffic model. The core idea of this algorithm is to find the minimum cut from the source node A to the rest of the network by repeatedly calculating the maximum flow. A minimum cut represents a set of edges in the network that, if these connections were severed, would have the least impact on network traffic. The nodes on either side of this set of edges become potential backup node candidates. The algorithm may require multiple iterations, with each iteration updating the network traffic state until the optimal cut set is found.
[0148] Through the minimum cut algorithm, the system may determine that nodes B and C are the most suitable devices to serve as backup nodes. This result means that B and C not only have sufficient resources and capabilities to take over the functions of A, but also maintain good connectivity in the network topology, minimizing the impact of network reconfiguration. In selecting B and C as backup nodes, the system also considers their redundancy, ensuring they have sufficient bandwidth and computing power to handle additional loads.
[0149] Once the backup node is identified, the system immediately performs a data redirection operation. All data streams originally transmitted through the lost node A will be rerouted to B and C. This process requires fine-grained traffic scheduling to ensure data transmission continuity and network load balancing. Based on the minimum cut result, the system allocates different data streams to B and C, mimicking the original functionality of A as closely as possible. Simultaneously, the system initiates the process of repairing the lost critical node A, which may include remote diagnostics, software restarts, or notifying technical personnel for on-site maintenance.
[0150] However, in some situations, B and C may not be able to fully handle the entire load of A. This could be due to their limited processing capacity or network topology limitations preventing the effective redirection of certain data flows. In such cases, the system initiates a load balancing mechanism. It analyzes other nodes in the network, such as D and E, and distributes some non-critical data flows to these nodes for processing. This dynamic load balancing strategy ensures that B and C can focus on processing the most critical business data while also making full use of other network resources.
[0151] Throughout the process, the system continuously monitors network performance metrics such as processing capacity, latency, and packet loss rate to ensure the effectiveness of the backup plan. If performance degradation is detected, the system will adjust the load balancing strategy in real time or consider enabling other backup nodes. This dynamic adjustment mechanism ensures that data transmission efficiency and reliability are maintained even in complex and ever-changing network environments.
[0152] By employing the minimum cut algorithm to determine backup nodes, the accuracy and efficiency of network recovery are improved. This allows for faster and more accurate selection of the most suitable backup node, thereby minimizing network downtime.
[0153] In summary, during the automatic recovery process of network devices after a loss of connection, the system adopted different handling strategies for critical and non-critical nodes, which stems from the differences in their importance and potential impact in the network.
[0154] For critical nodes, the system employs a minimum cut algorithm to determine backup nodes. This is because critical nodes typically play a core role in the network, and their loss of connection can significantly impact the connectivity and performance of the entire network. The minimum cut algorithm can find a set of edges with minimum capacity in the network flow graph, and cutting these edges effectively segments the network. In this embodiment, this method is used to identify the set of backup nodes best suited to take over the functions of the lost critical node. The minimum cut algorithm considers the overall network topology and traffic distribution, finding the optimal solution while maintaining network connectivity and minimizing performance impact. This method is particularly suitable for handling critical node loss because it comprehensively evaluates the network structure and finds the best backup strategy to ensure that the core functions and overall performance of the network are minimally affected.
[0155] In contrast, since the loss of non-critical nodes typically does not severely impact the entire network, more focus can be placed on local network characteristics and performance metrics. Therefore, the system employs a density-based approach to select backup nodes. This method comprehensively considers factors such as path length, bandwidth, and link load, quantifying the suitability of each potential backup node by calculating density. The density calculation allows the system to evaluate the suitability of backup nodes across multiple dimensions, taking into account both network topology and performance factors. This approach is more flexible, allowing for adjustments to the weights of various factors based on specific network environments and business requirements, thereby finding the most suitable backup node.
[0156] The two different methods described above demonstrate the system's intelligent and differentiated strategies for handling network device outages. Using the minimum cut algorithm on critical nodes ensures that the optimal solution can be found from a global perspective when dealing with situations that could significantly impact the entire network. Conversely, using a density-based method on non-critical nodes provides a more flexible and efficient solution, quickly finding locally optimal backup nodes while avoiding unnecessary computational complexity.
[0157] Reference Figure 2 This application also provides an automatic recovery system after a network device loses connection, comprising:
[0158] A network topology map acquisition module is used to acquire a network topology map corresponding to a target area, wherein the nodes of the network topology map correspond to the devices in the target area, and the edges of the network topology map correspond to the connection relationships of the devices;
[0159] The scheduling instruction application module is used to apply scheduling instructions to the lost node and repair the lost node if it is determined that there is a lost node in the network topology graph; wherein, the scheduling instructions are used to isolate the target data currently being transmitted by the lost node and schedule the transmission of the target data to a backup node in the network topology.
[0160] Based on the above embodiments, as an optional embodiment, the network topology map acquisition module is further configured to acquire the device types of the devices in the target area and the connection relationships of each device; determine nodes and edges according to the device types and connection relationships of each device; and generate a network topology map based on the nodes and edges; wherein the device types include a network management center and network devices, the network devices include a management control plane and a service forwarding plane, the network management center transmits control commands to the network devices through the management control plane, and transmits service data to the network devices through the service forwarding plane; wherein the connection relationships of the devices include: at least one network device establishing a communication connection with the network management center, and at least one network device interacting with information through the network center.
[0161] Based on the above embodiments, as an optional embodiment, the network topology acquisition module is further configured to determine key nodes and ordinary nodes according to the device type and connection relationship of each device, wherein the key nodes are more important than the ordinary nodes; determine backup nodes according to the key nodes and the ordinary nodes; and use the key nodes, the ordinary nodes and the backup nodes as nodes, and the connection relationship as edges.
[0162] Based on the above embodiments, as an optional embodiment, the network topology map acquisition module is further configured to calculate the importance of each device based on a first preset formula, according to the device type and connection relationship of each device; determine devices whose importance is greater than or equal to a threshold as key nodes, and determine devices whose importance is less than the threshold as ordinary nodes;
[0163] The first preset formula is:
[0164]
[0165] In the formula, R i T represents the importance of device i. i n represents the device type coefficient of device i. i m represents the number of devices directly connected to device i. i K represents the number of devices indirectly connected to device i. ij K represents the connection type coefficient where device i is directly connected to device j. ikThe connection type coefficient, C, for the indirect connection between device i and device k. j I represents the number of devices directly connected to device j. k λ represents the number of devices indirectly connected to device k, and μ represents the corresponding weight coefficients. The number of hops for the indirectly connected devices is less than or equal to 3.
[0166] Based on the above embodiments, as an optional embodiment, the network topology acquisition module is further configured to acquire the propagation path of the critical node; identify ordinary nodes that can replace the critical node in the propagation path as target nodes; acquire the redundancy of the target node; if there is a target node with redundancy greater than a threshold, then the target node with redundancy greater than the threshold is identified as a backup node; if there is no target node with redundancy greater than the threshold, then load balancing is performed on the critical node and the target node to make the redundancy of the target node greater than the threshold.
[0167] Based on the above embodiments, as an optional embodiment, the scheduling instruction application module is further configured to, if the lost node is not a critical node and there are at least two backup nodes corresponding to the lost node, calculate the density of each backup node based on at least two of the path length, bandwidth, and link load between each backup node and the lost node, and take the backup node with the highest density as the target backup node.
[0168] Based on the above embodiments, as an optional embodiment, the scheduling instruction application module is further configured to determine the backup node using the minimum cut algorithm if the disconnected node is a critical node.
[0169] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0170] This application also provides a computer storage medium that can store multiple instructions. These instructions are adapted to be loaded by a processor and executed as described in the above embodiments for the automatic recovery method after a network device loses connection. The specific execution process can be referred to in the detailed description of the illustrated embodiments, which will not be repeated here.
[0171] This application also discloses an electronic device. (See reference...) Figure 3 , Figure 3This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. The electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302.
[0172] The communication bus 302 is used to enable communication between these components.
[0173] The user interface 303 may include a display interface and a camera interface. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.
[0174] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0175] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by calling data stored in the memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface graphics, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also be implemented as a separate chip without being integrated into the processor 301.
[0176] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory 305 may include a non-transitory computer-readable storage medium. The memory 305 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. (Refer to...) Figure 3 The memory 305, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for an automatic recovery method after a network device loses connection.
[0177] exist Figure 3 In the illustrated electronic device 300, the user interface 303 is mainly used to provide an input interface for the user and to acquire user input data; while the processor 301 can be used to call an application stored in the memory 305 for an automatic recovery method after a network device loses connection. When executed by one or more processors 301, the electronic device 300 performs one or more of the methods described in the above embodiments. It should be noted that, for the foregoing method embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0178] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0179] In the various embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.
[0180] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0181] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0182] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0183] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practical application.
[0184] This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
Claims
1. A method for automatic recovery after a network device loses connection, characterized in that, include: Obtain the network topology map corresponding to the target area, wherein the nodes of the network topology map correspond to the devices in the target area, and the edges of the network topology map correspond to the connection relationships of the devices; If it is determined that there are disconnected nodes in the network topology graph, then a scheduling instruction is applied to the disconnected nodes and the disconnected nodes are repaired. The scheduling instruction is used to isolate the target data currently being transmitted by the disconnected node, and to schedule the target data to be transmitted to a backup node in the network topology. The step of obtaining the network topology map corresponding to the target area includes: Obtain the device types of the devices in the target area, as well as the connection relationships of each device; Nodes and edges are determined according to the device type and connection relationship of each device; Generate a network topology graph based on the nodes and edges; The equipment types include network management centers and network devices. The network devices include a management control plane and a service forwarding plane. The network management center transmits control commands to the network devices through the management control plane and transmits service data to the network devices through the service forwarding plane. The connection relationships of the devices include: at least one of the network devices establishing a communication connection with the network management center, and at least one of the network devices exchanging information through the network management center; The step of determining nodes and edges based on the device type and connection relationship of each device includes: Based on the device type and connection relationship of each device, key nodes and ordinary nodes are determined, with the key nodes being more important than the ordinary nodes; Based on the critical nodes and the ordinary nodes, the backup nodes are determined; The key nodes, ordinary nodes, and backup nodes are treated as nodes, and the connection relationships are treated as edges; The step of determining key nodes and ordinary nodes, as well as ordinary nodes and ordinary edges, based on the device type and connection relationship of each device includes: Based on the first preset formula, the importance of each device is calculated according to its device type and connection relationship; Devices whose importance is greater than or equal to the threshold are identified as critical nodes, and devices whose importance is less than the threshold are identified as ordinary nodes; The first preset formula is: ; In the formula, Indicates the importance of device i. This represents the device type coefficient of device i. This represents the number of devices directly connected to device i. This represents the number of devices indirectly connected to device i. This represents the connection type coefficient where device i is directly connected to device j. This represents the connection type coefficient between device i and device k. This indicates the number of devices directly connected to device j. This indicates the number of devices indirectly connected to device k. and These represent the corresponding weighting coefficients, where the number of hops for indirectly connected devices is less than or equal to 3.
2. The automatic recovery method for network device disconnection according to claim 1, characterized in that, The step of determining the backup node based on the critical node and the ordinary node includes: Obtain the propagation path of the key node; Ordinary nodes that can replace key nodes in the propagation path are used as target nodes; Obtain the redundancy of the target node; If there is a target node with redundancy greater than the threshold, then the target node with redundancy greater than the threshold is identified as a backup node. If there is no target node with redundancy greater than the threshold, then load balancing is performed on the critical node and the target node to make the redundancy of the target node greater than the threshold.
3. The automatic recovery method for network device disconnection according to claim 1, characterized in that, The method further includes: If the lost node is not a critical node, and there are at least two backup nodes corresponding to the lost node, then the density of each backup node is calculated based on at least two of the path length, bandwidth, and link load between each backup node and the lost node, and the backup node with the highest density is selected as the target backup node.
4. The automatic recovery method for network device disconnection according to claim 1, characterized in that, The method further includes: If the lost node is a critical node, then the minimum cut algorithm is used to determine the backup node.
5. An automatic recovery system for network devices after disconnection, used to implement the automatic recovery method for network devices after disconnection as described in any one of claims 1-4, characterized in that, include: A network topology map acquisition module is used to acquire a network topology map corresponding to a target area, wherein the nodes of the network topology map correspond to the devices in the target area, and the edges of the network topology map correspond to the connection relationships of the devices; The scheduling instruction application module is used to apply scheduling instructions to the lost node and repair the lost node if it is determined that there is a lost node in the network topology graph; wherein, the scheduling instructions are used to isolate the target data currently being transmitted by the lost node and schedule the transmission of the target data to a backup node in the network topology.
6. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-4.
7. A computer storage medium, characterized in that, The computer storage medium stores instructions that, when executed by a processor, perform the method as described in any one of claims 1-4.
Citation Information
Patent Citations
Brittleness risk analysis method and device for weapon equipment system
CN114881424A
Automatic troubleshooting and remedying network issues via connected neighbors
CN116708147A