Network recovery method and device, computer equipment, storage medium and program product
By obtaining network node parameters in real time and building an intelligent analysis mechanism, combining the mapping relationship library between fault information and recovery strategies, the rapid and accurate automatic recovery of network failures is achieved, and the problems of long detection cycles and low recovery efficiency caused by manual intervention in the existing technology are solved, and the stability and availability of network communication are improved.
Patent Information
- Application Number
- CN202510555812.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-18
AI Technical Summary
The existing network fault detection solutions rely on manual intervention, resulting in long detection cycles, low recovery efficiency, insufficient adaptability, and inability to meet the communication needs of ultra-high speed, ultra-large capacity, ultra-low latency and high reliability.
By obtaining the operating parameters of network nodes in real time, building an intelligent analysis mechanism, using the mapping relationship library of preset fault information and recovery strategies, the automated closed loop of fault location and recovery strategies is realized, and the network recovery is carried out using hierarchical progressive logic architecture and virtual address mapping technology.
Quickly and accurately identify network failures, shorten failure response time, reduce manual intervention needs, improve network communication stability and high availability, and ensure business continuity and data integrity.
Smart Images

Figure CN120343604A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mobile communication networks, and particularly to a network recovery method, apparatus, computer device, storage medium, and program product. Background Art
[0002] As the core infrastructure supporting key application scenarios such as future intelligent factories and smart cities, the network must meet multiple performance requirements such as ultra-high speed, ultra-large capacity, ultra-low latency, and high reliability. If the network fails, it may cause data loss and equipment anomalies at worst, and may even lead to production interruptions, resulting in significant losses and safety hazards. Therefore, timely detection and handling of faults are necessary prerequisites for ensuring the stable operation of critical services. However, the current fault detection solutions still rely on manual intervention, suffering from significant problems such as long detection cycles, low recovery efficiency, and insufficient adaptability. Summary of the Invention
[0003] In view of this, the present invention provides a network recovery method, apparatus, computer device, storage medium, and program product to solve the problems of long detection cycles, low recovery efficiency, and insufficient adaptability caused by the current network fault detection solutions relying on manual intervention.
[0004] In a first aspect, the present invention provides a network recovery method, including: obtaining network operation parameters of a current network node; analyzing the network operation parameters, and determining whether the current network node has a network fault based on the analysis result; if the analysis result indicates that the current network node has a network fault, determining target fault information of the network fault; determining a target recovery strategy corresponding to the target fault information according to a mapping relationship between the fault information and a network recovery strategy; and performing network recovery on the current network node according to the target recovery strategy.
[0005] The network recovery method provided by the embodiments of the present invention can quickly and accurately identify network faults by obtaining the operation parameters of network nodes in real time and constructing an intelligent analysis mechanism, overcoming the pain point of low efficiency in traditional manual troubleshooting. Based on a preset mapping relationship library between fault information and recovery strategies, an automated closed-loop for fault location and recovery strategy matching is realized, significantly shortening the fault response time. Therefore, this method adopts a hierarchical and progressive logical architecture, which not only ensures the accuracy of fault diagnosis but also enhances the system's self-healing ability through strategic recovery, greatly reducing the need for manual intervention and operation and maintenance costs, and effectively improving the communication stability and high availability of the network.
[0006] In an alternative embodiment, if the analysis result indicates that a network failure has occurred in the current network node, the target fault information of the network failure is determined, including: extracting abnormal operating parameters corresponding to the network failure from the network operating parameters; performing parameter matching between the abnormal operating parameters and the fault parameters in the fault identification conditions, and determining the target fault information corresponding to the network failure based on the fault matching result.
[0007] The network recovery method provided by the embodiments of the present invention can quickly lock the fault characteristics and accurately identify the fault type and fault point by accurately extracting abnormal operating parameters from the network operating parameters and performing intelligent parameter matching with predefined fault identification conditions, effectively avoiding the subjective misjudgment and delay problems in traditional manual troubleshooting.
[0008] In an alternative embodiment, extracting abnormal operating parameters corresponding to the network failure from the network operating parameters includes: comparing each parameter in the network operating parameters with its corresponding preset parameter threshold to determine the target operating parameter that exceeds the preset parameter threshold; obtaining the continuous duration for which the target operating parameter exceeds the preset parameter threshold; if the continuous duration exceeds the preset duration, determining the target operating parameter as the abnormal operating parameter.
[0009] The network recovery method provided by the embodiments of the present invention quickly filters potential abnormal operating parameters based on the preset parameter threshold, avoiding the computational redundancy of full-scale data analysis. By introducing the continuous duration criterion, it effectively distinguishes instantaneous fluctuations from persistent faults, solving the problem of false triggering caused by occasional peaks in the traditional threshold method. Through the dual verification mechanism, it not only ensures the quick locking of key abnormal operating parameters but also enhances the judgment reliability through time dimension verification, and can accurately filter out interference factors such as network jitter.
[0010] In an alternative embodiment, performing parameter matching between the abnormal operating parameters and the fault parameters in the fault identification conditions and determining the target fault information corresponding to the network failure based on the fault matching result includes: obtaining the parameter tree structure composed of the fault parameters; comparing the abnormal operating parameters and the fault parameters layer by layer according to the parameter tree structure to obtain the layer-by-layer comparison result of the parameters; analyzing the layer-by-layer comparison result of the parameters to determine the target fault information that matches the abnormal operating parameters.
[0011] The network recovery method provided by the embodiments of the present invention gradually narrows the fault range from global to local through the hierarchical screening characteristics of the parameter tree, avoiding the misjudgment risk caused by data redundancy or cross-interference in traditional single-parameter matching. During the layer-by-layer comparison process, each level conducts directional verification based on specific fault characteristics, effectively excluding the interference of irrelevant parameters, quickly locking the core abnormal indicators, and using the logical relevance of the tree structure to achieve multi-dimensional verification of the fault type. This hierarchical analysis mode not only significantly shortens the diagnosis time in complex fault scenarios but also enhances the interpretability and reliability of fault identification through the structured parameter matching path.
[0012] In an alternative embodiment, network recovery of the current network node is performed according to the target recovery strategy, including: determining whether the current network fault node is the master node; if the current network fault node is the master node, determining the standby node corresponding to the master node; synchronizing the service data of the master node to the standby node, and mapping the virtual address of the master node to the standby node, so that subsequent service data is directed to the standby node through the virtual address; and recovering the network of the master node at the standby node.
[0013] The network recovery method provided by the embodiments of the present invention adopts the master node status dynamic detection technology and combines the transparent migration strategy of virtual address mapping to complete seamless service switching on the premise of ensuring complete data synchronization, avoiding the risks of service interruption and data loss common in traditional master-slave switching. By decoupling the physical node topology at the virtual address layer, the fault transfer process is completely transparent to the upper-layer services, effectively ensuring the continuity of the user experience.
[0014] In an alternative embodiment, when the communication link interruption duration between the master node and the standby node exceeds the preset duration, the standby node is upgraded to the master node status; when the communication link between the master node and the upgraded standby node is restored, the target master node is re-determined from the master node and the upgraded standby node based on the preset priority; and the network element process status of the target master node is restored, and the user context information is restored to the target master node.
[0015] The network recovery method provided by the embodiments of the present invention automatically triggers the upgrade of the standby node to the master node when the master-slave node communication interruption times out, ensuring seamless migration of service traffic to the standby node and avoiding service interruption caused by link failures. After the communication is restored, the master node role is re-arbitrated based on the preset priority, effectively avoiding the risk of dual-master competition, and ensuring the integrity of service data and service consistency before and after the switch through the accurate restoration of user context information and the rapid recovery of the network element process status.
[0016] In an alternative embodiment, when the cache service of the master node is in a fault state, network recovery is performed on the current network node according to the target recovery policy, including: determining a new master node from multiple slave nodes by using a preset election method; and updating the master node addresses stored by the slave nodes other than the new master node so that the new master node takes over the cache service of the master node.
[0017] The network recovery method provided by the embodiments of the present invention significantly shortens the master-slave switchover time through a dynamic election strategy, avoiding the delay problem of traditional manual intervention. A distributed address synchronization mechanism is adopted to ensure that all slave nodes always point to the new master node in real time, eliminating the risk of data inconsistency. The election logic prioritizing performance is combined to optimize resource allocation, enhancing the overall load balancing ability and stability of the system.
[0018] In an alternative embodiment, determining whether the current network failure node is the master node includes: monitoring whether the standby node receives a heartbeat packet sent by the master node within a preset duration; if the standby node does not receive the heartbeat packet within the preset duration, determining that the current network failure node is the master node.
[0019] The network recovery method provided by the embodiments of the present invention constructs a real-time communication link monitoring system based on the lightweight heartbeat packet design. Through the dynamic setting of the preset duration, it can quickly capture the abnormal state of the master node being disconnected and effectively avoid the misjudgment risk caused by momentary network interruption. The active monitoring mode of the standby node is adopted to replace the traditional polling detection, significantly reducing the system resource consumption. At the same time, through the binary decision logic of the heartbeat packet reception status, the accuracy of the master node failure judgment is greatly improved.
[0020] In a second aspect, the present invention provides a network recovery device, including: an acquisition module, configured to acquire network operation parameters of the current network node; an analysis module, configured to analyze the network operation parameters and determine whether a network failure occurs in the current network node based on the analysis result; a first determination module, configured to determine the target failure information of the network failure if the analysis result indicates that a network failure occurs in the current network node; a second determination module, configured to determine a target recovery policy corresponding to the target failure information according to the mapping relationship between the failure information and the network recovery policy; and a recovery module, configured to perform network recovery on the current network node according to the target recovery policy.
[0021] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the network recovery method according to the first aspect or any corresponding embodiment thereof.
[0022] Fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the network recovery method of the first aspect or any corresponding embodiment thereof described above.
[0023] Fifth aspect, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the network recovery method of the first aspect or any corresponding embodiment thereof described above. Description of the Drawings
[0024] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0025] Figure 1 is a flowchart of the network recovery method according to an embodiment of the present invention;
[0026] Figure 2 is a flowchart of another network recovery method according to an embodiment of the present invention;
[0027] Figure 3 is a flowchart of yet another network recovery method according to an embodiment of the present invention;
[0028] Figure 4 is a structural block diagram of the network recovery device according to an embodiment of the present invention;
[0029] Figure 5 is a schematic hardware structure diagram of the computer device according to an embodiment of the present invention. Detailed Embodiments
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0031] In contemporary mobile communication networks, the Radio Access Network (RAN) is responsible for connecting user equipment to the Core Network (CN), while the core network is responsible for processing and transmitting user data and providing various communication services. The communication stability between the access network and the core network is crucial. Once a failure occurs, it may lead to large-scale interruption of user communication, seriously affecting the user experience.
[0032] As a key communication infrastructure for smart factories, if a failure occurs in the 6th Generation Mobile Communication Technology (6G) network, it will affect the connection between the RAN and the CN, thereby interfering with the production and management processes of the factory. Therefore, it is necessary to establish an efficient fault detection and automatic recovery mechanism to ensure the smooth operation of the production process.
[0033] However, the current fault detection and recovery mechanisms have a series of problems, such as low efficiency, untimely detection, long recovery time, poor accuracy, etc. These problems cannot meet the growing communication service demands. Therefore, there is an urgent need for a faster, more accurate, and efficient mechanism to ensure reliable communication between the RAN and the CN.
[0034] Currently, some patent solutions are dedicated to improving the fault detection and automatic recovery capabilities of the core network. For example, by collecting data through distributed links and combining key performance indicators, it is possible to quickly locate and troubleshoot faults, thereby improving the network repair efficiency. However, these solutions often introduce a distributed link tracing system and eBPF technology, increasing system complexity and requiring professional technical personnel for maintenance, resulting in increased manual intervention, decreased response speed, and accuracy. In addition, these solutions have weak customization capabilities and are difficult to be flexibly adjusted according to the needs of different departments or services. Their functions are relatively single and difficult to meet complex business processes and diverse customer requirements.
[0035] In view of this, the technical solution of the present invention provides a fault detection and recovery mechanism between the access network and the core network in a high-speed running network, which can accurately and quickly detect faults between the RAN and the CN, efficiently and automatically recover the environment, reduce operation and maintenance costs, and improve the communication stability between the access network and the core network.
[0036] According to an embodiment of the present invention, an embodiment of a network recovery method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0037] In this embodiment, a network recovery method is provided, which can be used in computer devices such as servers, network devices, or server clusters, etc. Figure 1 It is a flowchart of the network recovery method according to the embodiment of the present invention, as Figure 1 shown, and the process includes the following steps:
[0038] Step S101, obtain the network operation parameters of the current network node.
[0039] The current network node refers to the key device nodes in the core network (CN) or access network (RAN) of the 6G communication network, such as base station controllers, core network servers, edge computing nodes, etc. These nodes are responsible for data transmission, service processing, and user access, and are the core units for the stable operation of the network.
[0040] The network operation parameters refer to multi-dimensional performance data collected in real time from network nodes. For example, they can include: traffic metrics (such as bandwidth utilization, packet throughput), device status (such as CPU / memory utilization, hard disk health), link status (such as latency, packet loss rate, connection stability), network element process status (such as process survival status, response time), signal strength (signal quality in the radio access network), etc. These parameters are used to dynamically evaluate the health status of the node. Specifically, through a variety of data collection tools and protocols (such as SNMP, eBPF, API interfaces, etc.), the multi-dimensional network operation parameters of 6G network nodes can be continuously monitored. For example, for traffic metrics, bandwidth utilization, packet throughput, session connection count, etc. can be captured through deep packet inspection (DPI) or traffic probes; for device status, CPU / memory utilization, disk health, temperature, etc. can be obtained using hardware monitoring interfaces (such as IPMI); for link status, latency, packet loss rate, and link on / off status can be monitored through ICMP probes or BGP protocols; for network element process status, the survival status, response time, and resource occupancy of key processes (such as base station controller processes) can be checked by calling system APIs or custom scripts; for user behavior data, user access frequency, service request types, etc. can be extracted from the core network logs. The multi-dimensional network operation parameter data is summarized to the central database at a second-level or millisecond-level frequency and is displayed in real time through a visualization operation and maintenance platform, providing a basis for subsequent analysis.
[0041] Step S102, analyze the network operation parameters, and based on the analysis results, determine whether a network failure has occurred in the current network node.
[0042] The analysis result refers to the fault determination conclusion generated after comprehensively calculating the network operation parameters. Specifically, the acquisition tools (such as SNMP, eBPF, etc.) collect and store various network operation parameters in real time. By calculating and correlating these parameters, abnormal patterns can be identified. For example, if the bandwidth utilization rate continuously approaches or reaches the maximum value, or the device status (such as CPU / memory utilization rate) is too high, it may indicate that the device is overloaded, affecting normal operation. An increase in the packet loss rate or excessive delay in the link status may also be a sign of a network link failure. And the downtime or excessive response time of key processes in the process status may be a warning signal of a device or service failure. By combining multi-dimensional data and comprehensively analyzing these network operation parameters through set thresholds and predefined alarm policies, a conclusion can be drawn on whether there is a fault in the network node.
[0043] Step S103, if the analysis result indicates that a network fault has occurred in the current network node, determine the target fault information of the network fault.
[0044] The target fault information refers to the specific fault type and location information extracted from the analysis result. For example, it may include: fault type (such as hardware fault, software exception, configuration error, data inconsistency, etc.), fault location (such as specific network element, link, etc.), severity level (such as critical fault or minor fault), etc. Specifically, by matching the analysis result with the network topology and node status, the specific fault location can be located, and it can be determined whether it is a hardware fault (such as server failure, power supply failure, etc.), a software fault (such as service crash, abnormal restart, etc.), or a configuration or data error. This process can be completed by an automated analysis tool, which quickly generates a fault determination based on historical data and network models and provides specific target fault information, such as the judgment of fault type, location, and severity level.
[0045] Step S104, according to the mapping relationship between the fault information and the network recovery strategy, determine the target recovery strategy corresponding to the target fault information.
[0046] The mapping relationship refers to a preset rule library of the correspondence between fault information and recovery strategies. For example: in case of a hardware fault, trigger the standby machine network construction module and switch to the standby node; in case of a link interruption, dynamically adjust the routing strategy and enable redundant links; in case of a software exception, automatically restart the service or roll back to a stable version, etc. Specifically, this mapping relationship can be implemented through a configuration file or a database, supporting flexible expansion.
[0047] The target recovery strategy refers to the specific automated recovery operations formulated based on the fault information. Specifically, a mapping rule library of fault types and recovery strategies is preset. For example, for hardware faults, trigger the VRRP protocol to switch to the standby node; for link interruptions, enable redundant links and update the routing table; for data inconsistencies, start the MySQL active-active synchronization mechanism, etc. Retrieve the corresponding rule library according to the target fault information and select the optimal recovery strategy. When multiple strategies are applicable, perform priority arbitration based on preset weights (such as recovery speed, resource consumption), and finally generate a specific recovery strategy.
[0048] Step S105, perform network recovery on the current network node according to the target recovery strategy.
[0049] Automatically execute the target recovery strategy, including operations such as system configuration changes, routing adjustments, and service restarts, to quickly restore the network from the fault state to the normal operating state. This process is controlled and monitored through an automated operation and maintenance platform to ensure the efficiency and accuracy of network recovery.
[0050] The network recovery method provided by the embodiments of the present invention can quickly and accurately identify network faults by obtaining the running parameters of network nodes in real time and constructing an intelligent analysis mechanism, overcoming the pain point of low efficiency in traditional manual troubleshooting. Based on the preset mapping relationship library of fault information and recovery strategies, an automated closed-loop of fault location and recovery strategy matching is realized, greatly shortening the fault response time. Therefore, this method adopts a hierarchical and progressive logical architecture, which not only ensures the accuracy of fault diagnosis but also enhances the system's self-healing ability through strategic recovery, greatly reducing the need for manual intervention and operation and maintenance costs, and effectively improving the communication stability and high availability of the network.
[0051] In this embodiment, a network recovery method is provided, which can be used in computer devices such as servers, network devices, or server clusters, Figure 2 is a flowchart of the network recovery method according to the embodiments of the present invention, as Figure 2 shown, and this process includes the following steps:
[0052] Step S201, obtain the network running parameters of the current network node. For details, please refer to Figure 1 Step S101 of the embodiment shown here, which will not be elaborated further.
[0053] Step S202, analyze the network running parameters, and based on the analysis results, determine whether a network fault has occurred in the current network node. For details, please refer to Figure 1 Step S102 of the embodiment shown here, which will not be elaborated further.
[0054] Step S203, if the analysis results indicate that a network fault has occurred in the current network node, determine the target fault information of the network fault.
[0055] Specifically, the above-mentioned step S203 includes:
[0056] Step S2031: Extract abnormal operation parameters corresponding to network faults from network operation parameters.
[0057] Abnormal operation parameters refer to the parameters in network operation parameters that exceed the preset parameter threshold and whose continuous duration exceeds the preset duration. Specifically, for each network node, continuously collect and analyze the corresponding network operation parameters, and capture abnormal behaviors that may indicate faults in real time. For example, if the traffic metrics (such as bandwidth utilization) of a network node continuously approach its maximum capacity, or the CPU / memory utilization rate is abnormally high, it may indicate that there is an overload problem at this node, and the device's carrying capacity is insufficient, which may lead to service interruption or slow response. On the other hand, if the packet loss rate in the link state increases, the latency increases, or the stability of the connection decreases, this may indicate that there is a fault in the network link, such as a link disconnection or a physical layer problem. In addition, in the network element process state, if a key process crashes or the response time exceeds the threshold, it may also be a sign of a device or service failure. Signal strength problems may be manifestations of insufficient signal coverage or increased interference in the radio access network. Therefore, during the collection process of network operation parameters, the set threshold can be used to automatically determine which parameters exceed the normal range, and then determine which are abnormal operation parameters. Data collection tools such as Simple Network Management Protocol (SNMP), extended Berkeley Packet Filter (eBPF), or API interfaces can monitor network nodes in real time and summarize these parameters to a central database for subsequent analysis.
[0058] In some alternative embodiments, the above-mentioned step S2031 includes:
[0059] Step a1: Compare each parameter in the network operation parameters with its corresponding preset parameter threshold to determine the target operation parameters that exceed the preset parameter threshold.
[0060] The preset parameter threshold refers to the critical value of the normal range of the preset network operation parameters (such as upper or lower limits), such as the network element process response time threshold, link bandwidth utilization threshold, etc. The target operation parameters refer to the parameters in the network operation parameters that exceed the preset parameter threshold. Specifically, by collecting various network operation parameters of network nodes in real time, a normal range threshold is preset for each parameter (such as link latency threshold ≤ 50ms, CPU utilization threshold ≤ 80%, etc.). The collected parameters will be compared with the corresponding preset parameter thresholds in real time: for example, if the CPU utilization rate of a certain network element process reaches 90% (exceeding the preset 80%), it will be marked as a potential abnormal operation parameter, that is, the target operation parameter.
[0061] Step a2: Obtain the continuous duration for which the target operation parameters exceed the preset parameter threshold.
[0062] The continuous duration refers to the duration during which the target operating parameter exceeds the preset parameter threshold. Specifically, a timer mechanism is used to continuously track the target operating parameter. When a certain parameter first exceeds the threshold, the timer starts and records the start timestamp; if the parameter still continuously exceeds the threshold in subsequent monitoring cycles, the timer accumulates the duration. For example, if the link delay continuously exceeds 50 ms for 3 seconds, the timer records 3 seconds. If the parameter returns to normal during the monitoring cycle, the timer is reset. The calculation of the continuous duration is based on the timestamp difference, ensuring accuracy to the millisecond level.
[0063] Step a3, if the continuous duration exceeds the preset duration, determine the target operating parameter as an abnormal operating parameter.
[0064] The preset duration refers to the time threshold for determining whether the target operating parameter constitutes an abnormality, and can also be called the fault determination duration. Specifically, a fault determination duration is preset for each type of parameter (for example, the network element status needs to continuously exceed the threshold for 5 seconds for an abnormality). When the continuous duration recorded by the timer exceeds this preset duration, mark this parameter as an abnormal operating parameter. For example, if the network element status is abnormally continuous and exceeds the threshold for 6 seconds (exceeding the preset duration of 5 seconds), it is determined as abnormal. This mechanism avoids misjudgments caused by short-term fluctuations (such as instantaneous peaks), and only triggers the subsequent fault handling process for persistent abnormalities.
[0065] In the above embodiments, potential abnormal operating parameters are quickly screened based on the preset parameter threshold, avoiding the computational redundancy of full-scale data analysis. The introduction of the continuous duration criterion effectively distinguishes instantaneous fluctuations from persistent faults, solving the problem of false triggering caused by occasional peaks in the traditional threshold method. Through the dual verification mechanism, it not only ensures the rapid locking of key abnormal operating parameters, but also enhances the judgment reliability through time-dimensional verification, and can accurately filter out interference factors such as network jitter.
[0066] Step S2032, match the abnormal operating parameter with the fault parameter in the fault identification condition, and determine the target fault information corresponding to the network fault based on the fault matching result.
[0067] Fault identification conditions refer to a set of pre - set rules for identifying the types and locations of network faults, including fault parameters and their associated relationships. Fault parameters refer to the parameters used to describe specific fault characteristics in fault identification conditions (such as fault types, key indicators corresponding to fault points). Specifically, a clear fault identification model or rule base is established in advance. This rule base contains various fault identification conditions based on historical data, experience, and network topology, such as excessive bandwidth utilization, abnormal increase in packet loss rate, downtime of key processes, etc. By matching with these fault identification conditions, possible fault types can be identified from abnormal network operation parameters. For example, if the bandwidth utilization rate is close to 100% for a long time and is accompanied by high CPU or memory utilization at the same time, this abnormal pattern can be matched with the fault conditions related to hardware overload or traffic attacks to further determine whether this is the result of over - loaded operation of network resources. Similarly, if the packet loss rate of a link exceeds the set threshold and the delay increases sharply, this may match the fault conditions of link interruption or degraded link quality, thus determining it as a link fault. The core of fault matching is to compare abnormal operation parameters with predefined fault conditions one by one through dynamic calculation and analysis to quickly locate problems and determine the fault type (such as hardware, software, link, etc.) and its location (such as network element, device, link, etc.). According to the matching results, the severity of the fault (such as critical fault or minor fault) can be further determined, and corresponding target fault information can be generated, including detailed information such as fault type, location, and severity.
[0068] In some alternative embodiments, step S2032 described above includes:
[0069] Step b1, obtaining a parameter tree structure composed of fault parameters.
[0070] The parameter tree structure refers to a set of fault parameters organized in a tree - like hierarchical structure, used to hierarchically describe fault characteristics and associated relationships. Specifically, the parameter tree structure is a logical model pre - constructed through the fault tree analysis method (FTA), which hierarchically associates fault types with parameter abnormalities. For example, the root node is "network communication fault", and the child nodes can include "link interruption", "device overload", "configuration error", etc.; each child node is further associated with specific parameters (such as "link interruption" corresponding to abnormalities in "link delay" and "packet loss rate"). This tree - like structure is loaded through a configuration file or database to ensure that abnormal operation parameters can be hierarchically matched with potential fault causes during fault analysis.
[0071] Step b2, layer - by - layer comparing the abnormal operation parameters and fault parameters according to the parameter tree structure to obtain the layer - by - layer comparison result of the parameters.
[0072] The result of layer-by-layer parameter comparison refers to the intermediate result of hierarchical matching of abnormal operating parameters and fault parameters according to the parameter tree structure. Specifically, starting from the root node of the parameter tree, drill down layer by layer for matching. For example, if both "link delay" and "packet loss rate" are detected as abnormal, the "link interruption" branch is preferentially matched, and other child nodes (such as "equipment overload") are excluded. Each layer of comparison needs to verify whether the abnormal operating parameters meet the conditions of this branch: if the match is successful, enter the next level; if the match fails, backtrack to the previous level and try other branches. Finally, one or more possible fault paths (such as "link interruption → optical module failure") are generated as the result of layer-by-layer comparison.
[0073] Step b3, analyze the result of layer-by-layer parameter comparison to determine the target fault information that matches the abnormal operating parameters.
[0074] Based on the fault paths in the comparison result, combined with weight analysis (such as historical fault statistics, expert experience library) and context information (such as device logs, topology status), calculate the confidence level of each path. For example, if the optical module temperature is detected as abnormal under the "link interruption" path, the confidence level is increased to 90%. Finally, select the path with the highest confidence level as the target fault information (such as "hardware failure of a certain module").
[0075] In the above embodiment, through the hierarchical screening feature of the parameter tree, the fault range is gradually narrowed from the global to the local, avoiding the misjudgment risk caused by data redundancy or cross-interference in traditional single-parameter matching. During the layer-by-layer comparison process, each level conducts directional verification based on specific fault characteristics, effectively excluding the interference of irrelevant parameters, quickly locking the core abnormal indicators, and using the logical relevance of the tree structure to realize multi-dimensional verification of fault types. This hierarchical analysis mode not only greatly shortens the diagnosis time in complex fault scenarios but also enhances the interpretability and reliability of fault identification through the structured parameter matching path.
[0076] Step S204, according to the mapping relationship between the fault information and the network recovery strategy, determine the target recovery strategy corresponding to the target fault information. For details, please refer to Figure 1 Step S104 of the embodiment shown, which will not be elaborated here.
[0077] Step S205, perform network recovery on the current network node according to the target recovery strategy. For details, please refer to Figure 1 Step S105 of the embodiment shown, which will not be elaborated here.
[0078] The network recovery method provided by the embodiment of the present invention can accurately extract abnormal operating parameters from network operating parameters and perform intelligent parameter matching with predefined fault identification conditions, quickly lock the fault characteristics, accurately identify the fault type and fault point, and effectively avoid the subjective misjudgment and delay problems in traditional manual troubleshooting.
[0079] In this embodiment, a network recovery method is provided, which can be used in computer devices such as servers, network devices, or server clusters, etc. Figure 3 It is a flowchart of the network recovery method according to the embodiment of the present invention, as Figure 3 shown, the process includes the following steps:
[0080] Step S301, obtain the network operation parameters of the current network node. For details, please refer to Figure 2 step S201 of the embodiment shown, which will not be elaborated here.
[0081] Step S302, analyze the network operation parameters, and based on the analysis result, determine whether the current network node has a network fault. For details, please refer to Figure 2 step S202 of the embodiment shown, which will not be elaborated here.
[0082] Step S303, if the analysis result indicates that the current network node has a network fault, determine the target fault information of the network fault. For details, please refer to Figure 2 step S203 of the embodiment shown, which will not be elaborated here.
[0083] Step S304, according to the mapping relationship between the fault information and the network recovery policy, determine the target recovery policy corresponding to the target fault information. For details, please refer to Figure 2 step S204 of the embodiment shown, which will not be elaborated here.
[0084] Step S305, perform network recovery on the current network node according to the target recovery policy.
[0085] Specifically, the above step S305 includes:
[0086] Step S3051, determine whether the current network fault node is the master node.
[0087] The master node refers to the core node that undertakes the main service processing and data transmission in the 6G system, and is the main bearing node of the service traffic under normal conditions. Specifically, the master node is identified through the configuration file or the role setting of the network device. For example, the master node and the standby node are identified through protocols such as the Virtual Router Redundancy Protocol (VRRP). Determine whether the current faulty node is the master node by accessing the network topology information and querying the role and configuration of the network node. In addition, in the operation and maintenance platform of the 6G network, the master node will be identified as the main processing unit, with higher priority and resource usage. If a fault occurs, it can be determined whether the current network fault node is the master node through fault analysis and automatic identification of the node role.
[0088] In some alternative embodiments, the above step S3051 includes:
[0089] Step c1, monitor whether the standby node receives the heartbeat packet sent by the master node within a preset duration.
[0090] Step c2, if the standby node does not receive the heartbeat packet within the preset duration, determine that the current network failure node is the master node.
[0091] The standby node is the backup node of the master node and is usually in a standby state. When the master node fails, it automatically takes over the service. The heartbeat packet refers to the status information packet regularly sent by the master node to the standby node, which is used to verify the survival status of the master node and the effectiveness of the communication link. Specifically, the standby node realizes the monitoring of the heartbeat packet through a periodic listening mechanism. The master node sends a heartbeat packet to the standby node at a fixed time interval (such as once per second), and the standby node internally maintains a timer. Each time a heartbeat packet is received, the timer is reset. If the timer exceeds the preset duration (such as 3 seconds) and no new heartbeat packet is received, the standby node determines that the master node has failed. This process combines the state machine logic of the VRRP protocol to ensure the real-time and reliability of the monitoring.
[0092] Step S3052, if the current network failure node is the master node, determine the standby node corresponding to the master node.
[0093] The master node and the standby node are bound through a pre-configured redundancy group relationship. During system initialization, each master node is associated with one or more standby nodes and recorded in the configuration file. When the master node is determined to be faulty, the corresponding standby node is automatically selected according to the predefined mapping relationship (such as based on the VRRP group ID or node label). For example, in a dual-active hot standby architecture, the master node and the standby node belong to the same VRRP group, and the priority of the standby node is only second to that of the master node. By querying the priority list of the nodes in the group, the standby node with the highest priority is determined as the takeover candidate.
[0094] Step S3053, synchronize the service data of the master node to the standby node, and map the virtual address of the master node to the standby node, so that subsequent service data is directed to the standby node through the virtual address.
[0095] Service data refers to all data generated by users in the 6G system, including data that needs to be processed and transmitted in real time, such as registration information, communication data, and cached data. The virtual address refers to the virtual IP address (VRRP IP) based on the VRRP protocol, which is the external service address shared by the primary node and the standby node. Specifically, data synchronization is achieved through a dual-active architecture and a real-time replication mechanism. The service data of the primary node (such as database transactions and user session information) will be synchronized to the standby node in real time. For example, the MySQL dual-active architecture is used to ensure data consistency. The virtual address mapping depends on the VRRP protocol and the keepalived election mechanism. After the primary node fails, the standby node sends a gratuitous ARP broadcast to claim the takeover of the virtual IP address, causing network devices (such as switches and routers) to redirect traffic to the standby node. At the same time, the load balancer or DNS service will automatically update the routing table according to the change of the virtual IP to ensure that subsequent service traffic is seamlessly directed to the standby node.
[0096] Step S3054, restore the network of the primary node on the standby node.
[0097] After the standby node takes over, it reconstructs the network service through an automated recovery process. Specifically, the standby node activates the virtual IP and broadcasts ARP messages to update the routing table of the network device. Start the same network element processes as the original primary node (such as core network services and user authentication services), and load the synchronized service data. Check the integrity of data synchronization to ensure that the cache (such as Redis) and database (such as MySQL) are consistent with the state before the primary node failure. Switch the service traffic to the standby node through the load balancer or SDN controller to complete the network recovery.
[0098] For example, when the network fluctuates, the fault notification email mechanism of Keepalived can monitor the state changes of nodes in real time, determine the current role of the node (primary node or standby node), and trigger corresponding email notifications. At the same time, analyze the cause, point, and type of the fault through methods such as fault tree analysis and data analysis, and upload the data to the visualization operation and maintenance platform so that the operation and maintenance team can understand the network status in a timely manner and take corresponding measures.
[0099] In addition, the fault tolerance and reliability of the network can be enhanced by designing multiple link transmission mechanisms in the network, so as to meet scenarios with high requirements for data transmission continuity. At the same time, the elastic recovery mechanism of network function virtualization can also be adopted. Through virtualization technology, network functions are abstracted into software instances, reducing the dependence on hardware and improving resource utilization, which is suitable for scenarios that require high elasticity and flexibility.
[0100] The network recovery method provided by the embodiments of the present invention adopts the main node status dynamic detection technology, combines the transparent migration strategy of virtual address mapping, and completes the seamless service switching on the premise of ensuring the complete synchronization of data, avoiding the risks of service interruption and data loss commonly seen in traditional primary-backup switching. By decoupling the physical node topology at the virtual address layer, the failover process is completely transparent to the upper-layer services, effectively ensuring the coherence of the user experience.
[0101] In some alternative embodiments, the above step S305 further includes:
[0102] Step d1, when the duration of the communication link interruption between the primary node and the standby node exceeds a preset duration, upgrade the standby node to the primary node status.
[0103] When the communication link between the primary and standby nodes is interrupted for more than the preset duration (e.g., 10 seconds), the standby node autonomously upgrades to the "MASTER" status according to the timeout trigger mechanism. Specifically, the standby node detects the continuous loss of heartbeat packets and starts the timeout countdown. After the timeout, the standby node promotes to the primary node according to the VRRP protocol priority rules (compare the IP address sizes if the priorities are the same). The standby node sends a VRRP advertisement message, declaring itself as the new primary node, and takes over the virtual IP. Update the topology information and notify other network nodes (such as gateways, core network devices) to switch the traffic path.
[0104] Step d2, when the communication link between the primary node and the upgraded standby node is restored, re-determine the target primary node from the primary node and the upgraded standby node based on a preset priority.
[0105] The preset priority refers to the priority parameter pre-configured for the node. The target primary node refers to the only legal primary node determined through the preset priority or election mechanism during the primary-primary competition or fault recovery process. Specifically, the original primary node and the upgraded standby node (the current primary node) send heartbeat packets to each other and exchange priority information. If the original primary node has a higher priority, the "preemption mode" is triggered, and the original primary node becomes the MASTER again; if the priorities are the same, the primary node is determined according to the predefined rules (such as the IP address size). After determining the target primary node, the virtual IP switches back to the node with a higher priority, and the service data and user context information are synchronized. The non-primary nodes enter the "BACKUP" state, stop the service, and enter the listening mode to ensure the stability of the network topology.
[0106] Step d3, restore the network element process status of the target primary node and restore the user context information to the target primary node.
[0107] The network element process status refers to the running status of processes in a 6G system's network elements (such as servers and network devices), including whether there is overload, timeout without response, abnormal exit of child processes, etc. User context information refers to the session status data of users in the system, including registration information, connection status, service configuration, etc. Specifically, restore the network element processes (such as user authentication service, signaling processing service) before the fault according to the log or snapshot, and load the persistent data (such as the user session table in MySQL). Synchronize the latest user context information (such as online status, service permissions) from the standby node to ensure service continuity. For example, switch the cached data back to the primary node through the Redis sentinel mechanism. Compare the data differences between the primary and standby nodes, and complete the missing or conflicting data through incremental synchronization. Re-register with the core network and notify the surrounding nodes (such as base stations, gateways) to update the routing information, and gradually take over the service traffic.
[0108] For example, when the abnormal state of a network element process in the 6G system reaches a critical value (such as the network element status being abnormally continuous for more than the threshold for 6 seconds and exceeding the preset duration of 5 seconds), it will automatically trigger the standby network construction and automation mechanism, switch the service traffic to the standby node, and at the same time record the status information of the faulty node for subsequent analysis and repair. First, record the network element process to be detected according to parameter information such as processName and checkServices. Secondly, initialize the information of the network element status detection process, obtain the status information of each network element process, and record it in the log. If the network element process is in an abnormal state such as overload, timeout without response, abnormal exit of child processes, etc., and does not return to normal within the specified time, it will automatically trigger the standby network construction and automation mechanism, and switch the service traffic to the standby node. After the service is switched to the standby machine, end the network element process status fault detection program of the original host.
[0109] In the above embodiment, when the communication between the primary and standby nodes times out, it automatically triggers the upgrade of the standby node to the primary node to ensure seamless migration of the service traffic to the standby node and avoid service interruption caused by link failure. After the communication is restored, re-arbitrate the primary node role based on the preset priority, effectively avoiding the risk of dual-primary competition, and ensuring the integrity of service data and service consistency before and after the switch through the accurate restoration of user context information and the rapid recovery of network element process status.
[0110] In some optional embodiments, when the cache service of the primary node is in a fault state, step S305 further includes:
[0111] Step e1, determine a new primary node from multiple slave nodes using a preset election method.
[0112] The preset election method refers to the election mechanism used to determine the master node. Specifically, when the cache service (such as Redis) of the master node fails, the sentinel election mechanism is used to determine the new master node from multiple slave nodes. The sentinel node (monitoring instance) continuously detects the health status of the master node. After the master node loses connection, the sentinel initiates an election vote and screens candidate slave nodes according to preset rules (such as the priority of the slave node and the data synchronization offset). The slave node that obtains the consent of the majority of sentinels is promoted to the new master node and updates the system configuration. The sentinel broadcasts the address of the new master node to all slave nodes to ensure that the slave nodes switch data synchronization targets.
[0113] Step e2, updating the master node addresses stored in the slave nodes other than the new master node, so that the new master node takes over the cache service of the master node.
[0114] After the new master node is determined, the IP and port information of the new master node is sent to all slave nodes through Sentinel or configuration management tools (such as ZooKeeper). The slave node disconnects from the original master node and establishes a data synchronization link with the new master node. The new master node sends full or incremental data to the slave nodes to ensure data consistency. The client automatically routes the request to the new master node by subscribing to Sentinel notifications or querying the configuration center, completing a seamless switch.
[0115] In the above implementation, the dynamic election strategy significantly shortens the master-slave switching time and avoids the delay problem of traditional manual intervention. The distributed address synchronization mechanism is used to ensure that all slave nodes point to the new master node in real time, eliminating the risk of data inconsistency. The performance-first election logic is combined to optimize resource allocation and improve the overall load balancing capability and stability of the system.
[0116] In this embodiment, a network recovery device is also provided, which is used to implement the above embodiments and preferred implementations, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0117] This embodiment provides a network recovery device, such as Figure 4 As shown, including:
[0118] The acquisition module 401 is used to acquire the network operation parameters of the current network node;
[0119] The analysis module 402 is used to analyze the network operation parameters and determine whether a network failure occurs at the current network node based on the analysis results;
[0120] The first determination module 403 is configured to determine the target fault information of the network fault if the analysis result indicates that a network fault has occurred in the current network node;
[0121] The second determination module 404 is configured to determine the target recovery policy corresponding to the target fault information according to the mapping relationship between the fault information and the network recovery policy;
[0122] The recovery module 405 is configured to perform network recovery on the current network node according to the target recovery policy.
[0123] In some alternative embodiments, the first determination module 403 includes:
[0124] An extraction sub-module, configured to extract abnormal operation parameters corresponding to the network fault from the network operation parameters;
[0125] A matching sub-module, configured to perform parameter matching between the abnormal operation parameters and the fault parameters in the fault identification conditions, and determine the target fault information corresponding to the network fault based on the fault matching result.
[0126] In some alternative embodiments, the extraction sub-module includes:
[0127] A comparison unit, configured to compare each parameter in the network operation parameters with its corresponding preset parameter threshold, and determine the target operation parameters that exceed the preset parameter threshold;
[0128] A first acquisition unit, configured to acquire the continuous duration for which the target operation parameters exceed the preset parameter threshold;
[0129] A first determination unit, configured to determine the target operation parameters as abnormal operation parameters if the continuous duration exceeds the preset duration.
[0130] In some alternative embodiments, the matching sub-module includes:
[0131] A second acquisition unit, configured to acquire the parameter tree structure formed by the fault parameters;
[0132] A layer-by-layer comparison unit, configured to perform layer-by-layer comparison between the abnormal operation parameters and the fault parameters according to the parameter tree structure, and obtain the parameter layer-by-layer comparison result;
[0133] An analysis unit, configured to analyze the parameter layer-by-layer comparison result, and determine the target fault information that matches the abnormal operation parameters.
[0134] In some alternative embodiments, the recovery module 405 includes:
[0135] A judgment sub-module, configured to judge whether the current network fault node is the main node;
[0136] The first determination sub-module is used to determine a standby node corresponding to the master node if the current network failure node is the master node;
[0137] The transmission sub-module is used to synchronize the service data of the master node to the standby node and map the virtual address of the master node to the standby node, so that subsequent service data can be directed to the standby node through the virtual address;
[0138] The first recovery sub-module is used to recover the network of the master node at the standby node.
[0139] In some alternative embodiments, the recovery module 405 further includes:
[0140] The upgrade sub-module is used to upgrade the standby node to the master node status when the duration of the interruption of the communication link between the master node and the standby node exceeds a preset duration;
[0141] The second determination sub-module is used to re-determine the target master node from the master node and the upgraded standby node based on a preset priority when the communication link between the master node and the upgraded standby node is restored;
[0142] The second recovery sub-module is used to recover the network element process status of the target master node and restore the user context information to the target master node.
[0143] In some alternative embodiments, the recovery module 405 further includes:
[0144] The third determination sub-module is used to determine a new master node from multiple slave nodes by using a preset election method;
[0145] The update sub-module is used to update the master node address stored by other slave nodes except the new master node, so that the new master node takes over the cache service of the master node.
[0146] In some alternative embodiments, the judgment sub-module includes:
[0147] The monitoring unit is used to monitor whether the standby node receives a heartbeat packet sent by the master node within a preset duration;
[0148] The second determination unit is used to determine that the current network failure node is the master node if the standby node does not receive a heartbeat packet within a preset duration.
[0149] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be repeated here.
[0150] The network recovery device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0151] The network recovery device provided by the embodiment of the present invention can quickly and accurately identify network faults by obtaining the operation parameters of network nodes in real time and constructing an intelligent analysis mechanism, overcoming the pain point of low efficiency in traditional manual troubleshooting. Based on the mapping relationship library of preset fault information and recovery strategies, an automated closed-loop of fault location and recovery strategy matching is realized, greatly shortening the fault response time. Therefore, this device not only ensures the accuracy of fault diagnosis, but also enhances the system self-healing ability through strategic recovery, greatly reducing the need for manual intervention and operation and maintenance costs, and effectively improving the communication stability and high availability of the network.
[0152] The embodiment of the present invention also provides a computer device having the above Figure 4 shown network recovery device.
[0153] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention. As shown in Figure 5 , the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 5 In
[0154] Processor 10 can be a central processor, a network processor, or a combination thereof. Among them, processor 10 can further include a hardware chip. The above hardware chip can be an application specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device can be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.
[0155] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the methods shown in the above embodiments.
[0156] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely disposed relative to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0157] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memory.
[0158] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0159] The embodiments of the present invention further provide a computer-readable storage medium. The methods according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and to be stored in a local storage medium, so that the methods described herein can be stored in such software processes on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium may be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium may further include a combination of the above types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.
[0160] A part of the present invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the present invention through the operations of the computer. Those skilled in the art should understand that the forms in which computer program instructions exist in a computer-readable medium include but are not limited to source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0161] Although embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A network recovery method, characterized in that, The method includes: Obtaining the network operation parameters of the current network node; Analyzing the network operation parameters, and determining whether the current network node has a network fault based on the analysis result; If the analysis result indicates that the current network node has a network fault, determining the target fault information of the network fault; Determining the target recovery policy corresponding to the target fault information according to the mapping relationship between the fault information and the network recovery policy; Performing network recovery on the current network node according to the target recovery policy.
2. The method according to claim 1, characterized in that The step of, if the analysis result indicates that the current network node has a network fault, determining the target fault information of the network fault includes: Extracting abnormal operation parameters corresponding to the network fault from the network operation parameters; Performing parameter matching between the abnormal operation parameters and the fault parameters in the fault identification conditions, and determining the target fault information corresponding to the network fault based on the fault matching result.
3. The method according to claim 2, wherein The step of extracting abnormal operation parameters corresponding to the network fault from the network operation parameters includes: Comparing each parameter in the network operation parameters with its corresponding preset parameter threshold, and determining the target operation parameters that exceed the preset parameter threshold; Obtaining the continuous duration for which the target operation parameters exceed the preset parameter threshold; If the continuous duration exceeds the preset duration, determining the target operation parameters as the abnormal operation parameters.
4. The method according to claim 2, wherein The step of performing parameter matching between the abnormal operation parameters and the fault parameters in the fault identification conditions, and determining the target fault information corresponding to the network fault based on the fault matching result includes: Obtaining the parameter tree structure formed by the fault parameters; Performing layer-by-layer comparison of the abnormal operation parameters and the fault parameters according to the parameter tree structure to obtain the layer-by-layer comparison result of the parameters; Analyzing the layer-by-layer comparison result of the parameters to determine the target fault information that matches the abnormal operation parameters.
5. The method according to claim 1, characterized in that, The step of performing network recovery on the current network node according to the target recovery policy includes: Judging whether the current network fault node is the master node; If the current network fault node is the master node, determining the standby node corresponding to the master node; Synchronizing the service data of the master node to the standby node, and mapping the virtual address of the master node to the standby node, so that subsequent service data is directed to the standby node through the virtual address; Recovering the network of the master node on the standby node.
6. The method according to claim 5, wherein It further includes: When the communication link interruption duration between the master node and the standby node exceeds the preset duration, upgrading the standby node to the master node status; When the communication link between the master node and the upgraded standby node is restored, re-determining the target master node from the master node and the upgraded standby node based on the preset priority; Restoring the network element process status of the target master node, and restoring the user context information to the target master node.
7. The method according to claim 5, wherein When the cache service of the master node is in a fault state, the step of performing network recovery on the current network node according to the target recovery policy includes: Determining a new master node from multiple slave nodes by using a preset election method; Update the master node addresses stored in slave nodes other than the new master node, so that the new master node takes over the cache service of the master node.
8. The method according to claim 5, characterized in that, The determination of whether the current network failure node is the master node includes: Monitoring whether the standby node receives a heartbeat packet sent by the master node within a preset time period; If the standby node does not receive the heartbeat packet within the preset time period, determine that the current network failure node is the master node.
9. A network recovery device, characterized in that, The device includes: An acquisition module, configured to acquire network operation parameters of a current network node; An analysis module, configured to analyze the network operation parameters and determine whether a network failure occurs in the current network node based on the analysis result; A first determination module, configured to determine target failure information of the network failure if the analysis result indicates that a network failure occurs in the current network node; A second determination module, configured to determine a target recovery policy corresponding to the target failure information according to a mapping relationship between the failure information and a network recovery policy; A recovery module, configured to perform network recovery on the current network node according to the target recovery policy.
10. A computer device, characterized in that, Includes: A memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the network recovery method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause a computer to execute the network recovery method according to any one of claims 1 to 8.
12. A computer program product, characterized in that, Includes computer instructions, and the computer instructions are used to cause a computer to execute the network recovery method according to any one of claims 1 to 8.
Citation Information
Cited By
Dual-mode self-healing method of automatic test system based on virtualization hierarchical monitoring
CN120578561A
Gigabit passive optical network (GPON) fault intelligent diagnosis and self-healing system integrated with deep learning
CN120856538A
Network equipment fault intelligent detection method and system based on data transmission
CN120934995A
Intelligent alarm method and device based on keepalived
CN121664621A
10BASE-T1S network fault rapid positioning method and system
CN122093241A