Fault recovery method and device, computer device, readable storage medium and program product
By controlling the number and success rate of virtual network element restarts, the network impact problem of virtualized network elements when physical devices fail is solved, achieving stable network recovery and efficient fault handling.
Patent Information
- Application Number
- CN202411683846.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-11-22
AI Technical Summary
In the new metropolitan area network and converged edge scenarios of cloud-network convergence, virtualized network elements frequently restart when physical hardware devices fail, causing network shocks and secondary faults. Existing technologies are unable to effectively control the restart speed and connection process, affecting the reliability of network communication.
By determining the actual number of restarts and the success rate at the target time point, the restart process of virtual network elements is controlled using a preset number adjustment strategy until the fault conditions are met, thereby achieving smooth network reconnection and fault recovery.
It avoids virtual network element restart storms, stabilizes user online speeds, quickly restores network connections, and supports efficient fault recovery for various network application scenarios.
Smart Images

Figure CN119652734B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of cloud computing, and in particular to a fault recovery method and device, a computer device, a computer readable storage medium and a computer program product. BACKGROUND
[0002] In the scenario of new metropolitan area networks and converged edges in cloud network convergence, a large number of virtualized network elements are applied to the network at the operation level. When a physical hardware device in the network fails, the virtualized network elements in the network all initiate a reconnection request when determining that they lose network connection, that is, a large number of virtualized network elements restart and establish external connections. Since the virtualized network elements are restarted by cloud software, the restart speed and the connection initiation speed of the virtualized network elements are much higher than the restart speed of the physical hardware devices in the network, which causes an impact on the external network and leads to secondary failures. SUMMARY
[0003] Therefore, it is necessary to provide a fault recovery method, device, computer device, computer readable storage medium and computer program product capable of improving the reliability of communication between virtualized network elements and a network to solve the above technical problems.
[0004] In a first aspect, the present application provides a fault recovery method, comprising:
[0005] If the virtual network element in the network meets a fault condition, the actual restart number of a target time point is determined, the restart number of the target time point corresponding to an initial time point is determined based on the maximum restart rate of the network, the total number of network elements that need to be restarted and the target completion time length;
[0006] At the target time point, the restart of the virtual network element corresponding to the actual restart number of the target time point is initiated, and the restart success rate of the virtual network element at the target time point is determined;
[0007] Based on the restart success rate of the target time point and a preset number adjustment strategy, the actual restart number of a next target time point of the target time point is determined, and based on the actual restart number of the next target time point, the steps of initiating the restart of the virtual network element corresponding to the restart number of the target time point at the target time point and determining the restart success rate of the virtual network element at the target time point are re-executed until it is determined that the virtual network element does not meet the fault condition.
[0008] In one of the embodiments, the target time point is the initial time point, and the step of determining the actual restart number of the target time point if the virtual network element in the network meets the fault condition comprises:
[0009] In a case where it is determined that each virtual network element in the network disconnects from the network and that a fault condition is met, a theoretical restart number of the initial time point is determined, and the theoretical restart number of the initial time point is taken as an actual restart number of the initial time point.
[0010] In one of the embodiments, the method further comprises:
[0011] calculating an initial completion duration based on a maximum restart rate corresponding to the network and a total number of network elements that need to be restarted;
[0012] obtaining a target completion duration based on the initial completion duration and a preset buffer time;
[0013] processing the target time point and the target completion duration by a preset trigonometric function to obtain a first adjustment parameter, and obtaining a second adjustment parameter based on the total number of network elements that need to be restarted and the target completion duration, and determining a sum of the first adjustment parameter and the second adjustment parameter as a theoretical restart number corresponding to the target time point.
[0014] In one of the embodiments, determining an actual restart number of a next target time point of the target time point based on a restart success rate of the target time point and a preset number adjustment strategy comprises:
[0015] calculating an initial adjustment parameter based on the preset number adjustment strategy and a failure rate corresponding to the restart success rate of the target time point;
[0016] determining a target adjustment parameter based on a relationship between the initial adjustment parameter, the failure rate of the target time point and a preset failure rate threshold, and adjusting a theoretical restart number corresponding to the next target time point of the target time point based on the target adjustment parameter to obtain the actual restart number of the next target time point.
[0017] In one of the embodiments, determining a target adjustment parameter based on a relationship between the initial adjustment parameter, the failure rate of the target time point and a preset failure rate threshold comprises:
[0018] if the failure rate of the target time point is greater than or equal to the preset failure rate threshold, determining a difference between a target value and the initial adjustment parameter as the target adjustment parameter; or
[0019] if the failure rate of the target time point is less than the preset failure rate threshold, determining a sum between the target value and the initial adjustment parameter as the target adjustment parameter.
[0020] In one of the embodiments, the actual restart number of the next target time point of the target time point is determined based on the restart success rate of the target time point and a preset number adjustment strategy, and the step of restarting the virtual network element corresponding to the actual restart number of the target time point at the target time point and determining the restart success rate of the virtual network element of the target time point is re-executed based on the actual restart number of the next target time point until it is determined that the virtual network element does not satisfy the fault condition.
[0021] In the preset number adjustment strategy, a restart number matched with the failure rate is queried based on the failure rate corresponding to the restart success rate of the target time point, and the restart number matched with the failure rate is determined as the actual restart number of the next target time point of the target time point.
[0022] In a second aspect, the application further provides a fault recovery device, comprising:
[0023] A first determination module is configured to determine an actual restart number of a target time point if a virtual network element in a network satisfies a fault condition, wherein the restart number corresponding to the initial time point is determined based on a maximum restart rate corresponding to the network, a total number of network elements that need to be restarted, and a target completion time length.
[0024] A starting module is configured to start a restart of a virtual network element corresponding to the actual restart number of the target time point at the target time point, and determine a restart success rate of the virtual network element of the target time point.
[0025] A second determination module is configured to determine an actual restart number of a next target time point of the target time point based on the restart success rate of the target time point and a preset number adjustment strategy, and re-execute the step of starting the restart of the virtual network element corresponding to the actual restart number of the target time point at the target time point and determining the restart success rate of the virtual network element of the target time point based on the actual restart number of the next target time point until it is determined that the virtual network element does not satisfy the fault condition.
[0026] In a third aspect, the application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the following steps when executing the computer program:
[0027] If a virtual network element in a network satisfies a fault condition, an actual restart number of a target time point is determined, wherein the restart number corresponding to the initial time point is determined based on a maximum restart rate corresponding to the network, a total number of network elements that need to be restarted, and a target completion time length.
[0028] A restart of a virtual network element corresponding to the actual restart number of the target time point is started at the target time point, and a restart success rate of the virtual network element of the target time point is determined.
[0029] determining the actual restart number of the next target time point of the target time point based on the restart success rate of the target time point and the preset number adjustment strategy, and re-executing the steps of initiating the restart of the virtual network element corresponding to the restart number of the target time point at the target time point, determining the restart success rate of the virtual network element of the target time point based on the actual restart number of the next target time point of the target time point, until it is determined that the virtual network element does not satisfy the fault condition.
[0030] In a fourth aspect, the present application further provides a computer readable storage medium, having a computer program stored thereon, the computer program being executed by a processor to implement the following steps:
[0031] If the virtual network element in the network satisfies the fault condition, the actual restart number of the target time point is determined, the restart number corresponding to the initial time point is determined based on the maximum restart rate corresponding to the network, the total number of network elements to be restarted, and the target completion time length;
[0032] At the target time point, the restart of the virtual network element corresponding to the actual restart number of the target time point is initiated, and the restart success rate of the virtual network element of the target time point is determined;
[0033] determining the actual restart number of the next target time point of the target time point based on the restart success rate of the target time point and the preset number adjustment strategy, and re-executing the steps of initiating the restart of the virtual network element corresponding to the restart number of the target time point at the target time point, determining the restart success rate of the virtual network element of the target time point based on the actual restart number of the next target time point of the target time point, until it is determined that the virtual network element does not satisfy the fault condition.
[0034] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, the computer program being executed by a processor to implement the following steps:
[0035] If the virtual network element in the network satisfies the fault condition, the actual restart number of the target time point is determined, the restart number corresponding to the initial time point is determined based on the maximum restart rate corresponding to the network, the total number of network elements to be restarted, and the target completion time length;
[0036] At the target time point, the restart of the virtual network element corresponding to the actual restart number of the target time point is initiated, and the restart success rate of the virtual network element of the target time point is determined;
[0037] determine the actual restart number of the next target time point of the target time point based on the restart success rate of the target time point and the preset number adjustment strategy, and re-perform the steps of initiating the restart of the virtual network element corresponding to the restart number of the target time point at the target time point, determining the restart success rate of the virtual network element of the target time point based on the actual restart number of the next target time point of the target time point, until it is determined that the virtual network element does not satisfy the fault condition.
[0038] The above fault recovery method, device, computer equipment, computer readable storage medium and computer program product, wherein the method can comprise: if the virtual network element in the network satisfies the fault condition, determining the actual restart number of the target time point, the restart number corresponding to the initial time point being determined based on the maximum restart rate corresponding to the network, the total number of network elements that need to be restarted and the target completion time length; initiating the restart of the virtual network element corresponding to the actual restart number of the target time point at the target time point, determining the restart success rate of the virtual network element of the target time point; determining the actual restart number of the next target time point of the target time point based on the restart success rate of the target time point and the preset number adjustment strategy, and re-performing the steps of initiating the restart of the virtual network element corresponding to the restart number of the target time point at the target time point, determining the restart success rate of the virtual network element of the target time point based on the actual restart number of the next target time point of the target time point, until it is determined that the virtual network element does not satisfy the fault condition.
[0039] By adopting the method, the restart storm of the virtual network element after the local device in the network fails can be avoided, each virtual network element can be smoothly restarted to initiate network connection during fault recovery, the user online rate can be prevented from fluctuating, the backup device in the network can quickly process the restarted virtual network element, the fault can be quickly recovered, the number of restarts during the restart of the virtual gateway is controlled through the preset number adjustment strategy, the network reconstruction process is efficient, a plurality of number adjustment strategies are supported, and the actual application scenarios of a plurality of networks can be more effectively adapted. BRIEF DESCRIPTION OF DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other related drawings can be obtained by those skilled in the art without creating any inventive labor.
[0041] Figure 1 An application environment diagram of the fault recovery method in an embodiment;
[0042] Figure 2 A flowchart of the fault recovery method in an embodiment;
[0043] Figure 3 Flowchart for determining theoretical restart number of steps in one embodiment;
[0044] Figure 4 Flowchart for determining actual restart number in one embodiment;
[0045] Figure 5 Diagram for system architecture in one embodiment;
[0046] Figure 6 Diagram for curve of user online number per second in one embodiment;
[0047] Figure 7 Diagram for curve of theoretical restart number in one embodiment;
[0048] Figure 8 Diagram for curve of actual restart number in one embodiment;
[0049] Figure 9 Diagram for grouping of virtual network element in one embodiment;
[0050] Figure 10 Flowchart for fault recovery method in another embodiment;
[0051] Figure 11 Structural block diagram of fault recovery device in one embodiment;
[0052] Figure 12 Internal structure diagram of computer device in one embodiment. DETAILED DESCRIPTION
[0053] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not used to limit the present application.
[0054] The fault recovery method provided by the embodiments of the present application can be applied to, for example, Figure 1The application environment shown. The specific application scenario can be a new type of metropolitan area network and a fusion edge scenario in cloud network fusion, which includes a server cluster and a plurality of local terminal devices. The virtual network element (virtualized network element) 100 in the server cluster communicates with the local terminal device 200 in the network through the metropolitan area network. For example, one local terminal device 200 in the network can be connected to a plurality of virtual network elements 100, and the network is also configured with a backup local terminal device of the local terminal device 200. When the local terminal device fails to provide communication for the plurality of virtual network elements connected thereto, the network provided by the embodiment can control the restart speed and the number of restarts of the virtual network element, so that each virtual network element is accurately scheduled to quickly connect to the backup local terminal device and realize smooth restart of the virtual network element.
[0055] In one exemplary embodiment, as shown in Figure 2 A fault recovery method is provided, including steps 202 to 206. Among them:
[0056] Step 202, if the virtual network element in the network meets the fault condition, determine the actual number of restarts at the target time point.
[0057] Among them, the number of restarts corresponding to the initial time point at the target time point is determined based on the maximum restart rate of the network, the total number of network elements that need to be restarted, and the target completion time. The initial time point can be the first moment when the virtual network element in the network starts to restart after determining that the virtual network element meets the fault condition. The actual number of restarts at the target time point can be the number of virtual network elements that actually initiate restart at the target time point. The fault condition can be a judgment condition for determining that the virtual network element in the network has failed, and its content can be that each virtual network element connected to the local terminal device is disconnected, etc. The virtual network element can be a virtualized network element (NFV), which refers to a network device function realized by cloud software, and is usually used in cloud networks.
[0058] Specifically, the network includes a plurality of local terminal devices, each of which can be connected to a plurality of virtual network elements, and each local terminal device can be configured with a corresponding backup local terminal device. In this way, if it is detected that the local terminal device has failed and each virtual network element connected to the local terminal device has been disconnected, it can be determined that the virtual network element in the current network meets the fault condition. Then the actual number of restarts at the initial time point can be determined, for example, the maximum restart rate of the virtual network element corresponding to the network, the total number of virtual network elements that need to be restarted in the network, and the target completion time can be obtained. The target completion time can be the completion time required for restarting the virtual network element, or it can also be the calculated completion time required for restarting the virtual network element, which is at least determined based on the maximum restart rate of the network and the total number of network elements that need to be restarted.
[0059] Step 204, at the target time point, initiating the restart of the virtual network element corresponding to the actual restart number of the target time point, determining the restart success rate of the virtual network element at the target time point.
[0060] Specifically, at the moment of reaching the target time point, the restart of the virtual network element corresponding to the actual restart number of the target time point can be initiated, and the current restart success rate of the virtual network element is calculated during the restart process. For example, at the moment of reaching the initial time point, the virtual network element corresponding to the actual restart number of the initial time point can be determined in each virtual network element to be restarted, and the restart operation of the virtual network element corresponding to the actual restart number of the initial time point is initiated. After determining that the restart operation of the virtual network element corresponding to the actual restart number of the initial time point is initiated, the number of virtual network elements that have restarted successfully is determined, and the restart success rate of the virtual network element at the initial time point is calculated based on the number of virtual network elements that have restarted successfully and the actual restart number of the initial time point.
[0061] Step 206, based on the restart success rate at the target time point and the preset number adjustment strategy, determining the actual restart number of the next target time point at the target time point, and re-executing the step of initiating the restart of the virtual network element corresponding to the restart number of the target time point at the target time point, determining the restart success rate of the virtual network element at the target time point based on the actual restart number of the next target time point, until it is determined that the virtual network element does not meet the fault condition.
[0062] The preset number adjustment strategy can be a virtual network element grouping strategy, or a number adjustment strategy for adjusting the actual restart number from the theoretical restart number, etc.
[0063] Specifically, for each target time point, after determining the restart success rate of the target time point, the actual restart number corresponding to the next target time point of the target time point can be calculated based on the restart success rate of the target time point and the preset number adjustment strategy. And in the case where the virtual network element corresponding to the target time point is determined to have completed the restart, the actual restart number of the next target time point determined is taken as a new target time point, and the step of initiating the restart of the virtual network element corresponding to the restart number of the target time point at the target time point, determining the restart success rate of the virtual network element at the target time point is executed, until it is determined that the virtual network element does not meet the fault condition. The fault condition can be that each virtual network element in the current network is reconnected to the backup local device, or the current network fault is recovered, each virtual network element is connected to the original local device, or it can also be that each virtual network element has restarted successfully, and the communication connection between the virtual network element and the local device has been suggested, etc.
[0064] In one example, in a case that the virtual network elements in the network meet the fault condition, the actual restart number in a case that the target time point is the initial time point can be determined based on the maximum restart rate of the network corresponding, the total number of network elements to be restarted, and the target completion time length, and the restart of the virtual network elements corresponding to the actual restart number of the initial time point is initiated at the initial time point, it is judged whether the restart of each virtual network element is initiated, if the restart of each virtual network element is initiated, the number of virtual network elements restarted successfully is obtained, and the restart success rate corresponding to the initial time point is calculated based on the number of virtual network elements restarted successfully.
[0065] If the restart success rate is less than or equal to the preset restart success rate threshold, the restart of the virtual network elements corresponding to the actual restart number of the initial time point can be re-initiated, the restart success rate corresponding to the updated initial time point is obtained, if the restart success rate corresponding to the updated initial time point is greater than the preset restart success rate threshold, the actual restart number of the next target time point of the initial time point is obtained based on the actual restart number of the initial time point and the preset number adjustment strategy, the next target time point is taken as a new target time point, the step of determining the actual restart number of the next target time point of the target time point based on the restart success rate of the target time point and the preset number adjustment strategy is executed, until the restart of each virtual network element in the network is completed.
[0066] At the target time point, the restart of the virtual network elements (group B) corresponding to the actual restart number of the target time point is initiated, it is judged whether the restart of each virtual network element in the group B is initiated, if the restart of each virtual network element in the group B is initiated, the number of virtual network elements restarted successfully is obtained, and the restart success rate corresponding to the target time point is calculated based on the number of virtual network elements restarted successfully.
[0067] If the restart success rate is less than or equal to the preset restart success rate threshold, the restart of the virtual network elements corresponding to the actual restart number of the target time point in the group B can be re-initiated, the restart success rate corresponding to the updated target time point is obtained, if the restart success rate corresponding to the updated target time point is greater than the preset restart success rate threshold, the actual restart number of the next target time point of the target time point is obtained based on the actual restart number of the target time point and the preset number adjustment strategy, the next target time point is taken as a new target time point, the step of determining the actual restart number of the next target time point of the target time point based on the restart success rate of the target time point and the preset number adjustment strategy is executed, until the restart of each virtual network element in the network is completed.
[0068] The fault recovery method can avoid restart storm of the virtual network element after the local device in the network fails, can make each virtual network element smoothly re-initiate network connection during fault recovery, avoid user online rate shock, and can quickly process the restarted virtual network element by the backup device in the network, realize fast recovery of the fault, control the restart number in the virtual gateway restart process through the preset number adjustment strategy, realize high efficiency of the network reconstruction process, support multiple number adjustment strategies, and can more effectively adapt to multiple network actual application scenarios.
[0069] In one embodiment, the target time point is an initial time point, and the specific processing process of the step of "determining the actual restart number of the target time point if the virtual network element in the network meets the fault condition" can include:
[0070] When each virtual network element in the network is disconnected from the network, and it is determined that the fault condition is met, the theoretical restart number of the initial time point is determined, and the theoretical restart number of the initial time point is taken as the actual restart number of the initial time point.
[0071] The content of the fault condition can be that each virtual network element in the network is disconnected from the local device in the network.
[0072] Specifically, when it is determined that each virtual network element in the network is disconnected from the local device in the network, it can be determined that the fault condition is currently met, so that the theoretical restart number corresponding to the initial time point can be calculated, and the theoretical restart number of the initial time point is determined as the actual restart number of the initial time point.
[0073] Optionally, the specific process of calculating the theoretical restart number of the initial time point can be: processing the initial time point and the target completion time length by a preset trigonometric function to obtain a first adjustment parameter, and obtaining a second adjustment parameter based on the total number of network elements to be restarted and the target completion time length, and determining the sum of the first adjustment parameter and the second adjustment parameter as the theoretical restart number corresponding to the initial time point.
[0074] Specifically, the initial time point and the target completion time length are calculated by a preset trigonometric function, and the obtained trigonometric function value is determined as the first adjustment parameter. And calculate the quotient value between the total number of network elements to be restarted and the target completion time length, and determine the quotient value as the second adjustment parameter. In this way, the sum of the first adjustment parameter and the second adjustment parameter can be calculated, and the sum value is determined as the theoretical restart number corresponding to the initial time point.
[0075] In one example, based on the target completion time T and the first target value, a quotient value between the target completion time T and the first target value is calculated, and a sum value between the quotient value and the initial time point t is calculated. The sum value is processed by a preset trigonometric function to obtain a trigonometric function value. Optionally, the preset trigonometric function can be tan -1 () and the first target value can be 2.
[0076] In this embodiment, the restart of the virtual network element in the fault network can be initiated in time through the fault condition, the flexibility of the virtual network element control is improved, and the stability of the user online rate and the network communication is ensured.
[0077] In one example embodiment, as shown in Figure 3 the fault recovery method further includes:
[0078] Step 302, based on the maximum restart rate corresponding to the network and the total number of network elements to be restarted, the initial completion time is calculated.
[0079] Specifically, the constraint condition of the maximum restart rate of the network for the virtual network element and the total number of virtual network elements to be restarted in the current network are obtained, and the initial completion time is calculated. The maximum restart rate corresponding to the network can be the maximum peak rate of the network configured in advance or determined based on the actual situation of the network. Correspondingly, the total number of network elements to be restarted can be the total number of virtual network elements disconnected from the network in the current network. Optionally, the quotient value between the total number of network elements to be restarted and the maximum restart rate corresponding to the network can be calculated, and the quotient value is determined as the initial completion time.
[0080] Step 304, based on the initial completion time and the preset buffer time, the target completion time is obtained.
[0081] The preset buffer time can be a preconfigured buffer delay, and the specific value of the preset buffer time can be 1s, 2s, etc. The specific value of the preset buffer time is not limited in the present disclosure. The preset buffer time is a value greater than 0.
[0082] Specifically, the sum value between the initial completion time and the preset buffer time is calculated, and the sum value is determined as the target completion time.
[0083] Step 306, the target time point and the target completion time are processed by a preset trigonometric function to obtain a first adjustment parameter, and based on the total number of network elements to be restarted and the target completion time, a second adjustment parameter is obtained. The sum value of the first adjustment parameter and the second adjustment parameter is determined as the theoretical restart number corresponding to the target time point.
[0084] Specifically, trigonometric function calculations are performed on the target time point and the target completion time based on preset trigonometric functions, and the resulting trigonometric function values are determined as the first adjustment parameter. The quotient between the total number of network elements to be restarted and the target completion time is calculated, and this quotient is determined as the second adjustment parameter. Thus, the sum of the first and second adjustment parameters is calculated, and this sum is determined as the theoretical number of restarts corresponding to the target time point.
[0085] In one example, based on the target completion time T and a first target value, the quotient between the target completion time T and the first target value is calculated, and the sum of this quotient and the target time point t is calculated. This sum is then processed using a preset trigonometric function to obtain a trigonometric function value. Optionally, the preset trigonometric function could be tan... -1 ( ), the first target value can be 2, and the theoretical number of restarts V corresponding to the target time point can be calculated using the following formula:
[0086]
[0087] Where S is the total number of network elements that need to be restarted.
[0088] In this embodiment, the target completion time that meets network requirements is determined by a preset buffer time, and processed by a preset trigonometric function. This ensures that the theoretical number of restarts corresponding to each target time point is controllable, avoids restart storms of virtual network elements, and further improves the stability of the network during fault recovery.
[0089] In one embodiment, such as Figure 4 As shown, the specific processing steps for the step "determining the actual number of restarts at the next target time point based on the restart success rate at the target time point and the preset number adjustment strategy" may include:
[0090] Step 402: Calculate the initial adjustment parameters based on the preset number adjustment strategy and the failure rate corresponding to the restart success rate at the target time point.
[0091] The preset number adjustment strategy can be an adjustment parameter algorithm or an adjustment factor algorithm.
[0092] Specifically, for each target time point, the restart success rate corresponding to that target time point can be obtained, and the failure rate corresponding to that target time point can be obtained by the difference between the second target value and the restart success rate. Optionally, the second target value can be 1.
[0093] Thus, the failure rate is processed by the preset adjustment parameter algorithm to obtain an output result, and the output result is determined as the initial adjustment parameter. For example, the failure rate P of the target time point can be input into the preset adjustment parameter algorithm k() function to obtain the output result k(P) of the adjustment parameter algorithm, and the output result is determined as the initial adjustment parameter. Alternatively, the preset adjustment parameter algorithm can be a k() function, and the disclosure does not limit the specific type of the preset adjustment parameter algorithm. The k() function can be a linear function or a nonlinear function.
[0094] At step 404, based on the relationship between the initial adjustment parameter, the failure rate of the target time point, and the preset failure rate threshold, the target adjustment parameter is determined, and the theoretical restart number corresponding to the next target time point of the target time point is adjusted based on the target adjustment parameter to obtain the actual restart number of the next target time point.
[0095] The preset failure rate threshold can be determined based on the actual application scenario, preconfigured, or determined based on the performance of the current faulty network. For example, it can be five percent or the like.
[0096] Specifically, the target time point is compared with the preset failure rate threshold to obtain a comparison result, and the target adjustment parameter is calculated based on the comparison result, the first target value, and the initial adjustment parameter. In this way, the theoretical restart number of the next target time point can be calculated by the method described in the above embodiments, and the theoretical restart number of the next target time point is adjusted based on the target adjustment parameter to obtain the actual restart number of the next target time point. For example, the product value between the target adjustment parameter and the theoretical restart number of the next target time point can be calculated, and the calculated product value is determined as the actual restart number corresponding to the next target time point.
[0097] In this embodiment, the preset number adjustment strategy can effectively control the online number of virtual network elements at different time points, ensure that the number of restarted virtual network elements at each time point is controllable, and the network reconstruction process is also controllable, thereby achieving efficient restart of the virtual network element.
[0098] In one exemplary embodiment, the specific implementation process of the step of "determining the target adjustment parameter based on the relationship between the initial adjustment parameter, the failure rate of the target time point, and the preset failure rate threshold" can include:
[0099] If the failure rate of the target time point is greater than or equal to the preset failure rate threshold, the difference between the target value and the initial adjustment parameter is determined as the target adjustment parameter. Alternatively,
[0100] If the failure rate of the target time point is less than the preset failure rate threshold, a sum value between the target value and the initial adjustment parameter is determined as the target adjustment parameter.
[0101] Specifically, the target adjustment parameter can be obtained through a comparison result between the failure rate P of the target time point and the preset failure rate threshold p. For example, if the failure rate of the target time point is greater than or equal to the preset failure rate threshold, a difference value between the target value and the initial adjustment parameter k(P) can be calculated, and the difference value is determined as the target adjustment parameter. If the failure rate of the target time point is less than the preset failure rate threshold, a sum value between the initial adjustment parameter and the target value can be calculated, and the sum value is determined as the target adjustment parameter.
[0102] In one example, the actual restart number v corresponding to the next target time point can be determined, for example, by the following formula:
[0103]
[0104] Wherein, the target value can be 1, and v can be the theoretical restart number of the next target time point.
[0105] In the embodiment, the number adjustment strategy can effectively control the online number of the virtual network element at different time points, ensure that the restart number of the virtual network element at each time point is controllable, the network reconstruction process is also controllable, and efficient restart of the virtual network element is realized.
[0106] In one example embodiment, the specific processing process of the step of "determining the actual restart number of the next target time point of the target time point based on the restart success rate of the target time point and the preset number adjustment strategy" includes:
[0107] Based on the failure rate corresponding to the restart success rate of the target time point, the restart number matching the failure rate is queried in the preset number adjustment strategy, and the restart number matching the failure rate is determined as the actual restart number of the next target time point of the target time point.
[0108] Wherein, the preset number adjustment strategy can be a grouping matching strategy, and the virtual network elements to be restarted in the corresponding network can be divided into virtual network element groups containing different numbers of virtual network elements. The preset number adjustment strategy can include the restart number corresponding to different restart failure rates, that is, the corresponding relationship between the restart failure rate and the restart number.
[0109] Specifically, for each target time point, the restart success rate corresponding to the target time point can be obtained, and the failure rate corresponding to the target time point can be obtained through the difference between the second target value and the restart success rate. Optionally, the second target value can be 1. In the plurality of corresponding relationships contained in the preset number adjustment strategy, the restart number matching the failure rate is queried, and the restart number is determined as the actual restart number of the next time point of the current target time point.
[0110] In the embodiment, the online number of the virtual network element at different time points can be effectively controlled through the preset number adjustment strategy, so that the number of restarted virtual network elements at each time point is controllable, the network reconstruction process is also controllable, and efficient restart of the virtual network element is realized.
[0111] In the following, the specific implementation process of the above fault recovery method will be described in detail in combination with one specific embodiment, which can include:
[0112] The metropolitan area network can include a plurality of physical devices and a plurality of local devices, for example, can include physical device 1, physical device 2, physical device 3, each physical device can be connected with local device 1, when the local device 1 fails, the physical device under the local device 1 will be disconnected, in the traditional technology, all physical devices will initiate restart or reestablish connection after discovering that they are disconnected, the local device 2 can be a backup of the local device 1, the local device 2 cannot process a large number of device connection requests in time, and the reconnection of the physical device will time out. Since there is no coordination mechanism between the physical devices, the time backoff algorithm will cause the user online rate to fluctuate, the average rate is lower than the device performance limit, and the backup device of the local device is also prone to failure.
[0113] The fault recovery method provided in the embodiment is a method applied in the above-mentioned scenario, for example, can be a scenario of a large number of virtual network element devices restarting and establishing external connections during fault recovery, which realizes the scheduling of the virtual network element restart after failure, avoids the secondary failure problem caused by the virtual network element restart storm, and makes the network more stable during fault recovery. Figure 5 As shown in the figure, it can be a fault diagram of a virtual network element in a network:
[0114] The metropolitan area network includes a server cluster and a plurality of local devices, the server cluster can include virtual network element 1, virtual network element 2, and virtual network element 3, and the plurality of local devices can include local device 1 and local device 2. When the local device 1 fails, the virtual network element 1, the virtual network element 2, and the virtual network element 3 need to be restarted and connected with the local device 2.
[0115] In one example, the local terminal device 1 fails, and the virtual network elements 1, 2, and 3 will be disconnected. Under the unified control of the cluster, a small number of virtual network elements will attempt to redial, and most of the virtual network elements will remain silent. When a small number of virtual network elements dial successfully, the cluster will add more virtual network elements to initiate restart connection. As shown in Figure 6 , the online rate of the virtual network elements (the number of users online per second) will soon reach the online rate limit of the local terminal device (the online rate limit of the local terminal device), and can be stabilized at the highest online rate. The backup local terminal device 2 can quickly process the restarted virtual network elements to achieve rapid recovery from failure, while the number of users online is controllable and will not produce shock, and the backup device is not prone to new failures.
[0116] In one example, based on the target completion time T and the first target value, the quotient value between the target completion time T and the first target value is calculated, and the sum value between the quotient value and the target time point t is calculated. The sum value is processed by a preset trigonometric function to obtain a trigonometric function value. Optionally, the preset trigonometric function can be tan -1 , the first target value can be 2, and the theoretical restart number V corresponding to the target time point can be calculated by the following formula:
[0117]
[0118] , the relationship between each target time point and the theoretical restart number (theoretical calculation value) corresponding to each target time point can be as shown in Figure 7 , wherein U represents the maximum restart peak rate that the network can withstand.
[0119] In one example, the actual restart number v corresponding to the next target time point can be determined, for example, by the following formula:
[0120]
[0121] , the target value can be 1, and V can be the theoretical restart number of the next target time point. The relationship between each target time point and the actual restart number (actual value) corresponding to each target time point can be as shown in Figure 8 .
[0122] In one example, as shown in Figure 9 , the preset number adjustment strategy can be a group matching strategy, and the virtual network elements to be restarted in the corresponding network can be divided into virtual network element groups containing different numbers of virtual network elements. The preset number adjustment strategy can include the restart number corresponding to different restart failure rates, i.e., the corresponding relationship between the restart failure rate and the restart number. Each group can be group A, group B, group C, and the like.
[0123] For example, the number of virtual network elements included in group A can be 3, the number of virtual network elements included in group B can be 9, and the number of virtual network elements included in group C can be N. The server cluster fault recovery control plane can control the restart rate based on the recovery strategy, restart the virtual network elements corresponding to the groups, obtain a restart result feedback, and use the feedback result to control the recovery strategy again.
[0124] As shown in FIG. 8, it can be a specific execution process of the fault recovery method in an embodiment: Figure 10 When the network fails, the cluster control discovers that almost all the virtual network elements are disconnected.
[0125] According to the strategy, the restart of group A is initiated. The number of group A is generally small.
[0126] It is determined whether the dialing of the group is completed, that is, whether the restart of group A is all initiated. The initiation of dialing can be the restart of the virtual network element, including dialing, sending a network connection request, sending a message, and the like. If it is not completed, it is returned to re-determine whether the dialing is completed.
[0127] If it is completed, it is determined whether the dialing success rate of the group is greater than a threshold, that is, whether the current restart success rate is greater than or equal to a preset success rate threshold. If it is less than the preset success rate threshold, it is returned to re-execute the step of initiating the restart of group A according to the strategy.
[0128] If it is greater than or equal to, it is determined whether all the redials are completed, that is, whether all the virtual network elements to be restarted are restored to the connection with the local device. If it is not restored, the number of restarts in the next period is increased according to the strategy, for example, group B or C can be used for dialing, and it is returned to re-execute the step of determining whether the dialing is completed. If all the virtual network elements to be restarted are restored to the connection with the local device, the fault recovery is completed.
[0129] The fault recovery method in the embodiment can realize the re-online of the virtual network element device, avoid the secondary fault recovery failure caused by the connection storm during the burst restart, avoid the network shock during the burst dialing of a large number of users, and adjust the strategy by a preset number, so that the re-establishment process is controlled and an efficient re-establishment mechanism is realized through the strategy. Optionally, different strategy selection and parameter configuration are supported, which can more effectively adapt to different network states and expand the applicable application scenarios.
[0130]
[0131] It should be understood that although each step in the flowchart involved in the embodiments described above is shown in sequence according to the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in the embodiments described above can include multiple steps or stages, which are not necessarily executed at the same time point, but can be executed at different time points, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or steps or stages in other steps.
[0132] Based on the same inventive concept, the embodiments of the present application also provide a fault recovery device for implementing the fault recovery method described above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more fault recovery device embodiments provided below can refer to the limitations of the fault recovery method described above, which will not be repeated here.
[0133] In one exemplary embodiment, as shown in Figure 11 A fault recovery device 1100 is provided, comprising:
[0134] A first determination module 1102 is configured to determine the actual number of restarts at a target time point if the virtual network element in the network meets the fault condition, and the number of restarts corresponding to the target time point at the initial time point is determined based on the corresponding maximum restart rate of the network, the total number of network elements that need to be restarted, and the target completion time length.
[0135] An initiation module 1104 is configured to initiate the restart of the virtual network element corresponding to the actual number of restarts at the target time point at the target time point, and determine the restart success rate of the virtual network element at the target time point.
[0136] A second determination module 1106 is configured to determine the actual number of restarts at the next target time point based on the restart success rate at the target time point and the preset number adjustment strategy, and re-execute the steps of initiating the restart of the virtual network element corresponding to the number of restarts at the target time point at the target time point and determining the restart success rate of the virtual network element at the target time point based on the actual number of restarts at the next target time point until it is determined that the virtual network element does not meet the fault condition.
[0137] In one embodiment, the target time point is the initial time point, and the first determination module is specifically configured to:
[0138] In a case where it is determined that each virtual network element in the network disconnects from the network and a fault condition is met, a theoretical restart number of the initial time point is determined, and the theoretical restart number of the initial time point is taken as an actual restart number of the initial time point.
[0139] In one of the embodiments, the apparatus further comprises:
[0140] The first calculation module is configured to calculate an initial completion duration based on a maximum restart rate corresponding to the network and a total number of network elements that need to be restarted.
[0141] The second calculation module is configured to obtain a target completion duration based on the initial completion duration and a preset buffer time.
[0142] The third calculation module is configured to obtain a first adjustment parameter by processing the target time point and the target completion duration through a preset trigonometric function, obtain a second adjustment parameter based on the total number of network elements that need to be restarted and the target completion duration, and determine a sum of the first adjustment parameter and the second adjustment parameter as a theoretical restart number corresponding to the target time point.
[0143] In one of the embodiments, the second determination module is specifically configured to:
[0144] Calculate an initial adjustment parameter based on a preset number adjustment strategy and a failure rate corresponding to a success rate of the target time point.
[0145] Determine a target adjustment parameter based on a relationship between the initial adjustment parameter, the failure rate of the target time point and a preset failure rate threshold, and adjust a theoretical restart number corresponding to a next target time point of the target time point based on the target adjustment parameter to obtain an actual restart number of the next target time point.
[0146] In one of the embodiments, the second determination module is specifically configured to:
[0147] If the failure rate of the target time point is greater than or equal to the preset failure rate threshold, a difference between a target value and the initial adjustment parameter is determined as the target adjustment parameter. Or,
[0148] If the failure rate of the target time point is less than the preset failure rate threshold, a sum between the target value and the initial adjustment parameter is determined as the target adjustment parameter.
[0149] In one of the embodiments, the second determination module is specifically configured to:
[0150] Based on the failure rate corresponding to the target time point restart success rate, in the preset number of adjustment strategies, the number of restarts matched with the failure rate is queried, and the number of restarts matched with the failure rate is determined as the actual number of restarts at the next time point of the target time point.
[0151] Each module in the above fault recovery device can be implemented wholly or partially by software, hardware, and a combination thereof. The above modules can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the above modules.
[0152] In an exemplary embodiment, a computer device, which can be a server, is provided, and an internal structure diagram thereof can be as shown in Figure 12 The computer device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store data of a virtual network element. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through a network connection. The computer program is executed by the processor to implement a fault recovery method.
[0153] Those skilled in the art can understand that Figure 12 The structure shown in the above embodiment is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. Specifically, the computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0154] In an exemplary embodiment, a computer device is provided, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the above embodiments.
[0155] In an embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the steps of the above embodiments.
[0156] In one embodiment, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the steps of the above embodiments.
[0157] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.
[0158] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In the embodiments provided by the present application, any reference to the memory, database or other medium can include at least one of the non-volatile memory and the volatile memory. The non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (Resistive Random Access Memory, ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. The volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided by the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided by the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (Artificial Intelligence, AI) processor, etc., without being limited thereto.
[0159] The technical features of the above embodiments can be combined in any manner. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not contradict each other, they should be considered to be within the scope of the present application.
[0160] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the patent scope of the present application. It should be pointed out that, for ordinary skilled persons in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method of failure recovery, characterized by, The method comprises: if the virtual network element in the network meets the fault condition, determining the actual restart number of the target time point, the corresponding restart number of the initial time point being determined based on the maximum restart rate corresponding to the network, the total number of network elements to be restarted, and the target completion duration; at the target time point, initiating the restart of the virtual network element corresponding to the actual restart number of the target time point, determining the restart success rate of the virtual network element at the target time point; based on the restart success rate at the target time point and a preset number adjustment strategy, determining the actual restart number of the next target time point of the target time point, and based on the actual restart number of the next target time point, re-executing the steps of initiating the restart of the virtual network element corresponding to the restart number of the target time point at the target time point, and determining the restart success rate of the virtual network element at the target time point, until it is determined that the virtual network element does not meet the fault condition.
2. The method of claim 1, wherein, The target time point is the initial time point, and the step of determining the actual restart number of the target time point if the virtual network element in the network meets the fault condition comprises: in a case where it is determined that each virtual network element in the network is disconnected from the network and meets the fault condition, determining the theoretical restart number of the initial time point, and taking the theoretical restart number of the initial time point as the actual restart number of the initial time point.
3. The method of claim 1, wherein, The method further comprises: based on the maximum restart rate corresponding to the network and the total number of network elements to be restarted, calculating an initial completion duration; based on the initial completion duration and a preset buffer time, obtaining a target completion duration; processing the target time point and the target completion duration by a preset trigonometric function to obtain a first adjustment parameter, and based on the total number of network elements to be restarted and the target completion duration, obtaining a second adjustment parameter, and taking the sum of the first adjustment parameter and the second adjustment parameter as the theoretical restart number corresponding to the target time point.
4. The method of claim 2, wherein, The step of determining the actual restart number of the next target time point of the target time point based on the restart success rate at the target time point and a preset number adjustment strategy comprises: based on the preset number adjustment strategy and the failure rate corresponding to the restart success rate at the target time point, calculating an initial adjustment parameter; based on the relationship between the initial adjustment parameter, the failure rate at the target time point, and a preset failure rate threshold, determining a target adjustment parameter, and based on the target adjustment parameter, adjusting the theoretical restart number corresponding to the next target time point of the target time point to obtain the actual restart number of the next target time point.
5. The method of claim 4, wherein, The step of determining a target adjustment parameter based on the relationship between the initial adjustment parameter, the failure rate at the target time point, and a preset failure rate threshold comprises: if the failure rate at the target time point is greater than or equal to the preset failure rate threshold, taking the difference between a target value and the initial adjustment parameter as the target adjustment parameter; or If the failure rate of the target time point is less than a preset failure rate threshold, a sum value between a target value and the initial adjustment parameter is determined as a target adjustment parameter.
6. The method of claim 1, wherein, The actual restart number of the next target time point of the target time point is determined based on the restart success rate of the target time point and a preset number adjustment strategy, and the method comprises the steps of: Based on the failure rate corresponding to the restart success rate of the target time point, the restart number matched with the failure rate is queried in the preset number adjustment strategy, and the restart number matched with the failure rate is determined as the actual restart number of the next time point of the target time point.
7. A fault recovery apparatus characterized by comprising: The device comprises: A first determination module is configured to determine an actual restart number of a target time point if a virtual network element in a network meets a fault condition, wherein the actual restart number of the target time point is determined based on a maximum restart rate of the network, a total number of network elements that need to be restarted, and a target completion time length; An initiation module is configured to initiate a restart of a virtual network element corresponding to the actual restart number of the target time point at the target time point, and determine a restart success rate of the virtual network element at the target time point; A second determination module is configured to determine an actual restart number of a next target time point of the target time point based on the restart success rate of the target time point and a preset number adjustment strategy, and re-execute the steps of initiating the restart of the virtual network element corresponding to the restart number of the target time point at the target time point and determining the restart success rate of the virtual network element at the target time point based on the actual restart number of the next target time point until it is determined that the virtual network element does not meet the fault condition. 8.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-7. The processor executes the computer program to implement the steps of the method of any one of claims 1 to 6.
9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Automatic testing method and master control device
CN104216823A
Performing reboot cycles, a reboot schedule on on-demand rebooting
CN104937546A