Method and system for guaranteeing stability of management and control plane of wide-area cloud network

By weighting and classifying the wide area cloud network and automatically adjusting the inspection priority, combined with the fault handling mechanism, the problems of large inspection workload, high resource consumption and poor timeliness in the existing technology have been solved. Timely inspection of important business and rapid fault handling have been achieved, improving system stability and resource utilization efficiency.

WO2025227697A1PCT designated stage Publication Date: 2025-11-06CHINA TELECOM CLOUD TECH CO LTD

Patent Information

Application Number
PCT/CN2024/135840
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-30
Filing Date
2024-11-29
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Existing technologies in wide-area cloud network management and control suffer from problems such as large workload for inspection, high resource consumption, poor timeliness, strong monitoring lag, inability to provide real-time feedback on network stability, and inability of inspection systems to resolve issues in a timely manner, which may lead to greater failures in critical systems.

Method used

By assigning weights to business operations, setting priorities and weights for inspection tasks, and automatically adjusting inspection priorities using monitoring data, combined with a fault type handling mechanism, rapid fault identification and emergency avoidance can be achieved. The inspection process is optimized by employing modules such as an inspection use case platform, a queue of tasks to be issued, a policy setting platform, log collection and monitoring, a fault handling module, and data reconciliation.

Benefits of technology

It enables timely inspection and rapid fault handling of critical business operations, reduces online failures, improves system stability and resource utilization efficiency, and ensures the stability and real-time nature of control and management measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135840_06112025_PF_FP_ABST
    Figure CN2024135840_06112025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention are a method and system for guaranteeing the stability of a management and control plane of a wide-area cloud network. In the method, the priorities of inspection use cases are optimized, the weights and number of use cases in each round of inspection can be set according to the importance degrees of functions, and the priorities of related use cases can be automatically adjusted within a certain range by means of detected data, so that inspection resources can be utilized to the maximum extent. An inspection system dispatches an inspection use case strictly according to the utilization level of each machine under each management and control channel in a monitoring system, and if the utilization level of a related channel is too high, the use case enters a dispatch-pending queue and verification is performed again after a period of time, thereby avoiding affecting dispatch of online normal services. The inspection system focuses on real-time prioritized detection of faults in critical services, and critical services, services that recover from faults, and unstable services are prioritized for inspection and subject to multiple times of inspections under the combined action of weighting, monitoring and adjustments, allowing timely detection, prioritized avoidance and resolution of system issues.
Need to check novelty before this filing date? Find Prior Art

Description

Method and system for ensuring stability of wide-area cloud network management and control surface TECHNICAL FIELD

[0001] The present application relates to the technical field of network communication, and in particular to a method and system for ensuring stability of a wide-area cloud network management and control surface. BACKGROUND

[0002] With the high development of cloud computing technology, more and more businesses choose to manage in the cloud. With the increasing complexity of cloud business and the increasing size of business, the stability of business, the stability of business change, and the stability of network are increasingly required. In order to ensure the stability of the management and control channel, two schemes are usually used. One is an online inspection method, which actively verifies the stability of the channel through a test interface. The other is to infer whether the channel is stable by monitoring and counting various indicators. However, the above two methods have the following defects:

[0003] 1) The workload of online inspection will increase with the development of business. Due to the continuous expansion of the network and the increasing complexity of the business model, the workload of full verification increases exponentially. The increase in the workload of full verification also brings great difficulty to online inspection. If the inspection frequency is too high, it will occupy a lot of management and control channel resources. If the frequency is too low, it cannot reflect the problem in real time.

[0004] 2) The monitoring has a certain lag. If only the errors of the monitoring and issuing requests are used to deduce and locate the problem in the business, the problem has already occurred. If only the state of the monitoring machine and the water level capacity are monitored, many problems in the business cannot be found.

[0005] 3) The stability of the management and control channel has a certain uncertainty and cannot be real-time feedback like the data plane. Many scenarios need to trigger special scenarios online to verify the timeliness of the inspection.

[0006] 4) Many inspection systems only find problems but do not solve them in real time. If the fault occurs, the operation and maintenance personnel may cause greater failure in the middle of the way. Therefore, the inspection system needs to make some preliminary judgments and recoveries to solve the non-system defect problems. SUMMARY

[0007] This section is intended to summarize some aspects of the embodiments of the present application and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the Abstract and Title of the specification to avoid obscuring the summary of the disclosure. Such simplifications or omissions are not intended to limit the scope of the present application.

[0008] Therefore, in order to solve the above technical problems, the present application provides the following technical solutions: a method for ensuring the stability of the management and control surface of a wide-area cloud network, comprising the following specific steps:

[0009] S1: weight classification of the overall service;

[0010] S2: after the start of the inspection task, the task determines whether to issue according to the water level threshold of the management and control channel, and when the water level exceeds the set threshold, the task enters the pending issue queue;

[0011] S3: the priority and weight of the inspection task have two modes of automatic setting and manual setting; manual setting is to set the weight and number of tasks, and automatic setting is to calculate the unstable parameters of the interface and the channel coefficient according to the monitoring and historical data, and the two interact with each other to comprehensively determine the inspection priority of each interface;

[0012] S4: when a fault is detected, it is processed according to the type of the fault;

[0013] S5: after the fault is handled, the monitoring parameter affects the adjustment of the priority, and the use case is preferentially executed with other services of the service, the modification effect is verified in the first time, and the priority of the problem is improved within a certain time.

[0014] As a preferred scheme of the method for ensuring the stability of the management and control surface of a wide-area cloud network, in step S1, the weight classification is divided into multiple dimensions, and the order and number of single service operation during inspection are determined according to the priority and weight of the dimension and the influence of the environment.

[0015] As a preferred scheme of the method for ensuring the stability of the management and control surface of a wide-area cloud network, in step S2, the subsequent tasks are issued in sequence first, and then the tasks in the pending issue queue are forwarded in sequence after the set time.

[0016] As a preferred scheme of the method for ensuring the stability of the management and control surface of a wide-area cloud network, in step S3, the manual setting automatically increases the weight according to the influence of the service; the automatic setting periodically generates an optimized inspection task according to the current monitoring situation, and waits for the relevant personnel to determine whether to update the new task.

[0017] As a preferred scheme of the method for ensuring stability of a wide-area cloud network management and control plane, in step S4, the types of faults include link fault, data abnormality fault, system fault, etc., and the system processes the faults by combining different fault types with the characteristics of the system.

[0018] As a preferred scheme of the method for ensuring stability of a wide-area cloud network management and control plane, in step S5, when a fault occurs, if the fault cannot be solved temporarily, an emergency avoidance method can be used to temporarily change, so as to avoid the situation of online fault, and the specific steps are as follows:

[0019] S51, when a fault occurs, the interface is automatically issued three times to avoid the fault caused by network jitter; if the issuing is successful, the interface is recorded as an unstable interface, and no fault is recorded.

[0020] S52, if the issuing is not successful, it is judged that a fault occurs, and the system judges whether it is a link fault, a data abnormality fault, a system defect, etc.

[0021] S53, the link fault alarm is judged and processed.

[0022] S531, the link fault mainly shows timeout; if timeout occurs, the first priority is retry, and at the same time of retry, the system performs a ping test on the opposite end, judges whether the opposite end service is normal through heartbeat information of the opposite end machine and other multi-dimensional data, if the retry is not successful for three times, the system analyzes and concludes that the fault is caused by the link being not connected, at this time, the system switches to a standby link for issuing, and initiates a link alarm to notify the relevant operation and maintenance personnel to repair.

[0023] S532, if the standby link fails to issue, it means that the regional controller may have a fault, at this time, the standby regional controller corresponding to the regional controller is configured to issue; at the same time, an alarm is actively sent to notify the relevant operation and maintenance personnel to repair.

[0024] S533, if the standby regional controller also fails to issue, it is judged that the control plane network has a large-scale fault or a system defect, at this time, an emergency alarm is sent to the operation and maintenance personnel for key repair.

[0025] S54, data abnormality fault and solution.

[0026] If the interface reports an error, it is first judged whether the fault is caused by data abnormality.

[0027] The above data problem mainly refers to the case that the control plane data passed the check, but did not pass the check of the underlying network element layer, and such problem is mainly caused by inconsistency between upper and lower layers; at this time, the system re-triggers the automatic reconciliation function, and if the reconciliation data is inconsistent, the system sends the reconciliation result to the operation and maintenance personnel to verify whether the data needs to be synchronized; if the data needs to be synchronized, retry after data synchronization to check whether the problem is solved;

[0028] S55, system defect;

[0029] If the interface reports an error, neither the data nor the link is abnormal, and parameter check and service check operations have been performed, at this time, the system defect is judged, an alarm is given, and the relevant operation shift personnel are notified to locate and solve the problem.

[0030] The application also provides the above-mentioned system for stabilizing the management and control plane of a wide-area cloud network, characterized in that the system comprises an inspection case platform, a to-be-released queue, a policy setting platform, a log collection and monitoring, an inspection policy configuration module, a fault processing module and data reconciliation;

[0031] The inspection case platform is used for managing inspection cases, and users manage the inspection cases of each task through the inspection case platform and modify and add relevant cases to adapt to the changes of the platform.

[0032] The to-be-released queue is used for entering the inspection cases into the to-be-released queue when the management and control release channel water level is too high and is always in a busy state, so that the system re-processes the tasks in the to-be-released queue after a period of time;

[0033] The policy setting platform is used for users to set the relevant business priority through the platform, thereby affecting the setting of the entire inspection case platform.

[0034] The log collection and monitoring is mainly used for collecting the configurations of the global controller and the regional controller for each service release, thereby affecting the entire release task.

[0035] The inspection policy configuration module actively sets through the policy setting platform or calculates the optimized scheme according to the current situation through the monitoring system.

[0036] The fault processing module is mainly used for processing problems such as line disconnection and service failure when problems are found in the inspection by immediately starting the fault processing module.

[0037] The data reconciliation is mainly used for processing the influence caused by data inconsistency, dirty data or data loss.

[0038] As a preferred scheme of the system for ensuring the stability of the management and control plane of the wide-area cloud network, in the strategy setting platform, the user sets the weight, the number of times and other factors to calculate the priority of the use case, and the formula for calculating the score of the inspection use case is ((business weight / number of times)+channel coefficient)*unstable coefficient, and the higher the score calculated by the formula, the earlier the inspection use case is executed.

[0039] As a preferred scheme of the system for ensuring the stability of the management and control plane of the wide-area cloud network, in the strategy setting platform, the business weight and the number of times refer to the division of the business into core business, key business, general business and test business.

[0040] Among them, the core business refers to the business covering more than 60% of the customers and the business of more than 90% of the VIP customers, the core business weight is 1000, and the number of times is 10; the key business refers to the further business coverage of the core business, including more than 80% of the business and all VIP business, the key business weight is 400, and the number of times is 5; the general business includes all online businesses in use, the weight is 120, and the number of times is 2; and the test business refers to new business in the test process, the weight is 50, and the number of times is 1.

[0041] The channel coefficient is the importance of the link, which is determined according to the amount of business of the link, the coefficient of a particularly important link is 100, the coefficient of a generally important link is 50, and the coefficient of a link with a small number of uses is 10.

[0042] The unstable coefficient is determined by the monitoring key item, the number of failures and the average number of failure-free times, the initial value of the monitoring key item alarm is 1 and cannot be less than 1, the coefficient increases by 0.01, the coefficient increases by 0.02 for each request failure, and the coefficient decreases by 0.1 for each 6 hours of failure-free time.

[0043] As a preferred scheme of the system for ensuring the stability of the management and control plane of the wide-area cloud network, the inspection strategy configuration module is the core module of the entire system, and the optimized scheme needs to be confirmed by the operation and maintenance personnel before being switched.

[0044] The beneficial effects of the present application are:

[0045] 1. The priority of the inspection use case is optimized, the weight and the number of times of the use case in each round of inspection can be set according to the importance of the function, and the priority of the related use case can be automatically adjusted within a certain range through the monitoring data, so that the inspection resources can be operated in the largest range.

[0046] 2、The inspection system in the scheme issues inspection cases, which will strictly follow the water level of each machine in the monitoring system about each control channel to execute, if the relevant channel water level is too high, the case enters the pending queue, and rechecks after a period of time to avoid affecting the normal business of the online issue.

[0047] 3、The inspection system in the scheme focuses on the real-time priority discovery of important business faults, and through the joint action of weighting and monitoring modification, important business, fault recovery business and unstable business are given priority to inspection and multiple inspection, so that system problems can be discovered in time and avoided and solved in priority. BRIEF DESCRIPTION OF DRAWINGS

[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. Among them:

[0049] Fig. 1 is a flowchart of the cloud network control surface issue in the prior art.

[0050] Fig. 2 is a system architecture diagram of the present application.

[0051] Fig. 3 is a system flowchart of the present application.

[0052] Fig. 4 is a fault processing flowchart of the present application. DETAILED DESCRIPTION

[0053] In order to make the above-mentioned purposes, features and advantages of the present application more apparent and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification.

[0054] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present application, therefore the present application is not limited to the specific embodiments disclosed below.

[0055] Secondly, the "one embodiment" or "embodiment" referred to herein means that the specific features, structures or characteristics can be included in at least one implementation of the present application. In this specification, "in one embodiment" does not mean the same embodiment, nor is it an independent or selective embodiment that excludes other embodiments.

[0056] The scheme focuses on solving the problem of cloud network management and control surface stability. Due to the increasing complexity of the network and the sharp increase in business capacity, the workload and difficulty of online inspection are increasing. If frequent inspection is performed, it will occupy a large amount of management and control channel resources, thereby affecting the management and control surface issuing. At the same time, if the inspection frequency is too low, the effectiveness of the inspection will be reduced. Therefore, it is very important to design a reasonable inspection method.

[0057] The role of inspection is to actively discover problems in the system. In a large system, it is a valuable problem to use limited resources to discover problems in time. Discovering problems before customers can effectively reduce network faults and improve system stability. After relevant problems are found by inspection, a quick solution is needed to quickly solve the problem in a targeted manner. Only in this way can the fault be reduced. Inspection cannot affect online business. The system also needs to strictly design related mechanisms so that the inspection task does not affect the normal business issuing.

[0058] Referring to FIGS. 2-4, an embodiment of the present application provides a system for ensuring the stability of the wide-area cloud network management and control surface to solve the above technical problems, as shown in FIG. 2, and the specific implementation is as follows:

[0059] Inspection case platform: This platform is a platform for managing inspection cases. Users can manage the inspection cases of each task here and modify and add related cases to adapt to the changes of the platform.

[0060] Pending queue: Because the management and control issuing channel is always busy due to high water level, the inspection case cannot be executed in time. Therefore, it is first entered into the pending queue. The system will reprocess the tasks in the pending queue at intervals.

[0061] Strategy setting platform: Users can set the priority of related business through the platform, thereby affecting the setting of the entire inspection case platform.

[0062] Log collection and monitoring: It mainly collects the configuration of the global controller and regional controller for each business issuing, thereby affecting the entire issuing task.

[0063] Inspection strategy configuration module: It is the core module of the entire system. It can be actively set through the strategy setting platform or calculated through the monitoring system according to the current situation to obtain an optimized solution. However, the optimized solution needs to be confirmed by the operation and maintenance personnel before it can be switched.

[0064] Fault handling: If the inspection finds a problem, the fault handling module will be started immediately to timely handle problems such as line disconnection and business failure.

[0065] Data reconciliation: mainly deal with the impact of data inconsistency, or dirty data or data missing.

[0066] The operation of the system is illustrated by an embodiment as follows:

[0067] As shown in FIG. 1, the prior art initiates a task through a cloud platform, and then issues it to a global controller, which finds corresponding regional controllers through network element devices to which the configuration needs to be issued, and then issues the configuration to the corresponding network elements through the regional controllers.

[0068] First, the business is set, and the user calculates the priority of the use case by setting the weight and number of times and other factors. The formula for calculating the inspection case score is ((business weight / number of times)+channel coefficient)*unstable coefficient.

[0069] The higher the score calculated by the formula, the earlier the inspection case is executed.

[0070] The details of the scheme are introduced as follows:

[0071] 1. Business weight and number of times: the business is divided into core business, key business, general business, and test business. The core business generally covers more than 60% of customers and more than 90% of VIP customers. The core business weight is 1000, and the number of times is 10. The key business refers to further business coverage in addition to the core business, including more than 80% of the business and all VIP business. The key business weight is 400, and the number of times is 5. The general business includes all online businesses in use, with a weight of 120 and a number of times of 2. The test business refers to new businesses in the test process, with a weight of 50 and a number of times of 1.

[0072] 2. The channel coefficient is the importance of the link, which is determined according to the amount of business of the link. The coefficient of a particularly important link is 100, the coefficient of a generally important link is 50, and the coefficient of a link used less frequently is 10.

[0073] 3. The unstable coefficient is determined by the monitoring focus, the number of failures, and the average number of fault-free times. The initial value is 1 when the monitoring focus alarms once, and cannot be less than 1. The coefficient increases by 0.01 when a request fails once, and decreases by 0.1 when there is no fault for an average of 6 hours.

[0074] The above examples are only to better illustrate the details of the scheme, and the specific parameters can be adjusted according to the actual situation.

[0075] The specific implementation is as follows:

[0076] For example, there are two services, a service and b service, a service is the key service, we can set the weight to 400, and the number of times to 5, while b service is a general service, set the weight to 120, and the number of times to 2, assuming that the interface stability and channel of the two services are the same, the channel coefficient is 0, and the instability coefficient is 1, then the score of a service for 5 times is 400, 320, 240, 160, 80, and b is 120, 60, so we can perform a four times, then perform task b once, then perform a once, and finally perform b once, through this setting, we can make the important service execute first, and run multiple times in a period to ensure;

[0077] The channel coefficient refers to the weighting of the more important channel. When the channel is more important, the related weight will increase, and the execution priority of the related function will rise;

[0078] The instability coefficient is real-time feedback according to the system situation, the administrator can make the system automatically optimize according to the network situation in real time, and make the verification of the patrol case more inclined to the interface with unstable interface and poor channel coefficient, so as to find more problems.

[0079] For example, service a has two downlink links, which are downlinked to a1 and b1 network elements respectively, but according to the monitoring, the path to b1 network element is worse, so the instability of b1 network element will be higher, and the priority of b1 network element will be higher, so that the problem can be found earlier, and the problem can be found earlier. At the same time, with the improvement of the link, the fault-free time increases, and the instability coefficient will automatically decrease, so that the link returns to the normal predetermined priority;

[0080] After setting the priority of the task downlink, it is the process of downlink, but the downlink of patrol cannot affect the downlink of normal service, so the load level of related downlink channel will be checked before each patrol task downlink, when the related load level exceeds a certain threshold, the task will enter the pending downlink queue, and the judgment will be made after a period of time, if the task has not been downlinked for a certain number of times, it will also alarm, and the problem may exist when the load level of the control channel is always high;

[0081] When the task is downlinked, it can be found by the global controller that needs to be downlinked to the regional controller of the network element, and then downlinked to the specific network element by the regional controller, at the same time, the log and monitoring system will record the success rate, performance and other parameters of each downlink task, form related reports for output, and for important tasks and interfaces, related alarm information will also be output to facilitate timely problem discovery;

[0082] At the same time, the test information will also be combined with the data on the line to evaluate each interface and related control channels, recalculate the channel parameters and interface instability parameters, and thus affect the test priority of the related interface through the above formula;

[0083] The task is divided into manual setting and automatic setting. Manual setting is to set the weight and number of tasks, and automatic setting is to calculate the instability parameters of the interface according to the monitoring and historical data. The two interact with each other to determine the inspection priority of each interface.

[0084] As shown in Figure 4, when the inspection fails, if it cannot be solved at the moment, it can also be temporarily changed through emergency avoidance means to avoid the situation of online failure:

[0085] 1. Since there are relevant parameter verification and business verification in the front end, the request that actually reaches the business flow issuing process is considered to be a legal request. Therefore, the principle of the system is to issue the request as much as possible, and the system is generally fully verified in the online process, so it is considered that the real system defects are very few. The system focuses on identifying and automatically avoiding non-system defect problems to ensure normal business issuance of the system.

[0086] 2. The solution is as follows: When a failure occurs, the interface will automatically issue three times to avoid failures caused by network jitter. If there is success, it can be recorded as an unstable interface, and the failure is not recorded.

[0087] 3. When the system actually fails, the system will judge whether it is a link failure, data anomaly failure, system defect, etc. according to the error report.

[0088] 4. Link failure alarm judgment and processing:

[0089] 1) Link failure mainly shows timeout. If timeout occurs, the first priority is retry. While retrying, the system will perform ping on the opposite end, and through heartbeat information, the opposite end machine monitoring and other dimensions to judge whether the opposite end service is normal. If it fails after 3 retries, and combined with the system's judgment and comprehensive analysis that it is caused by link failure, switch to the backup link for issuance, and actively initiate link alarm to notify the relevant operation and maintenance personnel for repair;

[0090] 2) Another situation is that if the backup link does not work, it means that the regional controller may have a failure. At this time, configuration and issuance can be performed through the backup regional controller corresponding to the regional controller, and at the same time, actively send an alarm to notify the relevant operation and maintenance personnel for repair;

[0091] 3) If the standby regional controller also fails, it is likely that the control plane network has a large-scale failure or a system defect, in which case the operation and maintenance personnel will be urgently alerted to focus on repair.

[0092] 5. Data anomaly failure and solution: If the interface error is caused by data anomaly, the system can determine whether the data anomaly is the cause. The data problem here mainly refers to the data passing the control plane data verification but failing the bottom layer network element layer verification, which is mainly caused by inconsistent data between the upper and lower layers. At this time, the system will trigger automatic reconciliation again. If the reconciliation data is inconsistent, the system will send the reconciliation result to the operation and maintenance personnel to verify whether the data needs to be synchronized. If it is determined, the problem can be solved after synchronizing the data.

[0093] 6. System defect: If the interface error is not caused by data anomaly or link anomaly, and the parameter verification and business verification have been performed, it is likely that there is a real system defect, which will be alarmed to notify the relevant operation and maintenance personnel to locate and solve the problem.

[0094] In summary, the system can set corresponding weight values and times according to the importance of the business. The system can determine which use cases to execute first by comparing the final priority. Different delivery channels for the same use case also vary with the size of the user quantity and the importance of the user priority. Therefore, channel parameters need to be set to affect the priority of the test business delivery. At the same time, monitoring and logging can also change the related results according to the success rate of the channel and interface, so that interfaces and channels with poor stability can be tested first within a certain range to discover related problems in a timely manner.

[0095] At the same time, the system also provides emergency solutions to common problems, timely solves and corrects system errors, solves 100% of non-system defects in the operation process, and also tries to avoid system defects.

[0096] In summary, the system can set corresponding weight values and times according to the importance of the business. The system can determine which use cases to execute first by comparing the final priority. Different delivery channels for the same use case also vary with the size of the user quantity and the importance of the user priority. Therefore, channel parameters need to be set to affect the priority of the test business delivery. At the same time, monitoring and logging can also change the related results according to the success rate of the channel and interface, so that interfaces and channels with poor stability can be tested first within a certain range to discover related problems in a timely manner

[0097] At the same time, the system also provides emergency solutions to common problems, timely solves and corrects system errors, solves 100% of non-system defects in the operation process, and also tries to avoid system defects.

[0098] In the present scheme:

[0099] Wide area cloud network: refers to the network between local devices and remote cloud VPCs, including switches, network access devices, gateways, firewalls, and networks between VPCs and VPCs.

[0100] Control surface: refers to the control channel of the configuration process when the business changes, from the platform side, to the controller, and then to the network element device control and delivery channel.

[0101] Inspection: refers to the use case detection of the online system to find the problems on the line in real time.

[0102] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the present application, which should be covered by the scope of the claims of the present application.

Claims

1. A method for ensuring the stability of a wide-area cloud network control plane, characterized in that: The method comprises the following specific steps: S1: weight classification of the overall business; S2: after the inspection task starts, the task determines whether to issue according to the water level threshold of the control channel, and when the water level exceeds the set threshold, the task enters the pending issuance queue; S3: the priority and weight of the inspection task have two modes of automatic setting and manual setting; manual setting is to set the weight and number of times of the task, and automatic setting is to calculate the unstable parameters and channel coefficients of the interface according to the monitoring and historical data, and the two interact with each other to comprehensively determine the inspection priority of each interface; S4: when the inspection fails, the fault is processed according to the type; S5: after the fault is processed, the monitoring parameter influence weight priority is adjusted, the use case is preferentially executed with other business of the business, the modification effect is verified in the first time, and the priority of the problem is improved within a certain time.

2. The method for ensuring stability of a wide-area cloud network management and control plane as claimed in claim 1, characterized in that: In step S1, the weight classification is divided into multiple dimensions, and the order and number of times of the single business operation during the inspection are determined according to the priority and weight of the dimension and the influence of the environment.

3. The method for ensuring stability of a wide-area cloud network management and control plane as claimed in claim 2, characterized in that: In step S2, the subsequent tasks are issued in sequence first, and then forwarded in the pending issuance queue according to the sequence after the set time.

4. The method for ensuring stability of a wide-area cloud network management and control plane as claimed in claim 3, characterized in that: In step S3, the manual setting automatically increases the weight according to the influence of the business; the automatic setting periodically generates an optimized inspection task according to the current monitoring situation, and waits for the relevant personnel to determine whether to update the new task.

5. The method for ensuring stability of a wide-area cloud network management and control plane as claimed in claim 4, characterized in that: In step S4, the types of faults include link fault, data abnormality fault and system fault, and the system processes different fault types in combination with the characteristics of the system.

6. The method for ensuring stability of a wide-area cloud network management and control plane as claimed in claim 5, characterized in that: In step S5, when the inspection fails, if it cannot be solved at the moment, it can also be temporarily changed through emergency avoidance means to avoid the situation of online failure, and the specific steps are as follows: S51: when a fault occurs, the interface automatically issues three times to avoid the fault caused by network jitter; if the issuance is successful, it is recorded as an unstable interface and the fault is not recorded; S52: if the issuance is not successful, it is judged that a fault occurs, and the system judges whether it is a link fault, a data abnormality fault or a system defect according to the error; S53: judge and handle the link fault alarm; S531: the link fault mainly shows timeout; if timeout occurs, the first priority is retry, and at the same time of retry, the system performs a ping test on the opposite end, monitors multi-dimensional data of the opposite end machine through heartbeat information to judge whether the opposite end service is normal; if the retry is not successful for three times, the system analyzes and concludes that the fault is caused by the disconnection of the link, at this time, the system switches to the standby link for issuance, and initiates a link alarm to notify the relevant operation and maintenance personnel to repair; S532: if the standby link fails to issue, it means that the regional controller may have a fault, at this time, the configuration is issued through the standby regional controller corresponding to the regional controller; at the same time, an alarm is sent to notify the relevant operation and maintenance personnel to repair. S533, if the standby regional controller also fails to issue, it is determined that the control plane network is in large-scale failure or system defect, at this time, the emergency alarm operation and maintenance personnel are notified to carry out key repair; S54, data anomaly failure and solution; If the interface reports an error, it is determined whether the failure is caused by data anomaly; The above data problem mainly refers to the case that the uploaded control plane data passes the verification, but does not pass the verification of the bottom layer network element layer. Such a problem is mainly caused by inconsistency between upper and lower layers. At this time, the system re-triggers the automatic reconciliation function. If the reconciliation data is inconsistent, the system sends the reconciliation result to the operation and maintenance personnel to verify whether the data needs to be synchronized. If the data needs to be synchronized, retry after data synchronization to check whether the problem is solved. S55, system defect; If the interface reports an error, neither the data nor the link is abnormal, and parameter verification and business verification operations have been performed. At this time, it is determined that there is a system defect, an alarm is given, and the relevant operation on-duty personnel are notified to locate and solve the problem.

7. The system for ensuring stability of a wide-area cloud network management and control plane of claim 6, wherein: The system comprises a patrol case platform, a to-be-issued queue, a strategy setting platform, a log collection and monitoring, a patrol strategy configuration module, a fault processing module and data reconciliation. The patrol case platform is used to manage patrol cases. Users manage the patrol cases of each task through the patrol case platform, and modify and add related cases to adapt to the changes of the platform. The to-be-issued queue is used when the management and control issuing channel is in a busy state due to high water level, so that the patrol case cannot be executed in time. The patrol case is first entered into the to-be-issued queue, so that the system re-processes the tasks in the to-be-issued queue after a period of time. The strategy setting platform is used for users to set the priority of related businesses through the platform, thereby affecting the setting of the entire patrol case platform. The log collection and monitoring is mainly used to collect the configurations issued by the global controller and the regional controller for each business, thereby affecting the entire issuing task. The patrol strategy configuration module actively sets the strategy through the strategy setting platform or calculates the optimized scheme according to the current situation through the monitoring system. The fault processing module is mainly used to process the problems found in the patrol, such as line disconnection and business failure. Data reconciliation is mainly used to process the influence caused by inconsistent data, dirty data or data loss.

8. The system for ensuring stability of a wide-area cloud network management and control plane of claim 7, wherein: In the strategy setting platform, users set the weight, number of times and other factors to calculate the priority of the case. The formula for calculating the score of the patrol case is ((business weight / number of times)+channel coefficient)*unstable coefficient. The higher the score calculated by the formula, the earlier the patrol case is executed.

9. The system for ensuring stability of a wide-area cloud network management and control plane of claim 8, wherein: In the strategy setting platform, the business weight and the number of times refer to the division of businesses into core businesses, key businesses, general businesses and test businesses. Among them, the core business refers to the business covering more than 60% of customers and the business of more than 90% of VIP customers, the core business weight is 1000, and the number is 10 times; the key business refers to the further business coverage of the core business, including more than 80% of the business and all VIP business, the key business weight is 400, and the number is 5 times; the general business includes all online businesses in use, the weight is 120, and the number is 2 times; the test business refers to the new business in the test process, the weight is 50, and the number is 1 time; The channel coefficient is the importance of the link, which is determined according to the traffic of the link. The coefficient of a particularly important link is 100, the coefficient of a generally important link is 50, and the coefficient of a link used less frequently is 10; The unstable coefficient is determined by the monitoring key item, the number of failures and the average number of failures. Generally, the initial value of the coefficient is 1 when the monitoring key item alarms once, and the coefficient cannot be less than 1. The coefficient increases by 0.02 when the request fails once, and the coefficient decreases by 0.1 when there is no failure for an average of 6 hours.

10. The system for ensuring stability of a wide-area cloud network management and control plane of claim 9, wherein: The patrol strategy configuration module is the core module of the entire system. In the patrol strategy configuration module, the optimized scheme needs to be confirmed by the operation and maintenance personnel before switching.

Citation Information

Patent Citations

  • RAID inspection method and device, electronic equipment and readable storage medium

    CN110659135A

  • Cluster inspection method, device and system

    CN113472577A

  • Automatic inspection strategy system based on data-intensive batch task scheduling

    CN117236669A

  • Intelligent early warning and disposal method based on historical monitoring data

    CN117827608A

  • Method and system for guaranteeing stability of wide area cloud network management and control surface

    CN118488049A

Cited By

  • Automatic inspection and fault processing method and system for video conference

    CN121792723A