Network health check method, apparatus, device, medium, and program product
By combining anomaly alerts from network infrastructure with data center topology, and employing multi-category check scripts for network health checks, the low accuracy problem caused by reliance on prediction in existing technologies is resolved, enabling precise monitoring and rapid response of network devices and services.
Patent Information
- Application Number
- CN202411324485.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-23
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-09-23
AI Technical Summary
Existing network health check methods rely too heavily on prediction, which can lead to significant disruptions to other network devices or services and cannot accurately determine the normal state of the network environment, potentially causing losses.
By analyzing anomaly alerts based on network infrastructure, combining the data center topology with the region to which the anomaly alerted object belongs, we can identify associated devices and perform health checks using different types of inspection scripts, including deployment, configuration, and status checks, to accurately locate the anomaly object.
It improves the accuracy of network health checks, enabling timely detection and handling of anomalies, avoiding unnecessary losses, and enhancing the efficiency of fault location and handling.
Smart Images

Figure CN119363556B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of network technology and financial technology, specifically to network health check methods, devices, equipment, media, and program products. Background Technology
[0002] With the explosive growth of data centers in the financial sector, the scale and complexity of networks are increasing, making routine network health checks increasingly important. By checking whether network configurations and key elements meet expectations, it is possible to determine whether the current network equipment and network environment are functioning properly and can provide services to the outside world.
[0003] In the process of implementing the embodiments of this disclosure, it was found that the current methods for inspecting network devices and network environments are only for the network itself and rely too much on prediction. However, in addition to the specified switches, the interruption of other network devices or services also has a significant impact. If the prediction is missed, it can easily cause great losses. Summary of the Invention
[0004] In view of the above problems, this disclosure provides network health check methods, apparatus, equipment, media and program products.
[0005] According to a first aspect of this disclosure, a network health check method is provided, comprising: determining an anomaly alert object to which the anomaly alert information is directed based on anomaly alert information of a network facility, wherein the anomaly alert object is used to indicate an object providing different functions in the network facility; determining multiple associated devices associated with the anomaly alert object based on the topology of a data center and the region to which the anomaly alert object belongs, wherein the data center is used to provide services to the network facility; determining associated objects associated with the anomaly alert object from the multiple associated devices based on the hierarchy to which the anomaly alert object belongs, wherein the hierarchy is used to indicate the level of the anomaly alert object in the network architecture; and performing a health check on the associated objects based on pre-configured check scripts of different categories.
[0006] According to embodiments of this disclosure, determining an associated object from multiple associated devices based on the hierarchy to which the exception notification object belongs includes: determining the level above the hierarchy to which the exception notification object belongs; filtering target devices associated with the level above from the multiple associated devices; and determining the target device as the associated object.
[0007] According to embodiments of this disclosure, the hierarchical levels, from low to high, include single-line level, single-device level, high-availability group level, network area level, and network service level. The process of selecting target devices associated with a higher level from multiple associated devices includes: if the level to which the anomaly alert object belongs is determined to be single-line level, selecting network devices with direct connections to the network line from multiple associated devices to obtain target devices associated with the single-device level; if the anomaly alert object is determined to be a network device and the level to which it belongs is single-device level, selecting other network devices forming a high-availability group from multiple associated devices to obtain target devices associated with the high-availability group level; if the anomaly alert object is determined to be a network device and the level to which it belongs is high-availability group level, selecting other network devices in the network area where the network device is located from multiple associated devices to obtain target devices associated with the network area level; and if the anomaly alert object is determined to be a network device and the level to which it belongs is network area level, selecting other network devices providing network services from multiple associated devices to obtain target devices associated with the network service level.
[0008] According to embodiments of this disclosure, topology relationships are used to characterize the layout of network devices within data centers belonging to the same region and the network line connections between network devices, as well as the linkage relationships between data centers belonging to different regions; wherein, based on the topology relationships for data centers and the region to which the anomaly alert object belongs, determining multiple associated devices associated with the anomaly alert object includes: determining a target region matching the region to which the anomaly alert object belongs from the topology relationships; and determining the network devices within the data centers of the target region as multiple associated devices associated with the anomaly alert object.
[0009] According to embodiments of this disclosure, the network health check method further includes: when it is determined that the abnormality alert object is a network device, and the level to which the abnormality alert object belongs is a high availability group, and there is a linkage relationship between the data center of the region to which the network device belongs and the data center of other regions, the network device of the data center of other regions is identified as an associated object associated with the abnormality alert object.
[0010] According to embodiments of this disclosure, the inspection script includes: a deployment inspection script, a configuration inspection script, and a status inspection script. The deployment inspection script is used to check whether there are any abnormalities in the deployment of the associated object, the configuration inspection script is used to check whether there are any abnormalities in the configuration of the associated object, and the status inspection script is used to check whether there are any abnormalities in the operation of the associated object.
[0011] Network health check methods also include: when it is determined that the results of health checks on associated objects are all normal, sending an abnormality alert to a third-party device is a false alarm.
[0012] According to embodiments of this disclosure, the network health check method further includes: when it is determined that the result of a health check on an associated object is abnormal, generating a handling instruction according to the category of the check script that caused the abnormality, and sending the handling instruction to a third-party device; and when receiving confirmation information sent by the third-party device, issuing the handling instruction to the abnormal notification object so as to execute the handling instruction on the abnormal notification object.
[0013] A second aspect of this disclosure provides a network health check apparatus, comprising: a first determining module, configured to determine an abnormality alert object to which the abnormality alert information is directed based on abnormality alert information of a network facility, wherein the abnormality alert object is used to indicate an object providing different functions in the network facility; a second determining module, configured to determine multiple associated devices associated with the abnormality alert object based on the topology of a data center and the region to which the abnormality alert object belongs, wherein the data center is used to provide services to the network facility; a third determining module, configured to determine associated objects associated with the abnormality alert object from among the multiple associated devices based on the level to which the abnormality alert object belongs, wherein the level is used to indicate the level of the abnormality alert object in the network architecture; and a checking module, configured to perform health checks on the associated objects based on pre-configured check scripts of different categories.
[0014] A third aspect of this disclosure provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the network health check method described above.
[0015] A fourth aspect of this disclosure also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the network health check method described above.
[0016] The fifth aspect of this disclosure also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the network health check method described above.
[0017] According to embodiments of this disclosure, network health checks can be proactively initiated upon receiving anomaly alerts. During network health checks, based on the topology of the data center and the region to which the anomaly alert object belongs, multiple associated devices can be accurately identified. Furthermore, based on the hierarchical level of the anomaly alert object, associated objects can be determined from among these devices. This allows for the identification of associated objects from different layers within the computer network, improving the accuracy of network health checks and avoiding the low accuracy caused by considering only the dependencies between objects. Performing different types of checks on associated objects further enhances the accuracy of network monitoring, facilitating the identification of anomalies and enabling timely handling to prevent unnecessary losses. Moreover, the network health check method provided by this disclosure can further accurately locate anomalies when applications malfunction and devices with installed applications issue warnings, improving the efficiency of anomaly location and handling. Attached Figure Description
[0018] The foregoing contents, as well as other objects, features, and advantages of this disclosure, will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0019] Figure 1 The illustration schematically depicts application scenarios of network health check methods, apparatus, devices, media, and program products according to embodiments of the present disclosure;
[0020] Figure 2 A flowchart illustrating a network health check method according to an embodiment of the present disclosure is shown schematically.
[0021] Figure 3 A schematic diagram illustrating the hierarchy of a network architecture according to an embodiment of the present disclosure is provided.
[0022] Figure 4 A schematic diagram illustrating a data center topology according to an embodiment of the present disclosure is shown.
[0023] Figure 5 A flowchart illustrating a network health check method according to another embodiment of the present disclosure is shown schematically;
[0024] Figure 6 A schematic diagram illustrating the structure of a network health check device according to an embodiment of the present disclosure is shown; and
[0025] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing a network health check method according to an embodiment of the present disclosure. Detailed Implementation
[0026] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0027] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0028] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0029] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0030] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.
[0031] In the technical solution disclosed herein, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation entry points are provided for users to choose to authorize or refuse.
[0032] In scenarios involving automated decision-making using personal information, the methods, devices, and systems provided in this disclosure all offer users corresponding entry points for choosing to agree to or reject the automated decision-making results. If the user chooses to reject, the process proceeds to the expert decision-making stage. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and then making a decision. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.
[0033] In implementing the embodiments of this disclosure, it was found that current methods for inspecting network devices and the network environment only target the network itself and rely excessively on prediction. However, besides specifying switches, interruptions in other network devices or services also have a significant impact, and omissions in prediction can easily lead to substantial losses. For example, in related examples, although the method of using upstream network layer devices to perform health checks only on adjacent downstream network layer devices addresses the technical problem of excessive performance overhead caused by health checks to downstream devices, it mainly considers the relationship between networks and cannot determine the current network environment or even the application's situation.
[0034] Embodiments of this disclosure provide network health check methods, apparatus, devices, media, and program products that can be applied to the fields of network technology and financial technology. The method includes: determining the abnormal object to which the abnormality alert information is addressed based on abnormality alert information from network facilities, wherein the abnormality alert object is used to indicate objects providing different functions within the network facility; determining multiple associated devices related to the abnormality alert object based on the topology of a data center and the region to which the abnormality alert object belongs, wherein the data center provides services to the network facility; determining associated objects from the multiple associated devices based on the hierarchy to which the abnormality alert object belongs, wherein the hierarchy indicates the level of the abnormality alert object in the network architecture; and performing a health check on the associated objects based on pre-configured check scripts of different categories.
[0035] Figure 1 The illustration schematically depicts application scenarios of network health check methods, apparatus, devices, media, and program products according to embodiments of the present disclosure.
[0036] like Figure 1As shown, application scenario 100 according to this embodiment may include a first network device 101, a second network device 102, a third network device 103, a network 104, and a server 105. Network 104 serves as a medium for providing communication links between network devices 101, 102, 103, and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0037] The first network device 101, the second network device 102, and the third network device 103 interact with the server 105 through the network 104 to receive or send messages, etc. The first network device 101, the second network device 102, and the third network device 103 can be physical infrastructure belonging to the computer network, including but not limited to routers, switches, hubs, gateways, firewalls, wireless access points, servers, workstations, network storage devices, etc.
[0038] Server 105 can be a server that provides various services, such as a backend management server that supports the first network device 101, the second network device 102, and the third network device 103 (for example only). The backend management server can analyze and process the received abnormal prompts and feed back the processing results to a third-party device. The third-party device can be an operating device used by network operation and maintenance personnel to control the first network device 101, the second network device 102, and the third network device 103.
[0039] It should be noted that the network health check method provided in this embodiment can generally be executed by server 105. Correspondingly, the network health check device provided in this embodiment can generally be located in server 105. The network health check method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first network device 101, the second network device 102, the third network device 103, and / or server 105. Correspondingly, the network health check device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first network device 101, the second network device 102, the third network device 103, and / or server 105.
[0040] It should be understood that Figure 1 The number of network devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of network devices, networks, and servers can be included.
[0041] The following will be based on Figure 1 The described scene, through Figures 2-5 The network health check method of the disclosed embodiments is described in detail.
[0042] Figure 2 A flowchart illustrating a network health check method according to an embodiment of the present disclosure is shown schematically.
[0043] like Figure 2 As shown, the network health check method of this embodiment includes operations S210 to S240.
[0044] When operating S210, determine the target of the abnormality message based on the abnormality message from the network facility.
[0045] In operation S220, based on the topology of the data center and the region to which the anomaly alert object belongs, multiple associated devices are identified that are associated with the anomaly alert object.
[0046] In operation S230, based on the level to which the exception notification object belongs, the associated object is determined from multiple associated devices.
[0047] In operation S240, health checks are performed on associated objects based on pre-configured check scripts for different categories.
[0048] According to embodiments of this disclosure, network facilities may include physical infrastructure belonging to a computer network, such as, but not limited to, network devices and network lines. Network lines can be used to connect network devices.
[0049] According to embodiments of this disclosure, anomaly alert information can be used to indicate a suspected anomaly in a network facility that provides different functions. This disclosure does not specifically limit the method for obtaining anomaly alert information for network facilities; for example, it may include, but is not limited to, generating anomaly alert information when a monitoring program installed on the network device detects that the performance of the network device or network line does not match the expected performance.
[0050] According to embodiments of this disclosure, the exception notification object can be used to indicate objects that provide different functions in the network facility, such as including but not limited to network devices, network lines, network services, etc.
[0051] According to embodiments of this disclosure, the region to which an anomaly alert object belongs can be determined, and the data center to which the anomaly alert object belongs can be determined based on the region. The data center to which the anomaly alert object belongs is determined from the data center topology, and based on the data center topology, network devices within the data center to which the anomaly alert object belongs are identified as multiple associated devices related to the anomaly alert object.
[0052] According to embodiments of this disclosure, a hierarchy can be used to indicate the level of an exception notification object within the network architecture. For example, the hierarchy to which a network device belongs can be a single device level, a high availability group level, a network area level, or a network service level, etc.; the hierarchy to which a network service belongs can be a network service level; and the hierarchy to which a network line belongs can be a single line level, etc.
[0053] According to embodiments of this disclosure, a data center is used to provide services to network facilities. The topology relationships for a data center may include topology relationships between regional data centers and topology relationships between network devices within a data center. Associated objects may be used to indicate objects that are associated with anomaly alert objects at different layers within the computer network.
[0054] According to embodiments of this disclosure, the topology of network devices within a data center within a region of the data center to which the anomaly alert object belongs can be determined based on the topology of the data center. Associated objects are then determined from the topology of network devices within the region of the data center to which the anomaly alert object belongs, based on the hierarchy to which the anomaly alert object belongs.
[0055] According to embodiments of this disclosure, different categories of inspection scripts may include metrics for network management and monitoring, such as, but not limited to, deployment, configuration, and status metrics. Pre-configured inspection scripts of different categories can be retrieved and run, and the results can be used to determine whether the anomaly alerts are false alarms.
[0056] According to embodiments of this disclosure, network health checks can be proactively initiated upon receiving anomaly alerts. During network health checks, based on the topology of the data center and the region to which the anomaly alert object belongs, multiple associated devices can be accurately identified. Furthermore, based on the hierarchical level of the anomaly alert object, associated objects can be determined from among these devices. This allows for the identification of associated objects from different layers within the computer network, improving the accuracy of network health checks and avoiding the low accuracy caused by considering only the dependencies between objects. Performing different types of checks on associated objects further enhances the accuracy of network monitoring, facilitating the identification of anomalies and enabling timely handling to prevent unnecessary losses. Moreover, the network health check method provided by this disclosure can further accurately locate anomalies when applications malfunction and devices with installed applications issue warnings, improving the efficiency of anomaly location and handling.
[0057] In implementing the concept disclosed herein, it was discovered that although network topology can be used to identify objects that have dependencies on the abnormality alert object and then perform health checks, service interruptions or high availability anomalies can affect the accuracy of health check results. Therefore, considering only the dependencies between objects themselves will result in low accuracy of health check results.
[0058] Based on this, this embodiment addresses the above. Figure 2 The operation S230 shown, which determines the associated object from multiple associated devices based on the hierarchy to which the exception notification object belongs, may include the following operations: determining the level above the hierarchy to which the exception notification object belongs; filtering the target devices associated with the level above from the multiple associated devices; and determining the target devices as associated objects.
[0059] According to embodiments of this disclosure, the network architecture may include multiple different levels of hierarchy. The next higher level is at a higher level than the hierarchy to which the exception notification object belongs. For example, if the exception notification object is determined to be a network line, the hierarchy to which the exception notification object belongs may be the single-line level. The next higher level above the hierarchy to which the exception notification object belongs may be the single-device level.
[0060] According to embodiments of this disclosure, a target device associated with a higher level is used to indicate that the function provided by the target device belongs to the higher level, or that the target device belongs to the higher level, etc.
[0061] According to embodiments of this disclosure, by considering the topology of a data center and the object's level within the network architecture, associated objects related to the anomaly alert object can be identified from different layers of the computer network, improving the accuracy of network health checks. This at least partially solves the technical problem that considering only the dependencies between objects themselves leads to low accuracy in health check results.
[0062] According to another embodiment of this disclosure, if it is determined that there is an abnormality in the result of a health check on an associated object, the level to which the current abnormality-prone object belongs can be updated to the previous level in the last health check.
[0063] Repeat the following steps: Determine the level above the current exception object. Based on the data center topology, identify multiple associated devices related to the exception object. Filter the multiple associated devices to find the target device associated with the previous level. Finalize the target device as the associated object.
[0064] According to embodiments of this disclosure, by performing health checks layer by layer from low to high, abnormal objects can be accurately identified so as to provide early warnings and resolve risks and faults in their infancy.
[0065] Figure 3 A schematic diagram illustrating the hierarchy of a network architecture according to an embodiment of the present disclosure is shown.
[0066] like Figure 3 As shown, the network architecture hierarchy, from low to high, can include single-line level 310, single-device level 320, high-availability group level 330, network area level 340, and network service level 350.
[0067] In one embodiment, filtering target devices associated with the previous level from multiple associated devices may include the following operation: if it is determined that the level to which the abnormal prompt object belongs is single-line level 310, filtering network devices that have a direct connection relationship with the network line from multiple associated devices to obtain target devices associated with single-device level 320.
[0068] According to embodiments of this disclosure, when the level to which the anomaly alert object belongs is single-line level 310, the corresponding anomaly alert object can be a network line. By performing a health check on the network devices connected to both ends of the network line, if it is determined that the network devices connected to both ends of the network line are in a normal state, it can be determined that the network line is in a normal state, and the anomaly alert information is a false alarm. If it is determined that the network devices connected to both ends of the network line are in an abnormal state, the target device can be further determined based on the high availability group level 330 above the single device level 320 and the network devices connected to both ends of the network line. For example, it can be determined whether the network devices connected to both ends of the network line are high availability network devices. If so, other network devices constituting the high availability group level can be determined as target devices associated with the high availability group level 330 above the single device level 320.
[0069] According to embodiments of this disclosure, the network devices connected at both ends of a network line can be determined based on the topology for a data center.
[0070] In another embodiment, when the abnormality prompt information is used to indicate that the network device has a suspected abnormality, it can be determined that the abnormality prompt information is directed to the network device.
[0071] If the anomaly is identified as a network device and the anomaly belongs to the single device level, other network devices that make up the high availability group can be selected from multiple associated devices to obtain the target device associated with the high availability group level.
[0072] If the anomaly is identified as a network device and the anomaly belongs to a high availability group, other network devices in the network area where the network device is located are selected from multiple associated devices to obtain the target device associated with the network area level.
[0073] If the anomaly is identified as a network device and the level to which the anomaly belongs is the network area level, other network devices that provide network services are selected from multiple associated devices to obtain the target device associated with the network service level.
[0074] According to embodiments of this disclosure, other network devices for providing high availability can be determined based on the mapping relationship between multiple associated devices and layers, and these other network devices together with the network device form a high availability group.
[0075] According to embodiments of this disclosure, other network devices for providing network services to external devices can be determined based on the mapping relationship between multiple associated devices and hierarchies.
[0076] According to embodiments of this disclosure, considering that anomalies at different levels in the network architecture can also affect the accuracy of health check results, the accuracy of health check results is improved by screening target devices associated with the previous level from multiple associated devices for health checks. This is beneficial for accurately locating anomalies when network device anomalies are indicated, but other anomaly information of the specific network device is unclear, thereby improving the efficiency of fault location and handling, and helping operation and maintenance administrators narrow down the scope of fault root cause location.
[0077] Figure 4 A schematic diagram of a data center topology according to an embodiment of the present disclosure is shown.
[0078] According to embodiments of this disclosure, a data center may include a management data center and data centers in different regions. The following refers to... Figure 4 The examples provided further illustrate the topology of a data center, but are merely illustrative and not intended to limit the scope of this disclosure.
[0079] like Figure 4 As shown, the topology of a data center can include a hierarchical topology. The management data center is the core layer, and data centers in different regions, such as data centers in Region 1, Region 2, and Region 3, form the distribution layer. For example, in the core layer, aggregation switches belonging to different regions can be defined, followed by core switches. In the distribution layer, switch devices can be defined, based on their purpose, vendor, or business area (production, management, etc.). The topology of the data center can be obtained by acquiring data from actual network devices and mapping them to relevant definitions according to their purpose, type, and network architecture hierarchy.
[0080] According to embodiments of this disclosure, topology can be used to characterize the layout of network devices within a data center belonging to the same region and the network line connections between network devices, as well as the linkage between data centers belonging to different regions.
[0081] Based on the data center topology and the region to which the anomaly alert object belongs, multiple associated devices are identified. This may include the following operations: determining a target region that matches the region to which the anomaly alert object belongs from the topology; and identifying network devices within the data center of the target region as the multiple associated devices associated with the anomaly alert object.
[0082] According to embodiments of this disclosure, multiple associated devices related to the abnormality alert object can be identified through region matching, which facilitates the filtering of associated devices within the region at the hierarchical level.
[0083] According to another embodiment of this disclosure, the topology of a data center can be constructed by defining an internal view of each data center. For example, the switching devices within each region of the data center can be defined according to different functional service areas. Then, the network devices are mapped to the relevant definitions based on their purpose, type, and network architecture hierarchy to obtain the topology of the data center.
[0084] According to another embodiment of this disclosure, the topology of a data center can be constructed by defining a network area view. For example, a network area can be composed of one or more computer rooms, such as a network composed of core switches, aggregation switches, and aggregation / access switches. The topology of the data center is obtained by mapping the switches to the switches according to their purpose, type, and network architecture level.
[0085] According to another embodiment of this disclosure, in addition to operations S210 to S240, the network health check method may further include the following operation after determining the associated objects associated with the abnormality alert object from multiple associated devices based on the level to which the abnormality alert object belongs: if it is determined that the abnormality alert object is a network device, and the level to which the abnormality alert object belongs is a high availability group level, and there is a linkage relationship between the data center of the region to which the network device belongs and the data centers of other regions, then the network devices of the data centers of other regions are also determined as associated objects associated with the abnormality alert object.
[0086] According to embodiments of this disclosure, it is possible to determine whether there is a linkage relationship between data centers in the region where network devices belong and data centers in other regions by managing the data center. This linkage relationship can be used to indicate network interaction between data centers.
[0087] According to embodiments of this disclosure, since network device anomalies can also affect network device anomalies within data centers that are interconnected, by jointly identifying network devices in other interconnected data centers and associated objects determined from multiple associated devices that are associated with the anomaly alert object as associated objects, the accuracy of network health checks can be improved by considering the hierarchical levels within the same region and combining the levels of different regions.
[0088] Figure 5 A flowchart illustrating a network health check method according to another embodiment of the present disclosure is shown schematically.
[0089] According to embodiments of this disclosure, the inspection script may include: a deployment inspection script, a configuration inspection script, and a status inspection script. The deployment inspection script can be used to check whether there are any abnormalities in the deployment of the associated object, the configuration inspection script can be used to check whether there are any abnormalities in the configuration of the associated object, and the status inspection script can be used to check whether there are any abnormalities in the operation of the associated object.
[0090] According to embodiments of this disclosure, checking for anomalies in the deployment of associated objects may include deployment-related indicators such as hardware / software consistency, high availability of location deployment, consistency of device quantity, and consistency of device components. Checking for anomalies in the configuration of associated objects may include configuration-related indicators such as functional-structural configuration, functional-service configuration, and non-functional-management configuration. Checking for anomalies in the operational status of associated objects may include status-related indicators such as operational status, table entries, performance capacity, and traffic balancing.
[0091] like Figure 5 As shown, the network health check method of this embodiment includes operations S510 to S560.
[0092] When operating S510, pre-configured deployment class check scripts, configuration class check scripts, and status class check scripts are invoked.
[0093] When operating S520, deployment check scripts, configuration check scripts, and status check scripts are run based on the data of associated objects to obtain the results of health checks.
[0094] When operating S530, check whether the results of the health check on the associated objects are all normal.
[0095] According to an embodiment of this disclosure, if the results of health checks on the associated objects are all normal, operation S540 is executed; otherwise, operation S550 is executed.
[0096] When operating the S540, a false alarm message was sent to a third-party device indicating an anomaly.
[0097] When operating the S550, a handling instruction is generated based on the category of the check script that corresponds to the anomaly, and the handling instruction is sent to a third-party device.
[0098] When operating the S560, upon receiving confirmation information from a third-party device, a handling instruction is sent to the exception notification object so that the handling instruction can be executed on the exception notification object.
[0099] According to embodiments of this disclosure, the third-party device may be an operating device used by network operation and maintenance personnel to control network devices.
[0100] According to embodiments of this disclosure, when the health check results corresponding to the three types of check scripts for the associated object are all normal, the abnormality alert is determined to be a false alarm. The abnormality alert can then be sent to the object's operations and maintenance administrator via a third-party device. Alternatively, the health check information can also be sent to the operations and maintenance administrator via a third-party device. Based on the obtained information and the judgment conclusion, the operations and maintenance administrator can further determine whether to adjust the check scripts.
[0101] According to embodiments of this disclosure, considering that the anomaly alert object is generally the core of the relationship, if the anomaly alert object is abnormal, then all surrounding services should also be abnormal. Therefore, if some surrounding objects are normal and some are abnormal, a handling instruction can be sent to a third-party device. The operations and maintenance administrator can organize group discussions with other operations and maintenance personnel regarding the associated objects of the anomaly and the category of the anomaly inspection script to determine whether there are potential risks and whether the automatically generated handling instructions are reasonable, thus resolving risks and faults in their early stages.
[0102] According to the embodiments of this disclosure, when the health check results of the three types of check scripts corresponding to the associated object are all abnormal, it is highly likely that the abnormality is caused by the abnormality of the object. After the operation and maintenance administrator of the abnormality object confirms the automatically generated handling instructions, the handling instructions are sent to the abnormality object.
[0103] According to embodiments of this disclosure, the disposal instructions may include isolation, restart, etc.
[0104] According to embodiments of this disclosure, the efficiency and effectiveness of pre-fault handling of abnormal objects achieved through health checks can be improved by using multi-category inspection scripts.
[0105] Based on the above-described network health check method, this disclosure also provides a network health check device. The following will be combined with... Figure 6 The device is described in detail.
[0106] Figure 6 A schematic block diagram of a network health check device according to an embodiment of the present disclosure is shown.
[0107] like Figure 6 As shown, the network health check device 600 of this embodiment includes a first determining module 610, a second determining module 620, a third determining module 630, and a check module 640.
[0108] The first determining module 610 is used to determine the abnormal object to which the abnormal prompt information is addressed based on the abnormal prompt information of the network facility, wherein the abnormal prompt object is used to indicate the object providing different functions in the network facility. In one embodiment, the first determining module 610 can be used to perform the operation S210 described above, which will not be repeated here.
[0109] The second determining module 620 is used to determine multiple associated devices related to the anomaly alert object based on the topology of the data center and the region to which the anomaly alert object belongs, wherein the data center is used to provide services to network facilities. In one embodiment, the second determining module 620 may be used to perform the operation S220 described above, which will not be repeated here.
[0110] The third determining module 630 is used to determine the associated object from multiple associated devices based on the level to which the exception notification object belongs, wherein the level is used to indicate the level of the exception notification object in the network architecture. In one embodiment, the third determining module 630 can be used to perform the operation S230 described above, which will not be repeated here.
[0111] The inspection module 640 is used to perform health checks on associated objects based on pre-configured inspection scripts of different categories. In one embodiment, the inspection module 640 can be used to perform the operation S240 described above, which will not be repeated here.
[0112] According to embodiments of this disclosure, the second determining module 620 may include: a first sub-determining unit, a filtering unit, and a second sub-determining unit.
[0113] The first sub-determination unit is used to determine the level above the level to which the exception notification object belongs. The filtering unit is used to filter target devices associated with the level above from multiple associated devices. The second sub-determination unit is used to determine the target device as the associated object.
[0114] According to embodiments of this disclosure, the hierarchical levels, from low to high, may include single-line level, single-device level, high-availability group level, network area level, and network service level. The process of selecting target devices associated with a higher level from multiple associated devices includes: if the level to which the anomaly alert object belongs is determined to be single-line level, selecting network devices with direct connections to the network line from multiple associated devices to obtain target devices associated with the single-device level; if the anomaly alert object is determined to be a network device and the level to which it belongs is single-device level, selecting other network devices forming a high-availability group from multiple associated devices to obtain target devices associated with the high-availability group level; if the anomaly alert object is determined to be a network device and the level to which it belongs is high-availability group level, selecting other network devices in the network area where the network device is located from multiple associated devices to obtain target devices associated with the network area level; and if the anomaly alert object is determined to be a network device and the level to which it belongs is network area level, selecting other network devices providing network services from multiple associated devices to obtain target devices associated with the network service level.
[0115] According to embodiments of this disclosure, topology relationships can be used to characterize the layout of network devices within data centers belonging to the same region and the network line connections between network devices, as well as the linkage relationships between data centers belonging to different regions; wherein, based on the topology relationships for data centers and the region to which the anomaly alert object belongs, determining multiple associated devices associated with the anomaly alert object includes: determining a target region matching the region to which the anomaly alert object belongs from the topology relationships; and determining the network devices within the data centers of the target region as multiple associated devices associated with the anomaly alert object.
[0116] According to embodiments of this disclosure, the network health check device 600 may further include a fourth determining module. The fourth determining module is used to determine network devices in other data centers as associated objects related to the abnormal notification object when it is determined that the abnormal notification object is a network device, the abnormal notification object belongs to a high availability group level, and there is a linkage relationship between the data center in the region to which the network device belongs and data centers in other regions.
[0117] According to embodiments of this disclosure, the inspection script may include: a deployment inspection script, a configuration inspection script, and a status inspection script. The deployment inspection script is used to check whether there are any abnormalities in the deployment of the associated object, the configuration inspection script is used to check whether there are any abnormalities in the configuration of the associated object, and the status inspection script is used to check whether there are any abnormalities in the operation of the associated object.
[0118] The network health check device 600 may also include a sending module. The sending module is used to send a notification to a third-party device indicating that the abnormality was a false alarm, provided that the health check results for the associated objects are all normal.
[0119] According to embodiments of this disclosure, the network health check device 600 may further include a generation module and a distribution module. The generation module, upon determining that an anomaly exists in the result of a health check on an associated object, generates a handling instruction based on the category of the corresponding abnormal check script and sends the handling instruction to a third-party device. The distribution module, upon receiving confirmation information from the third-party device, distributes the handling instruction to the abnormal notification object so that the handling instruction can be executed on the abnormal notification object.
[0120] According to embodiments of this disclosure, any plurality of modules among the first determining module 610, the second determining module 620, the third determining module 630, and the checking module 640 may be combined into one module, or any one of these modules may be split into multiple modules. Alternatively, at least a portion of the functionality of one or more of these modules may be combined with at least a portion of the functionality of other modules and implemented in one module. According to embodiments of this disclosure, at least one of the first determining module 610, the second determining module 620, the third determining module 630, and the checking module 640 may be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in any one of software, hardware, and firmware methods, or in a suitable combination of any of these methods. Alternatively, at least one of the first determining module 610, the second determining module 620, the third determining module 630, and the checking module 640 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.
[0121] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing a network health check method according to an embodiment of the present disclosure.
[0122] like Figure 7As shown, an electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage portion 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0123] RAM 703 stores various programs and data required for the operation of electronic device 700. Processor 701, ROM 702, and RAM 703 are interconnected via bus 704. Processor 701 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 702 and / or RAM 703. It should be noted that programs may also be stored in one or more memories other than ROM 702 and RAM 703. Processor 701 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in one or more memories.
[0124] According to embodiments of this disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to a bus 704. The electronic device 700 may also include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 710 as needed so that computer programs read from it can be installed into the storage section 708 as needed.
[0125] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.
[0126] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this disclosure, the computer-readable storage medium may include ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703 described above.
[0127] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to enable the computer system to implement the network health check method provided in the embodiments of this disclosure.
[0128] When the computer program is executed by the processor 701, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0129] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 709, and / or installed from a removable medium 711. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0130] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, it performs the functions defined in the system of this disclosure embodiment. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0131] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on a user's computing device, partially on a user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0133] Those skilled in the art will understand that the features described in the various embodiments of this disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments of this disclosure can be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0134] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. A method for checking network health, characterized in that, The method includes: Based on the abnormal prompt information of the network facility, determine the abnormal prompt object to which the abnormal prompt information is directed, wherein the abnormal prompt object is used to indicate the object that provides different functions in the network facility; Based on the topology of the data center and the region to which the anomaly alert object belongs, multiple associated devices are identified that are associated with the anomaly alert object, wherein the data center is used to provide services to the network infrastructure; Based on the hierarchy to which the anomaly alert object belongs, associated objects are determined from among the multiple associated devices, wherein the hierarchy indicates the level of the anomaly alert object in the network architecture; and Health checks are performed on the associated objects based on pre-configured check scripts for different categories.
2. The method according to claim 1, characterized in that, The step of determining the associated object from among the multiple associated devices based on the hierarchy to which the exception notification object belongs includes: Determine the level above the level to which the exception notification object belongs; Filter target devices associated with the previous level from among the multiple associated devices; and The target device is identified as the associated object.
3. The method according to claim 2, characterized in that, The hierarchy, from lowest to highest, includes single-line level, single-device level, high-availability group level, network area level, and network service level. The step of filtering the target device associated with the previous level from the plurality of associated devices includes: If the level to which the abnormal prompt object belongs is determined to be the single-line level, network devices with direct connection to the network line are selected from multiple associated devices to obtain the target device associated with the single-device level; If the abnormality alert object is determined to be a network device and the level to which the abnormality alert object belongs is the single device level, other network devices that form a high availability group are selected from multiple associated devices to obtain the target device associated with the high availability group level; If the anomaly alert target is determined to be the network device, and the anomaly alert target belongs to the high availability group level, other network devices in the network area where the network device is located are filtered from multiple associated devices to obtain the target device associated with the network area level; and If the abnormality alert object is determined to be the network device, and the level to which the abnormality alert object belongs is the network area level, other network devices that provide network services are selected from multiple associated devices to obtain the target device associated with the network service level.
4. The method according to claim 1, characterized in that, The topology is used to characterize the layout of network devices within data centers belonging to the same region and the network line connections between network devices, as well as the linkage between data centers belonging to different regions. The step of determining multiple associated devices related to the anomaly alert object based on the topological relationship of the data center and the region to which the anomaly alert object belongs includes: Determine the target region that matches the region to which the anomaly alert object belongs from the topological relationship; and The network devices within the data center of the target area are identified as multiple associated devices related to the anomaly alert object.
5. The method according to claim 4, characterized in that, The method further includes: If the abnormality alert object is determined to be the network device, and the level to which the abnormality alert object belongs is the high availability group level, and there is a linkage relationship between the data center of the region to which the network device belongs and the data center of other regions, then the network devices of the data centers of other regions are determined as the associated objects associated with the abnormality alert object.
6. The method according to any one of claims 1 to 5, characterized in that, The inspection scripts include: a deployment inspection script, a configuration inspection script, and a status inspection script. The deployment inspection script is used to check whether there are any abnormalities in the deployment of the associated object. The configuration inspection script is used to check whether there are any abnormalities in the configuration of the associated object. The status inspection script is used to check whether there are any abnormalities in the operation of the associated object. The method further includes: If the results of health checks on the associated objects are all normal, a notification is sent to the third-party device indicating that the abnormality alert is a false alarm.
7. The method according to claim 6, characterized in that, The method further includes: If an anomaly is detected in the health check result of the associated object, a handling instruction is generated based on the category of the check script corresponding to the anomaly, and the handling instruction is sent to the third-party device; and Upon receiving confirmation information from the third-party device, the handling instruction is sent to the exception notification object so that the handling instruction can be executed on the exception notification object.
8. A network health check device, characterized in that, The device includes: The first determining module is used to determine the abnormal prompt object to which the abnormal prompt information is directed based on the abnormal prompt information of the network facility, wherein the abnormal prompt object is used to indicate the object that provides different functions in the network facility; The second determining module is used to determine multiple associated devices associated with the anomaly alert object based on the topology of the data center and the region to which the anomaly alert object belongs, wherein the data center is used to provide services to the network facility; The third determining module is configured to determine, from among the multiple associated devices, an associated object related to the anomaly alert object based on the hierarchy to which the anomaly alert object belongs, wherein the hierarchy indicates the level of the anomaly alert object in the network architecture; and The inspection module is used to perform health checks on the associated objects based on pre-configured inspection scripts of different categories.
9. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and equipment for monitoring network quality
CN108900388A
Network operation and maintenance troubleshooting method, system and related equipment
CN117768299A