Fault diagnosis method, apparatus, and system
By obtaining the topological relationship and reliability status of node objects in the communication system, and using topological relationships to derive the root cause of failure, the diagnosis failure caused by the missing event relationship in the communication system is solved, and the success rate of fault diagnosis is improved.
Patent Information
- Application Number
- PCT/CN2024/141616
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-23
- Publication Date
- 2025-07-03
AI Technical Summary
When an existing communication system lacks a certain link in the event relationship in troubleshooting, it is impossible to accurately determine the root cause of the fault, resulting in a diagnosis failure and a low success rate.
By obtaining the topological relationship and reliability status between multiple node objects, using topological relationships to deduce the missing event link, perform fault diagnosis, and improve success rate.
Even if there are missing links in the event relationship, the root cause of the fault can be accurately derived and the success rate of fault diagnosis can be improved.
Smart Images

Figure CN2024141616_03072025_PF_FP_ABST
Abstract
Description
Fault diagnosis method, device and system
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 29, 2023, with application number 202311869433.6 and application name “Fault Diagnosis Method, Device and System”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of communication technology, and in particular to a fault diagnosis method, device and system. Background Art
[0003] Currently, communication systems can rely on the relationships between events associated with objects to perform fault diagnosis. For example, in Figure 1A, object 2 is deployed on object 1, and object 3 is deployed on object 2. Event A (represented by a circle) associated with host 1 is the high CPU usage of host 1, and event C associated with host 1 is packet loss on host 1; event B associated with virtual machine 1 is the high CPU usage of virtual machine 1, and event D associated with virtual machine 1 is packet loss on the network port of virtual machine 1; event E associated with service instance 1 is the high service failure rate of service instance 1. For example, object 1 is host 1, object 2 is virtual machine (VM) 1 deployed on host 1, and object 3 is service instance 1 deployed on virtual machine 1.
[0004] In some scenarios, when object 1 generates event A, object 2 generates event B, and object 3 generates event C, the communication system can determine, based on the relationship between events A, B, and C, that the root cause of the fault that caused object 3 to generate event C is object 1's excessive CPU usage. However, in some cases, when an event link is missing from the event relationship, the communication system cannot determine the root cause based on the missing event. For example, the communication system detects event A from object 1, event C from object 3, and other events from other objects; but it does not detect event B from object 2 (for example, due to judgment conditions, alarm thresholds, or cycle factors). In other words, events A and E are detected, but event B is not. The missing link in the event relationship between events A and B makes it impossible to determine the root cause of event E, resulting in a failed diagnosis. This shows that with current technology, the probability of fault diagnosis failure is high. Summary of the Invention
[0005] The present application provides a fault diagnosis method, device and system that can improve the success rate of fault diagnosis.
[0006] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions.
[0007] In a first aspect, the present application provides a fault diagnosis method, the execution subject of which can be a communication device or a component located in the communication device (for example, a chip, a chip system or a processor, etc.). The following description is based on the execution subject being a communication device. The method may include: obtaining relationship information between multiple node objects, and performing fault diagnosis on the multiple node objects based on the relationship information between the multiple node objects. The relationship information includes the topological relationship between the multiple node objects and the reliability status of the multiple node objects.
[0008] Compared with the related art that relies heavily on the relationship between abnormal events to diagnose faults, which leads to diagnosis failure when there are missing links in the event relationship, the solution of the embodiment of the present application performs fault diagnosis based on the relationship between objects (such as topological relationships), and is applicable to a variety of scenarios. Even if there are missing links in the event relationship, the root cause of the fault can be deduced, thereby improving the success rate of fault diagnosis.
[0009] For example, as shown in Figure 1A, the node detects abnormal event A of object 1 and abnormal event E of object 3. Although the node does not detect abnormal event B on object 2, based on the deployment relationship between object 2 and object 1, the abnormality of object 1 is likely to cause the abnormality of object 2. Combined with the deployment relationship between object 2 and object 3, the node can infer that the cause of the abnormality of object 3 is the abnormality of object 2. That is, in this application, even if a certain abnormal event link (such as abnormal event B) is missing, it will not affect the fault diagnosis of the node. The node can make up for the missing abnormal event link based on the relationship information between the objects, that is, deduce the correlation between the system abnormal events, and then continue to perform fault diagnosis.
[0010] Moreover, in some cases, the node may not be able to obtain the underlying fault information. In this case, the method of the present application can be used to deduce some fault information based on the relationship between node objects, without relying on the underlying fault information, and fault diagnosis can also be achieved.
[0011] In one possible design, the plurality of node objects include a first node object and at least one second node object;
[0012] The at least one second node object is deployed on the first node object, and the topological relationship between the first node object and the at least one second node object is a deployment relationship;
[0013] Or, the first node object includes the at least one second node object, and the topological relationship between the first node object and the at least one second node object is a combination relationship;
[0014] Or, the first node object includes the at least one second node object, the topological relationship between the at least one second node object is a redundant relationship, and the topological relationship between the first node object and the at least one second node object is an aggregation relationship;
[0015] Alternatively, a communication relationship exists between the first node object and the at least one second node object, and the topological relationship between the first node object and the at least one second node object is an upstream-downstream relationship.
[0016] In one possible design, the reliability status of any node object among the multiple node objects includes normal or abnormal; the abnormality includes any of the following states: warning, damage, overload, and failure.
[0017] In a possible design, the topological relationship between the first node object and the at least one second node object is a deployment relationship, and when the reliability status of the first node object is abnormal, the reliability status of the at least one second node object is also abnormal;
[0018] Alternatively, the topological relationship between the first node object and the at least one second node object is a combination relationship, and when there is a second node object with an abnormal reliability status among the at least one second node object, the reliability status of the first node object is abnormal;
[0019] Alternatively, the topological relationship between the first node object and the at least one second node object is an aggregation relationship, and when the reliability status of the at least one second node object is abnormal, the reliability status of the first node object is abnormal.
[0020] In one possible design, the reliability status of any node object among the multiple node objects is determined based on indicators of the node object; the indicators include at least one of the following indicators: the traffic of the node object, the overload level of the node object, the response time of the node object, and the data error rate of the node object.
[0021] In one possible design, the multiple node objects include resource objects and / or business objects; the business objects are node objects with business processing functions; and the resource objects are node objects that provide resources.
[0022] In one possible design, if at least one indicator of the node object's traffic, overload level, and response time is abnormal, and the node object's data error rate is normal, the node object's reliability status is warning;
[0023] Alternatively, if at least one of the traffic flow, overload level, and response time of the node object is abnormal, the data error rate of the node object is abnormal, and the node object has not lost its node function, the reliability status of the node object is damaged;
[0024] Or, if the overload degree of the node object is abnormal and the node object has not lost its node function, the reliability status of the node object is overload;
[0025] Alternatively, if the traffic, overload level, response time, and data error rate of the node object are all abnormal, and the node object loses its node function, the reliability status of the node object is faulty.
[0026] In one possible design, it also includes:
[0027] Send the fault diagnosis result and / or fault analysis diagram.
[0028] In one possible design, the fault analysis graph includes a topology of node objects associated with a root cause of the fault.
[0029] In one possible design, the topology of the node objects associated with the root cause of the fault is determined based on multiple paths, and the multiple paths are determined based on relationship information between the multiple node objects; each of the multiple paths includes at least one node object.
[0030] In one possible design, information of the multiple paths is sent.
[0031] In one possible design, it also includes:
[0032] Information on indicators of the plurality of node objects and diagnostic rules corresponding to the indicators of the plurality of node objects are received, where the indicators of the plurality of node objects and the diagnostic rules are used to determine reliability states of the plurality of node objects.
[0033] In a possible design, obtaining relationship information between multiple node objects includes:
[0034] Relationship information between the multiple node objects is received.
[0035] In one possible design, it also includes:
[0036] Relationship information between the multiple node objects is sent, where the relationship information is used to perform fault diagnosis on the multiple node objects.
[0037] In a second aspect, a fault diagnosis method is provided. The method may be performed by a communication device or a component within the communication device (e.g., a chip, a chip system, or a processor). The following description uses the communication device as an example. The method may include: sending relationship information between multiple node objects, the relationship information including the topological relationship between the multiple node objects and the reliability status of the multiple node objects; using the relationship information to perform fault diagnosis on the multiple node objects; and receiving the results of the fault diagnosis.
[0038] In one possible design, the plurality of node objects include a first node object and at least one second node object;
[0039] The at least one second node object is deployed on the first node object, and the topological relationship between the first node object and the at least one second node object is a deployment relationship;
[0040] Or, the first node object includes the at least one second node object, and the topological relationship between the first node object and the at least one second node object is a combination relationship;
[0041] Or, the first node object includes the at least one second node object, the topological relationship between the at least one second node object is a redundant relationship, and the topological relationship between the first node object and the at least one second node object is an aggregation relationship;
[0042] Alternatively, a communication relationship exists between the first node object and the at least one second node object, and the topological relationship between the first node object and the at least one second node object is an upstream-downstream relationship.
[0043] In one possible design, the reliability status of any node object among the multiple node objects includes normal or abnormal; the abnormality includes any of the following states: warning, damage, overload, and failure.
[0044] In a possible design, the topological relationship between the first node object and the at least one second node object is a deployment relationship, and when the reliability status of the first node object is abnormal, the reliability status of the at least one second node object is also abnormal;
[0045] Alternatively, the topological relationship between the first node object and the at least one second node object is a combination relationship, and when there is a second node object with an abnormal reliability status among the at least one second node object, the reliability status of the first node object is abnormal;
[0046] Alternatively, the topological relationship between the first node object and the at least one second node object is an aggregation relationship, and when the reliability status of the at least one second node object is abnormal, the reliability status of the first node object is abnormal.
[0047] In one possible design, the reliability status of any node object among the multiple node objects is determined based on indicators of the node object; the indicators include at least one of the following indicators: the traffic of the node object, the overload level of the node object, the response time of the node object, and the data error rate of the node object.
[0048] In one possible design, the multiple node objects include resource objects and / or business objects; the business objects are node objects with business processing functions; and the resource objects are node objects that provide resources.
[0049] In one possible design, if at least one indicator of the node object's traffic, overload level, and response time is abnormal, and the node object's data error rate is normal, the node object's reliability status is warning;
[0050] Alternatively, if at least one of the traffic flow, overload level, and response time of the node object is abnormal, the data error rate of the node object is abnormal, and the node object has not lost its node function, the reliability status of the node object is damaged;
[0051] Or, if the overload degree of the node object is abnormal and the node object has not lost its node function, the reliability status of the node object is overload;
[0052] Alternatively, if the traffic, overload level, response time, and data error rate of the node object are all abnormal, and the node object loses its node function, the reliability status of the node object is faulty.
[0053] In one possible design, it also includes:
[0054] A fault analysis graph is sent, wherein the fault analysis graph includes a topology of node objects associated with a root cause of the fault.
[0055] In one possible design, the topology of the node objects associated with the root cause of the fault is determined based on multiple paths, and the multiple paths are determined based on relationship information between the multiple node objects; each of the multiple paths includes at least one node object.
[0056] In one possible design, information of the multiple paths is sent.
[0057] In one possible design, it also includes:
[0058] Information on indicators of the plurality of node objects and diagnostic rules corresponding to the indicators of the plurality of node objects are received, where the indicators of the plurality of node objects and the diagnostic rules are used to determine reliability states of the plurality of node objects.
[0059] In a possible design, obtaining relationship information between multiple node objects includes:
[0060] Relationship information between the multiple node objects is received.
[0061] In a third aspect, a fault diagnosis device is provided, which has the function of implementing the method described in any of the above aspects and any possible implementation thereof. The function can be implemented in hardware, or the corresponding software can be executed by hardware. The hardware or software includes one or more modules corresponding to the above functions.
[0062] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program (also referred to as instructions or code), which, when executed by a fault diagnosis device, causes the fault diagnosis device to perform the method of any of the above aspects or any embodiment of any of the aspects.
[0063] In a fifth aspect, a computer program product is provided. When the computer program product runs on a fault diagnosis device, the fault diagnosis device executes a method of any aspect or any embodiment of any aspect.
[0064] In a sixth aspect, a circuit system is provided, the circuit system including a processing circuit, the processing circuit being configured to execute the method of any aspect or any embodiment of any aspect.
[0065] In the seventh aspect, a chip system is provided, comprising at least one processor and at least one interface circuit, wherein the at least one interface circuit is used to perform transceiver functions and send instructions to the at least one processor. When the at least one processor executes the instructions, the at least one processor executes the method of any aspect or any embodiment of any aspect.
[0066] In an eighth aspect, a fault diagnosis device is provided, comprising a processor and a memory. The memory is configured to store a computer program (also referred to as instructions or code), and the processor is configured to execute the computer program so that the fault diagnosis device performs the method of any of the above aspects or any embodiment of any of the aspects.
[0067] In a ninth aspect, a fault diagnosis system is provided, comprising a first diagnostic device and a second diagnostic device. The first diagnostic device is configured to execute the method described in the first aspect and any possible implementation thereof. The second diagnostic device is configured to execute the method described in the second aspect and any possible implementation thereof.
[0068] In a tenth aspect, a fault diagnosis method is provided, which includes: a first diagnostic device executing the method described in the first aspect and any possible implementation thereof, and a second diagnostic device executing the method described in the second aspect and any possible implementation thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] FIG1A is a schematic diagram of a fault diagnosis scenario provided by related art;
[0070] FIG1B is a schematic diagram of a fault diagnosis scenario provided by related art;
[0071] FIG2 is a schematic diagram of a topological relationship provided in an embodiment of the present application;
[0072] 3A-3B are schematic diagrams of the system architecture provided in an embodiment of the present application;
[0073] Figures 4 and 5 are schematic diagrams of the architecture of the system provided in the embodiments of the present application;
[0074] FIG6 is a schematic diagram of a scenario of a fault diagnosis method provided in an embodiment of the present application;
[0075] FIG7 is a schematic diagram of a flow chart of a fault diagnosis method provided in an embodiment of the present application;
[0076] 8A and 8B are flowchart diagrams of a fault diagnosis method according to an embodiment of the present application;
[0077] 8C and 8D are schematic diagrams of scenarios of the fault diagnosis method provided in an embodiment of the present application;
[0078] FIG9 is a schematic diagram of a business object and a resource object provided in an embodiment of the present application;
[0079] FIG10 is a schematic diagram of an indicator for determining reliability status provided in an embodiment of the present application;
[0080] FIG11 is a schematic diagram of the reliability status of an object provided in an embodiment of the present application;
[0081] FIG12 is a schematic diagram of a scenario of a method for determining a reliability status provided in an embodiment of the present application;
[0082] FIG13 is a schematic diagram of a topological relationship between objects provided in an embodiment of the present application;
[0083] FIG14 is a schematic diagram of a fault analysis diagram provided in an embodiment of the present application;
[0084] 15A-15D are flowcharts of a fault diagnosis method according to an embodiment of the present application;
[0085] 16A-16E are schematic diagrams of scenarios of a fault diagnosis method provided in an embodiment of the present application;
[0086] FIG17 is a schematic diagram of the structure of the device provided in an embodiment of the present application;
[0087] FIG18 is a schematic diagram of the structure of the device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0088] Currently, fault diagnosis can be performed using fault tree analysis (FTA) or knowledge graph methods. Both of these diagnostic methods strongly rely on the relationship between events of an object.
[0089] Among them, in the fault tree analysis method, the cause of the fault (which can be referred to as the root cause of the fault) can be analyzed step by step in a tree-like manner based on the hierarchical characteristics of the fault and the relationship between the cause and result of the fault.
[0090] A fault tree is a diagram that represents logical cause-and-effect relationships. It can include events and logic gates. Events can be used to describe the fault states of systems and components. Logic gates can be used to correlate events and represent the logical relationships between them.
[0091] Events can include top events, intermediate events, or bottom events. A top event (also called a root event) corresponds to an undesirable device fault state, such as excessive CPU usage. Intermediate events: Based on the top event, possible events leading to the top event are queried and analyzed. These possible events include intermediate events. Events directly related to the top event are called first-level intermediate events. For example, event 2 in Figure 1B is a first-level intermediate event.
[0092] A bottom event is the lowest level fault event that leads to a top event. A bottom event can also be called a basic event that leads to a top event. After analysis, as shown in FIG1B , the bottom event that leads to the top event is event 7.
[0093] Fault trees contain the concepts of cut sets and minimal cut sets. A cut set refers to a collection of one or more bottom events in a fault tree. When these events occur simultaneously, the communication system can determine the occurrence of the top event. A minimal cut set is a cut set that, if any one of the bottom events is removed, no longer forms a cut set.
[0094] In the fault tree, the minimum cut set can be matched. If a certain minimum cut set is satisfied and the events contained in the path where the minimum cut set is located all represent faults, the root cause of the fault can be found based on this.
[0095] A fault knowledge graph is a graph-based data structure consisting of nodes and edges. Each node represents an entity in the real world (such as an instantiation of a network element), and each edge represents the relationship between entities.
[0096] In the fault knowledge graph, by defining the conduction relationship between fault events and based on the relationship between events, the health status of the object (such as whether it is a fault) is determined, and the abnormal propagation path is identified to facilitate fault identification and recovery.
[0097] In the fault knowledge graph method or the fault tree analysis method, fault diagnosis depends on the relationship between events. In some cases, if one or more event links are missing among multiple events, the current fault diagnosis method will fail.
[0098] In order to solve the above technical problems and improve the success rate of fault diagnosis, the present application provides a fault diagnosis method. First, the terms involved in the present application are introduced:
[0099] 1. Topological relationships between objects
[0100] The topological relationship between objects includes, but is not limited to, any of the following relationships: composition relationship, aggregation relationship, deployment relationship, redundancy relationship, and upstream and downstream relationship.
[0101] The topological relationship is introduced by taking an example where multiple node objects to be diagnosed include a first node object and at least one second node object.
[0102] At least one second node object is deployed on a first node object, and the topological relationship between the first node object and at least one second node object is a deployment relationship. In other words, a deployment relationship refers to one or more objects being deployed on another object. For example, in Figure 2, service instance 1 of the microservice layer is deployed on POD1 of the container layer, and service instance n is deployed on POD2. For another example, in Figure 2, VM1 and VM2 of the virtual machine (VM) layer are deployed on host 1 of the computing resource layer.
[0103] Or, the first node object includes at least one second node object, and the topological relationship between the first node object and the at least one second node object is a combination relationship. In an embodiment of the present application, in the combination relationship, the life cycle of the first node object is consistent with the life cycle of the second node object that constitutes the first node object. For example, in Figure 2, the service cluster and the load balancing (LB) cluster are combined to form a network function (NF), and the relationship between the NF and these two clusters is a combination relationship. Among them, the service cluster may include one or more services, and the LB cluster may include one or more LBs. When a service fails in the service cluster, the NF fails accordingly.
[0104] Or, the first node object includes at least one second node object, the topological relationship between at least one second node object is a redundant relationship, and the topological relationship between the first node object and at least one second node object is an aggregation relationship. In an embodiment of the present application, in the aggregation relationship, the life cycles of the first node object and the second node object may be inconsistent. For example, in Figure 2, the service cluster of the microservice layer is formed by the aggregation of service instance 1 to service instance n, and there is a redundant relationship between service instance 1 to service instance n. When a service instance (such as service instance 1) fails, since there are still other non-faulty service instances that can provide services, the service cluster may not fail and the service cluster is still available.
[0105] Alternatively, a first node object and at least one second node object have a communication relationship, and the topological relationship between the first node object and the at least one second node object is an upstream-downstream relationship. For example, in Figure 2, host network port 1 of the computing resource layer has a communication relationship with top of rack (TOR) downlink network port 1 of the network resource layer. The topological relationship between host network port 1 and TOR downlink network port 1 can be called an upstream-downstream relationship.
[0106] It should be noted that the hierarchical and topological relationships between objects shown in FIG2 are merely examples.
[0107] 2. Cross-domain management
[0108] Cross-domain management enables unified management of multiple network element management systems. Logical management functions deployed at the cross-domain layer, such as the network slice management function (NSMF), provide various management services, which can be exposed to the public through the EGMF.
[0109] 3. Domain management
[0110] Single-domain management enables management of fifth-generation (5G) base stations or 5G core networks, such as radio access network (RAN) domain management and core network (CN) domain management. Logical management functions deployed at the single-domain layer, such as the network slice subnet management function (NSSMF) and the management data analytics function (MDAF), are used to implement management services for 5G base stations or the 5G core network. These services can be made available to the public through the exposure governance management function (EGMF).
[0111] For example, Figure 3A shows an example of a network architecture. As shown in Figure 3A, the network architecture may include an operation support system (OSS), a network management system, network elements, and a telecom cloud.
[0112] OSS includes a fault model orchestration module, a cross-domain fault recovery system, and a fault visualization module. The fault model orchestration module can be used to perform cross-domain fault diagnosis and recovery. The fault visualization module supports querying and subscribing to the reliability status and fault self-recovery information of network element objects and cloud resource objects, enabling visualization of the fault diagnosis and self-recovery process. After subscription, the diagnosis and self-recovery related interfaces can be visualized. The cross-domain fault recovery system can be used for cross-domain fault diagnosis and recovery.
[0113] In one or more embodiments of the present application, the diagnostic node may also record diagnostic results in a log. The diagnostic node includes one or more of an OSS, a network management system, a network element, or a telecom cloud. In the embodiments of the present application, the diagnostic node may also be referred to as a diagnostic device, a fault diagnosis device, or other names. The diagnostic device may include a first diagnostic device. The diagnostic device may also include a second diagnostic device.
[0114] The network management system can communicate with the OSS. The network management system includes but is not limited to the operation management center (OMC) and MDAF. Optionally, the network management system may include a fault model orchestration module, a fault visualization module, and a cross-network element and cross-layer fault recovery system. Among them, the introduction of the fault model orchestration module and the fault visualization module can refer to the introduction of the corresponding modules in the OSS. The cross-network element and cross-layer fault recovery system can be used for cross-layer or cross-network element fault diagnosis and fault recovery. Cross-layer or cross-network element fault diagnosis can be single-domain fault diagnosis. The cross-network element and cross-layer fault recovery system can also be called a network fault recovery system.
[0115] In the embodiment of the present application, fault recovery may also be referred to as fault self-healing.
[0116] The network management system can be connected and communicated with the network element. The network element can include a network element fault recovery system for single network element fault diagnosis and fault recovery.
[0117] The network management system can also communicate with the telecom cloud. The telecom cloud can include a telecom cloud fault recovery system for fault diagnosis and fault recovery within a single telecom cloud platform.
[0118] In the embodiments of the present application, different levels of fault diagnosis and recovery can be supported. For example, the above-mentioned cross-domain fault diagnosis and recovery, single-domain fault diagnosis and recovery (e.g., cross-network element or cross-layer), single-network element fault diagnosis and recovery, and single-telecom cloud fault diagnosis and recovery. Fault diagnosis other than the single-network element level and single-telecom cloud level can be collectively referred to as network-level fault diagnosis.
[0119] For example, if there is sufficient remaining capacity within a network element (such as an AMF), the microservice with packet loss on the network element can be isolated, and the remaining microservices on the network element can continue to serve, that is, network element-level fault recovery is performed. For another example, if there are many bad points in the network, and the bad points are located at a higher level, such as a switch at a higher level that is broken, network element-level fault recovery can no longer solve the network problem, then network-level recovery can be performed to isolate the entire network element at the bad point. For another example, the network element fault recovery system uses the services within the network element or the redundancy within the network element to perform fault recovery. For another example, when the network element fault recovery system cannot recover, it will be restored by the "network fault recovery system", such as the network fault recovery system recovering through traffic scheduling around the faulty network element.
[0120] For example, FIG3B shows another example of the system architecture of an embodiment of the present application. As shown in FIG3B , the fault recovery system in the OSS, network management, network element, and telecom cloud may include an intelligent fault closed-loop service. Through the intelligent fault closed-loop service, the fault recovery system can select a fault recovery method based on factors such as the number of faults and the severity of the faults, and call the corresponding atomic capabilities to perform fault recovery. For example, when isolating a single module of a single network element can no longer resolve the anomaly, the entire network element can be isolated. For example, supporting the abnormal propagation path of cloud and network collaboration, realizing fault self-healing determined by cloud-network collaboration.
[0121] For example, Figure 4 shows an example of the architecture of the intelligent fault closed-loop service. The architecture of the intelligent fault closed-loop service is shown in Table 1 below:
[0122] Table 1
[0123] Among them, the fault rules may include diagnostic rules and / or recovery rules. The reliability model may include relationship information between node objects. In each embodiment of the present application, the reliability model may also be called an object relationship model, a fault model, a first model, a state relationship model, or other names.
[0124] The embodiment of the present application can define complex fault scenarios (UseCase), which include but are not limited to sub-health scenarios and resource abnormality scenarios. In different scenarios, the indicators that need to be collected may be different to support fault diagnosis in the corresponding scenarios. Compared with the fault tree, it is necessary to arrange the cause and effect relationship, arrangement, conditions, etc. of some fault events, which is highly complex. In the embodiment of the present application, no complex arrangement is required. When a new scenario appears, the corresponding indicator can be specified in the reliability model. The relationship information between node objects remains unchanged. It can be seen that the solution of the embodiment of the present application has a low implementation complexity.
[0125] Atomic capabilities can refer to various fault recovery methods, such as reset, migration, reconstruction, isolation, and repair.
[0126] Optionally, the OSS, network management system, network elements, and telecom cloud fault recovery systems can also obtain corresponding fault models (reliability models) and scenario (use case) information and perform fault diagnosis based on the reliability models and scenario information. For example, the OSS obtains cross-domain fault models and scenario information and performs fault diagnosis accordingly. The network management system obtains cross-network element fault models and cross-layer fault models and scenario information and performs fault diagnosis accordingly. The network elements obtain network element fault models and scenario information and perform fault diagnosis accordingly. The telecom cloud obtains cloud resource fault models and scenario information and performs fault diagnosis accordingly.
[0127] As shown in Figure 3B, the management plane provides a unified management UI, supports multi-version reliability model management, and fault self-healing use case management. It can also be used to deliver adaptation packages to the cloud platform or network elements.
[0128] The cloud platform can be used for model analysis and storage. Based on intelligent fault closed-loop services, the telecom cloud fault recovery system supports fault detection, diagnosis, and self-healing. For example, based on scenario-based fault models, model-driven single-domain fault diagnosis at the resource and business layers is implemented. Fault reporting is also supported. Query and subscription methods are provided to obtain the reliability status of each object, enabling management-plane fault visualization.
[0129] For example, Figure 5 shows a schematic diagram of the connections between logical management functions (such as network management) in a management plane. As shown in Figure 5, the management plane includes NSMF, NSSMF, MDAF, EGMF, communication service management function (CSMF), network function management function (NFMF), and NF.
[0130] Among them, NSMF, NSSMF, MDAF, EGMF, NFMF and NF provide different types of MnS respectively. EGMF can open the MnS provided by various logical management functions to the outside world.
[0131] For a detailed description of the architecture of the network elements involved in this solution, please refer to relevant technologies, such as the reference standard 3GPP TS23.501, and the embodiments of this application will not be repeated here.
[0132] The terms "first" and "second" in the specification and drawings of this application are used to distinguish different objects, or to distinguish different treatments of the same object. Words such as "first" and "second" can distinguish between identical or similar items with substantially the same functions and effects. For example, the first device and the second device are merely used to distinguish different devices and do not limit their order. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity or execution order, and words such as "first" and "second" do not necessarily limit differences.
[0133] "At least one" means one or more, and "a plurality" means two or more.
[0134] "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or plural.
[0135] Furthermore, the terms "including," "having," and any variations thereof, as used in the description of this application are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.
[0136] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being more preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0137] In the embodiments of the present application, when A sends a message to B, A can send it directly to B, or A can send it to B through one or more other devices, such as A sending it to C, which then sends it to B. Optionally, when forwarding messages between different devices, the intermediate devices can process the messages. For example, if A sends a message to C, C processes the message and then sends the processed message (such as a control instruction) to B. The format of the message received by B may be different from the format of the message sent by A.
[0138] For example, as shown in Figure 6, a virtual machine is deployed on a host, and the virtual machine and the host are in a deployment relationship. Service instance 1 and service instance 2 are deployed on virtual machines, and the topological relationship between the virtual machine and the two service instances is also a deployment relationship.
[0139] In some cases, the fault recovery system detects that event A occurs on the host: packet loss on the host, and the fault recovery system also detects that event E occurs on service instance 1: packet loss on service instance 1. Based on the deployment relationship between the host and the virtual machine, the fault recovery system can infer that when event A occurs on the host, the virtual machine deployed on the host also experiences an abnormality accordingly, such as generating event B: packet loss. The fault recovery system can also infer that the cause of event E (packet loss) generated by service instance 1 is packet loss in the virtual machine (event B) based on the deployment relationship between service instance 1 and the virtual machine. In other words, even if there is a missing event link in the event relationship (such as event B), the fault recovery system can deduce the transmission relationship of the abnormal state between objects based on the topological relationship between the objects, supplement the missing event link, and determine the root cause of the fault.
[0140] It can be seen that compared with the related art that relies heavily on the relationship between abnormal events to diagnose faults, which leads to diagnosis failure when there are missing links in the event relationship, the solution of the embodiment of the present application performs fault diagnosis based on the relationship between objects (such as topological relationships), which is applicable to a variety of scenarios. Even if there are missing links in the event relationship, the root cause of the fault can be deduced, thereby improving the success rate of fault diagnosis.
[0141] The following is a detailed description of the method of the embodiment of the present application. FIG7 shows an example of the process of the diagnostic method of the embodiment of the present application. The method may include:
[0142] S101: A diagnosis node obtains relationship information between multiple node objects.
[0143] The relationship information includes the topological relationship between multiple node objects and the reliability status of multiple node objects.
[0144] The reliability status may also be referred to as the health status, or health level, or healthiness, or credibility, or device status, or fault status or other names, and the embodiments of the present application do not limit this.
[0145] As a possible implementation, the diagnostic node may include the above-mentioned fault recovery system. The fault recovery system executes S101. The diagnostic node may include at least one of the following nodes: OSS, or network management, or network element or a node in the telecom cloud (such as a server).
[0146] For example, Figure 8A illustrates an example process for obtaining relationship information between object nodes in S101. In this example, the fault model orchestration capability is deployed on the OSS. The fault model orchestration module in the OSS obtains relationship information between multiple node objects (corresponding to S101a1 in Figure 8A). The OSS network fault recovery system can then obtain this relationship information from the fault model orchestration module (corresponding to S101b1 in Figure 8A).
[0147] As shown in Figure 8A , the OSS can send relationship information between multiple node objects to the network manager. For example, the OSS sends this relationship information to the network manager's fault model orchestration module via the fault model orchestration module (corresponding to S101c1 in Figure 8A ). The network manager's network fault recovery system obtains this relationship information from the fault model orchestration module (corresponding to S101d1 in Figure 8A ). For another example, the OSS sends this relationship information to the network manager via the network fault recovery system.
[0148] As shown in Figure 8A, the network management sends the relationship information of multiple node objects to the network element and the telecom cloud. For example, the network management sends the relationship information through the fault model orchestration module (corresponding to S101e1 and S101f1 in Figure 8A). After the network element receives the relationship information, it can perform fault diagnosis and fault recovery through the network element fault recovery system in combination with the relationship information. Similarly, after the telecom cloud receives the relationship information, it can perform fault diagnosis and fault recovery through the telecom cloud fault recovery system in combination with the relationship information. In Figure 8A, the fault model orchestration module in the network management is optional. In some examples, if the network management includes a fault model orchestration module, the network management can receive the above-mentioned relationship information through the fault model orchestration module, and the network fault recovery system in the network management can obtain the relationship information from the fault model orchestration module.
[0149] As another example, Figure 8B illustrates another example process for obtaining relationship information between object nodes in S101. Fault model orchestration capabilities are deployed in the network management system. The fault model orchestration module in the network management system obtains relationship information between multiple node objects. The network management system's network fault recovery system can obtain this relationship information from the fault model orchestration module. As shown in Figure 8B , the network management system can send this relationship information to network elements and the telecom cloud. In Figure 8B , the fault model orchestration module in the OSS is optional.
[0150] Taking the fault model orchestration capability deployed on OSS as an example, where OSS is used to obtain relationship information between multiple node objects, as a possible implementation method, OSS obtains a reliability model (which may be referred to as a first model). The first model includes the relationship information between the above-mentioned multiple node objects. Exemplarily, the OSS fault model orchestration module loads the established first model or orchestrates the first model online. Afterwards, OSS can generate a model adaptation package (which may be referred to as an adaptation package) based on the first model. The adaptation package contains relevant information in the first model. In some examples, the OSS fault model orchestration module imports the adaptation package into the OSS network fault recovery system.
[0151] As shown in Figure 8C, the OSS can send the adaptation package to the network manager. For example, the OSS fault model orchestration module imports the adaptation package into the fault model orchestration module of each network manager. The network manager's fault model orchestration module imports the adaptation package into the network manager's fault recovery system.
[0152] The NMS can send the adaptation package to the NE and the telecom cloud. For example, the NMS's fault model orchestration module imports the adaptation package into the NE's fault recovery system and into the telecom cloud's fault recovery system.
[0153] In one or more embodiments of the present application, the adaptation packet may also be referred to as a message or other names.
[0154] Taking the fault model orchestration capability deployed in the network management as an example, Figure 8D illustrates another example process for obtaining an adaptation package. The network management's fault model orchestration module loads an established first model or orchestrates the first model online. The network management then generates an adaptation package based on the first model, containing relevant information about the first model. As shown in Figure 8D , the network management's fault model orchestration module imports the adaptation package into the network management's fault recovery system. The network management's fault model orchestration module then imports the adaptation package into the network element's fault recovery system and into the telecom cloud's fault recovery system.
[0155] Optionally, in addition to the relationship information between multiple node objects, the first model may also include at least one of the following information: the type of the node object; an indicator for determining the reliability status; the reliability status; a diagnostic rule for the reliability status; and a transfer relationship between reliability statuses. The following describes each piece of information in the first model:
[0156] 1. Business objects and resource objects
[0157] Depending on the function or capability of the node object, the node object may include a business object and / or a resource object.
[0158] A business object is an object with independent business logic functions, or a node object with business processing capabilities. For example, a business object can be a module within a network element. For example, a business object can have one or more business functions such as forwarding, communication links, user registration, session, and phone calls.
[0159] Resource objects are node objects that can obtain allocated resources from a resource pool. Resource objects can provide resources to other node objects. For example, a resource object can be a host that provides general functions, does not handle business logic, and provides resources to network elements, such as software and operating systems that can be installed on the host.
[0160] For example, as shown in Figure 9, business objects may include LB clusters in AMF, LB instances deployed in PODs, link clusters, service clusters, and database (DB) clusters. These objects are used to support AMF business processing. Business objects may also include corresponding modules (or instantiated objects) in SMF, UDM, and IMS.
[0161] As shown in Figure 9, resource objects may include hosts, virtual machines, POD instances, storage, and network resources. For example, these resource objects can be used to process the resources required by network elements such as AMF for business processing.
[0162] For example, taking the 5G core network (5GC) network element as an example, the corresponding modules in the core network element can be defined as resource objects or business objects. Assuming that the NF shown in Figure 2 is AMF, as shown in Figure 2, in some examples, the modules in the AMF located at the network element layer and microservice layer are business objects, used by the AMF to process services. For example, the AMF, service cluster, LB cluster, service instance, etc. are business objects. The modules in the AMF located at the container layer, virtual machine layer, computing resource layer, and network resource layer are resource objects, used to provide the resources required for the AMF to process services.
[0163] It's important to note that resource objects and business objects are not mutually exclusive concepts. In some cases, the same module can be both a resource object and a business object. As shown in Figure 2, the PODs at the container layer and the Fabric plane can be both business objects and resource objects, enabling them to process business and provide the resources needed for that business.
[0164] 2. Indicators for determining reliability status
[0165] The indicator of a node object, which may also be called a status indicator, a device indicator, etc., is used to determine the reliability status of the node object.
[0166] Optionally, the indicators include the following indicators: traffic of the node object, saturation of the node object, response time of the node object, and data error rate of the node object.
[0167] The traffic of a node object can represent the amount of data. The traffic can be the traffic flowing into the node object and / or the traffic flowing out of the node object. For example, the traffic is related to the request rate (number of requests per second).
[0168] The response time of a node object, also known as latency, can include the time data spends in a queue or waiting to be processed. For example, response time is measured in milliseconds. For example, as shown in Figure 10, for receiving data, response time can refer to the time between the data entering the node object and the processed data leaving the node object.
[0169] Data errors in a node object can be caused by errors in the data it generates (such as processing errors) or errors in the processing of received data. These erroneous data can be used by the node object itself or sent to other node objects for use. The data error rate of a node object affects the reliability status of the node object. For example, when the node object has sufficient resources, the data error rate is low, the reliability of the node object is high, and the node object works normally. For example, when a node object processes data (such as received data or data generated by itself), data errors may occur due to insufficient resources or other reasons, and the reliability of the node object is low.
[0170] The degree of overload of a node object, also known as saturation, is related to the resource utilization of the node object. For example, the degree of overload can be determined by the queue depth (or concurrency level). For example, when a node object approaches overload, the queue depth of the node object changes from zero to non-zero. For example, the degree of overload can be implemented using a counter.
[0171] In other embodiments, different indicators can be collected for different types of node objects, so that indicators applicable to the corresponding node objects are collected for the corresponding node objects, thereby accurately determining the reliability status of the corresponding node objects. Alternatively, the above indicators can be refined into different sub-indicators for different types of node objects, and the reliability status of the node objects can be determined based on the sub-indicators of the node objects.
[0172] For example, for a business object, the data error rate can be determined based on the sub-indicator of the business object's key performance indicator (KPI) success rate, and the degree of overload can be determined based on the resource utilization rate of the business software. For example, an instance (an example of business software) can handle the business of 50,000 terminals and apply for a corresponding amount of memory. When the instance has used 80% of the memory, it can be determined that the instance is overloaded. In one or more embodiments of the present application, the key KPI success rate of the node object can also be replaced by a service KPI.
[0173] For example, for resource objects, the overload level can be determined based on resource occupancy, resource application success rate, and storage input / output (IO) latency, while the data error rate can be determined based on communication success rate and packet loss rate. For example, if a host's CPU occupancy is high, the success rate of business objects requesting CPU resources from that host may decrease. In this case, the host is considered overloaded.
[0174] Optionally, the resources include but are not limited to at least one of the following resources: computing, storage, and network.
[0175] Of course, the indicators used to determine the reliability status are not limited to the above-mentioned ones. For example, the reliability status of the node object can also be determined in combination with the communication link information between the node object and other node objects.
[0176] 3. Reliability status
[0177] Optionally, the objects defined in the reliability model are objects for which indicators can be obtained and reliability states can be defined. Optionally, the objects can have replaceable units for self-healing, and the replaceable units are in a redundant relationship with the objects.
[0178] The reliability status of a node object can be normal or abnormal. Based on the evolution and severity of the diagnosed abnormality, as shown in Figure 11, the abnormality can include any of the following states: warning, damage, overload, and failure.
[0179] Warning: When a node object is in this state, its node function is not affected, but there is a certain risk. For business objects, the node function refers to the business processing function; for resource objects, the node function refers to the function of providing resources to other node objects.
[0180] In other words, for a business object, a warning state means that the service processing function of the business object is still normal. For a resource object, a warning state means that the resource object can still provide resources to other objects normally. For example, the usage rate of the service session table or roaming number resources is abnormal. Another example is that the service volume does not match the traffic model. In another example, the load of a node object exceeds the resource threshold but has not yet affected the service, requiring operation and maintenance (such as capacity expansion).
[0181] Impaired: Also known as a sub-healthy state. When a node object is in this state, its functionality is impaired, but the object is not completely non-functional. For business objects, an impaired state means a decrease in their ability to process business operations. For resource objects, an impaired state means a decrease in their ability to provide resources to other objects.
[0182] In some cases, a damaged node object causes a decrease in the KPI of the network in which the node object resides. In other cases, the functionality of a node object in a specific data network (DN) or cell is impaired, and while the KPI of the network in which the node object resides does not significantly decrease, the fault recovery system can generate an alarm for the node object.
[0183] Overload: Due to the heavy load (or overload) on a node object, node functionality is impaired, but the node object does not completely lose functionality. For service objects, an overload state means that the service object's load exceeds the threshold and its service processing capacity is reduced. For resource objects, an overload state means that the resource object's load exceeds the threshold and its ability to provide resources to other objects is reduced. For example, if a network element's load exceeds the threshold, the network element is overloaded.
[0184] Failure: This refers to a state where a node object completely loses its functionality. For business objects, a failure means that the business object is no longer capable of processing business operations. For resource objects, a failure means that the resource object is no longer able to provide resources to other objects.
[0185] The reliability status may also include more or fewer states. For example, the fault status may be further divided into: primary fault, severe fault. The embodiment of the present application does not limit the definition method and number of the reliability status.
[0186] The following describes the relationship between reliability status and indicators (such as traffic, overload level, etc.):
[0187] Optionally, as shown in Table 2, when the traffic, overload level, response time, and data error rate of the node object are all normal, the reliability status of the node object is normal. A certain indicator is normal, which may mean that the value of the indicator is within the allowable range. Conversely, if the value of the indicator exceeds the allowable range, the indicator is abnormal. For example, if the traffic of the node object is within the traffic range, the overload level is within the load range, the response time is within the delay range, and the data error rate is within the fault tolerance range, the reliability status of the node object is normal. If there are abnormal indicators in the traffic, overload level, response time, and data error rate of the node object, the reliability status of the node object is abnormal.
[0188] Table 2
[0189] Or, optionally, as shown in Table 2, when at least one indicator of the node object's traffic, overload level, and response time is abnormal and the node object's data error rate is normal, the node object's reliability status is warning.
[0190] Or, optionally, as shown in Table 2, at least one indicator of the node object's traffic, overload level, and response time is abnormal, the node object's data error rate is abnormal, and the node object has not lost its node function, the node object's reliability status is damaged.
[0191] Or, optionally, as shown in Table 2, when the overload degree of the node object is abnormal and the node object has not lost its node function, the reliability status of the node object is overload.
[0192] Or, optionally, as shown in Table 2, when all indicators of the node object are abnormal and the node object has not lost its node function, the reliability status of the node object is faulty.
[0193] The above is only one example of how the four indicators affect the reliability status. In other embodiments, the four indicators can also be used in different combinations to affect the reliability status. For example, the damage status can be defined as: abnormal response time and abnormal at least one of the indicators of traffic flow, overload level, and data error rate. Alternatively, the reliability status can be indicated by more or fewer indicators without limitation. For example, the reliability status can be determined based on the response time, traffic flow, and overload level. For another example, the reliability status can be determined based on the response time, traffic flow, overload level, and link status.
[0194] In some embodiments, the indicator of a node object includes an indicator corresponding to at least one of the functions, resources, and data of the node object. In other words, the indicator of the node object can be determined based on the indicator corresponding to at least one of the functions, resources, and data of the node object. Optionally, the function of the node object includes at least one of the following functions: business function, communication function, and resource provision function. Exemplarily, the communication function can be communication within a network element or communication between network elements. Exemplarily, the data includes routing data, such as routing data within a network element or routing data between network elements.
[0195] For example, taking a business object as a node object, Figure 12 illustrates the relationship between the indicators corresponding to functions, resources, and data, and the reliability status of the business object. As shown in Figure 12, at least one of the business functions and business communications is associated with a fault state. For example, if the success rate of reading and writing data for a business object exceeds the allowable range, and the failure rate of communication between the business object and other node objects exceeds a threshold, the business object is considered faulty.
[0196] Exemplarily, as shown in FIG12 , at least one of the service function, service communication, and service data is related to the impairment state. Exemplarily, as shown in FIG12 , the service resource is related to the overload state. Exemplarily, as shown in FIG12 , the service resource is related to the warning state.
[0197] 4. Diagnostic rules for reliability status
[0198] The diagnostic rules and the indicators of the plurality of node objects are used to determine the reliability status of the plurality of node objects.
[0199] Exemplarily, the diagnostic rules are defined in the first model: if the success rate of SMF 5G standalone architecture (SA) session establishment drops by more than 1%, and the number of SMF 5G SA session establishment requests is greater than 200, then the SMF is determined to be service function damaged, and the reliability status of the SMF is correspondingly determined to be a damaged state.
[0200] In an embodiment of the present application, a diagnostic node can receive information about indicators of multiple node objects and diagnostic rules corresponding to the indicators of multiple node objects, and determine the reliability status of the multiple node objects based on the indicators and diagnostic rules of the multiple node objects. For example, the diagnostic node can generate an instantiated reliability model based on the network topology of the system. Afterwards, the diagnostic node collects information about one or more of the above indicators based on the reliability model, performs fault diagnosis in combination with the diagnostic rules, and updates the reliability model to a twin model.
[0201] 5. Transmission relationship between reliability states
[0202] In an embodiment of the present application, based on the topological relationship between multiple node objects and the reliability status of the multiple node objects, the transmission relationship of the reliability status between the multiple node objects can be derived to form an abnormal propagation path so as to perform root cause judgment of fault diagnosis.
[0203] The abnormal propagation path includes one or more node objects. There is at least one abnormal node object in the abnormal propagation path. Taking the fault analysis diagram in Figure 14 as an example, there is an abnormal Fab1 in path 2, and path 2 is an abnormal propagation path.
[0204] In one or more embodiments of the present application, the exception propagation path may also be referred to as an exception propagation chain, or an exception path, or an exception chain, or other names.
[0205] Optionally, at least one second node object is deployed on the first node object, and the topological relationship between the first node object and the at least one second node object is a deployment relationship. When the reliability status of the first node object is abnormal, the reliability status of at least one second node object is abnormal. For example, Table 3 shows an example of deriving a transfer relationship of reliability status based on the topological relationship and reliability status between node objects. As shown in Table 3, a, b, and c are deployed on A. When A is abnormal, it can be deduced that a, b, and c deployed on A are also affected and abnormal.
[0206] Alternatively, if the topological relationship between a first node object and at least one second node object is a composite relationship, and if one of the at least one second node objects has an abnormal reliability status, the reliability status of the first node object is abnormal. As shown in Table 3, in a composite relationship, a, b, and c form A. Because the lifecycles of the local and overall components of a composite relationship are consistent, any abnormal node object among a, b, and c can affect the reliability status of A, causing A to become abnormal. If a, b, and c are all normal, A is normal or its reliability status remains unchanged.
[0207] Alternatively, if the topological relationship between a first node object and at least one second node object is an aggregation relationship, and the reliability status of at least one second node object is abnormal, the reliability status of the first node object is abnormal. As shown in Table 3, a, b, and c are aggregated to form A. When a, b, and c are all abnormal, A is inferred to be abnormal. Furthermore, in Table 3, the aggregation relationship of a or b or c abnormal -> A possibly abnormal can mean: if a or b or c is a warning, A is inferred to be a possible warning; if a or b or c is damaged, A is inferred to be possibly damaged; if a or b or c is overloaded, A is inferred to be possibly overloaded; if a or b or c is faulty, A is inferred to be possibly faulty.
[0208] Table 3
[0209] In one or more embodiments of the present application, as shown in Figure 13 , a composition relationship can be a composition relationship between business objects or a composition relationship between resource objects. For example, multiple business objects B form business object C. For another example, multiple resource objects D form resource object E. This embodiment of the present application does not impose any restrictions on this.
[0210] Similarly, in one or more embodiments of the present application, as shown in FIG13 , the aggregation relationship may be a relationship between business objects or between resource objects, and the embodiments of the present application do not impose any restrictions on this.
[0211] Similarly, in one or more embodiments of the present application, as shown in FIG13 , a deployment relationship can be a relationship between a business object and a resource object, or between resource objects. For example, a business object is deployed on a resource object. This embodiment of the present application does not limit this.
[0212] Similarly, in one or more embodiments of the present application, as shown in FIG13 , the upstream and downstream relationship may be, for example, a message passing relationship between business objects or a communication relationship between resource objects, but the present application does not limit this.
[0213] In one or more embodiments of the present application, the redundancy relationship may be, for example, a relationship between business objects or between resource objects, but the present application does not impose any limitation on this.
[0214] It should be noted that as the network evolves, more topological relationships between node objects can be defined, and this embodiment of the application does not limit this. For example, when node object A is abnormal, node object B is also abnormal, and the topological relationship between objects A and B can be defined as relationship A'.
[0215] After introducing the information contained in the first model, the following examples illustrate the format of some information in the first model:
[0216] Optionally, the node objects and topological relationships may be defined using Table 4. As shown in Table 5, the topological relationship between the SM2-cluster and the VNF may be a combination relationship.
[0217] Optional, such as the relationship additional attributes in Table 4, refer to some attributes related to the topological relationship, such as the resource utilization of the instantiated node object, which can be displayed through the interface. The relationship additional attributes can be, for example, reserved fields.
[0218] Table 4
[0219] For example, an example of the information in Table 4 is as follows:
[0220] Based on the above data, the topological relationship between SM2-Cluster and VNF is a composite relationship. The topological relationships between other object nodes can also be defined in a similar way.
[0221] Optionally, the definitions in Table 5 may be used to define indicators, custom data, etc. for determining reliability status.
[0222] Table 5
[0223] For example, multiple indicators to be collected can be separated by #, such as 192…126#192…128. For example, the metricID of the same node object is unique, so the specific indicator to be collected for the node object can be distinguished based on the metricID.
[0224] For example, the unit of the collection period can be seconds (s) or minutes. Other periods are also possible. For example, indicators can be collected every 60 seconds, or every 5 minutes. For example, multiple periods can be separated by a comma (,), for example, 5, 60, 300.
[0225] Illustratively, the value range of the network management object ID is 0 to 4294967294.
[0226] For example, the node object type can be set to be valid only for a "1-minute collection period" and not for other collection periods. For example, the moiID of the node object is consistent with the moiID used when marking the node object, so as to collect indicators. For example, for a POD, the moiID value indicates the POD type, and the specific instance ID can be matched based on the moiID.
[0227] For example, an example of the information in Table 5 is as follows:
[0228] Optionally, Table 6 can be used to define diagnostic rules for reliability status, such as the rule files required to determine the reliability status, the parameters (or indicators) used by the rules, and the fault diagnosis and recovery scenarios (use cases). Use cases are also called fault diagnosis and recovery use cases.
[0229] Table 6
[0230] The reliability status is calculated based on one or more metrics of a node object. For example, the health_degree of a node is calculated based on traffic, data error rate, overload level, and response time. Another example is the health_degree_service_KPI, which is calculated based on the node object's service KPI. Another example is the health_degree_service_res, which is calculated based on the node object's service resources (compute, storage, and network resources). Another example is the health_degree_compute, which is calculated based on the node object's compute resources.
[0231] Status description, such as damage caused by overload, business sub-health, or process reset.
[0232] The rule file includes one or more diagnostic rules. For example, the rule_args in the rule file can be represented by a json string. For example, the rule file can be a text file or a file in another format, which is not limited in the present embodiment.
[0233] For example, when rule_args includes pod_kpi_succ_rate > 99, the metric to be determined, rule_data, is pod_kpi_succ_rate. In some examples, metricID can be used to indicate pod_kpi_succ_rate, such as using wholeSystem_192…92 as the metricID for pod_kpi_succ_rate. As shown in Table 4, the corresponding rule_data for the SMF can be {kpi:{wholeSystem_192…92,wholeSystem_192…77},alarms:[100339,xxx],custom_metrics:[xx1,xx2]}. Among them, wholeSystem_192…92 can be used to indicate the pod_kpi_succ_rate parameter in the SMF's rule_args, and wholeSystem_192…77 can be used to indicate that pod_kpi_msg_count in the rule_args is greater than 200.
[0234] For example, the UCID is smf_24.0_fabric_unhealth, the node object type is SMF, the reliability status is health_degree, the status description is: damage caused by communication sub-health, rule_args is: {kpi:{pod_kpi_succ_rate>99,pod_kpi_msg_count>200}alarm:{alarmed:xx,alarmid:xx}}, and top_fault is 1.
[0235] For example, the UCID is smf_24.0_fabric_unhealth, the node object type is SM2-cluster, the reliability state is health_degree, the state description is: SM2 cluster fault, the rule file is: sm2_cluster.txt, rule_args is: {process_state: fault}, and top_fault is 0.
[0236] For example, the UCID is smf_24.0_fabric_unhealth, the node object type is SM2-cluster, the reliability status is health_degree_service_KPI, the status description is: SM2 cluster service KPI is completely lost, the rule file is: sm2_cluster.txt, rule_args is: {kpi:{pod_kpi_succ_rate>99,pod_kpi_msg_count>200}…, and top_fault is 0.
[0237] For example, an example of the information in Table 6 is as follows:
[0238] S102: The diagnosis node performs fault diagnosis on multiple node objects based on relationship information between the multiple node objects.
[0239] For example, as shown in Figure 6, a virtual machine is deployed on a host, and the virtual machine and the host are in a deployment relationship. Service instance 1 and service instance 2 are deployed on virtual machines, and the topological relationship between the virtual machine and the two service instances is also a deployment relationship.
[0240] In some cases, the fault recovery system detects that event A occurs on the host: packet loss on the host, and determines that the reliability status of the host is a fault. The fault recovery system also detects that event E occurs on service instance 1: packet loss on service instance 1, and determines that the reliability status of service instance 1 is a fault. Based on the deployment relationship between the host and the virtual machine, the fault recovery system can infer that when event A occurs on the host, the virtual machine deployed on the host also experiences an abnormality accordingly, such as generating event B: packet loss. The fault recovery system can also infer that the cause of event E (packet loss) generated by service instance 1 is packet loss in the virtual machine (event B) based on the deployment relationship between service instance 1 and the virtual machine. In other words, even if there is a missing event link in the event relationship (such as event B), the fault recovery system can deduce the transmission relationship of the reliability status between objects based on the topological relationship between the objects, supplement the missing event link, and determine the root cause of the fault.
[0241] The solution of the embodiment of the present application performs fault diagnosis based on the relationship between objects (such as topological relationship), and is applicable to a variety of scenarios. Even if there are missing links in the event relationship, the root cause of the fault can be deduced, thereby improving the success rate of fault diagnosis.
[0242] In some embodiments, S101 may be implemented as follows: the diagnosis node generates a fault analysis graph based on the relationship information between multiple node objects, and determines the fault diagnosis result according to the fault analysis graph. The fault analysis graph is introduced as follows:
[0243] Among them, the fault analysis graph includes the topology of node objects associated with the root cause of the fault. Topology is also called a topology graph. Figure 14 shows an example of a fault analysis graph, which includes path 6. The node objects included in path 6 are: Fab1, POD8, SM2_0POD, SM service, SMF. Among them, Fab1, POD8, SM2_0POD, and SMF are all in a damaged state. Based on this, the diagnosis node can infer that the node objects associated with the root cause of the fault include Fab1, which means that Fab1 is one of the root causes of system damage. In Figure 14, the topology of node objects associated with the root cause of the fault includes path 6. In some examples, the diagnosis node can send a fault analysis graph. For example, when the diagnosis node is not sufficient to support fault diagnosis, the diagnosis node can report the fault analysis graph to the superior node, and the superior node can perform fault diagnosis based on the fault analysis graph. For another example, the diagnosis node reports the fault analysis graph, and the superior node presents the fault analysis graph in an interface.
[0244] Optionally, the topology of the node objects associated with the root cause of the fault is determined based on multiple paths in the fault analysis graph, and the multiple paths are determined based on the relationship information between the multiple node objects; each of the multiple paths includes at least one node object. The reliability states of the node objects on the path are related, and the reliability state of the node object can affect the reliability state of the adjacent node objects. Exemplarily, the transmission relationship of the reliability state is shown in Table 3 above. Still as shown in Figure 14, the fault analysis graph includes multiple paths, such as paths 1-9. The diagnostic node can determine the topology of the node objects associated with the root cause of the fault from the multiple paths. For example, it is determined that the node objects associated with the root cause of the fault include Fab1, and the topology associated with Fab1 is determined.
[0245] Optionally, the fault analysis graph may further include: information of the aforementioned multiple paths. In some examples, the diagnosis node may send a fault analysis graph including the aforementioned multiple path information.
[0246] As a possible implementation manner, the diagnosis node determines the topology of the node objects associated with the root cause of the fault according to the reliability status of the node objects on each path among the multiple paths.
[0247] Optionally, the diagnosis node calculates the reliability status of the path according to the reliability status of each node object on the path.
[0248] Optionally, the reliability status can be evaluated using a score. Optionally, the higher the reliability score of a node object (which may be referred to as the reliability score), the higher the reliability of the node object and the lower the probability of an abnormality of the node object. Optionally, the scores corresponding to normal, warning, damaged, overloaded, and fault states decrease in descending order.
[0249] In some examples, the diagnostic node obtains multiple links based on the topological relationship between node objects, where the reliability status scores of Fab1, POD8, SM2_0POD, SM service, and SMF on link 6 are respectively Score1-Score5. Score1-Score5 can be multiplied to obtain the reliability score of path 6. Of course, Score1-Score5 can also be subjected to other operations to obtain the reliability score of path 6, such as weighted summation of Score1-Score5. Similarly, the diagnostic node scores other links and determines the root cause of the fault based on the reliability scores of multiple links. The embodiments of the present application do not limit the specific algorithm.
[0250] In some examples, the root cause of the fault can also be determined by combining the convergence relationship between different paths. A convergence relationship can mean that a node object is located on at least two abnormal propagation paths. Still referring to Figure 14, the SM service is on paths 6 and 7. Paths 6 and 7 are both abnormal propagation paths where abnormal objects exist. The SM service aggregates abnormal paths including paths 6 and 7. In some examples, the more abnormal paths a node object aggregates, the greater the probability that the node object is associated with the root cause of the fault, and the node object can be determined as the root cause of the fault.
[0251] Optionally, in the fault analysis diagram, different reliability states can correspond to different UI styles. For example, different reliability states are marked with different colors, such as normal object nodes marked green, warning object nodes marked yellow, damaged object nodes marked orange, and faulty object nodes marked red. For example, in Figure 14, the nodes outlined by the dashed box of Style 1 are damaged nodes, and the nodes outlined by the dashed box of Style 2 are faulty nodes.
[0252] Optionally, in some embodiments, the diagnosis node may transmit the relationship information of the multiple node objects. For example, when the diagnosis node itself cannot independently complete the fault diagnosis, it may report the relationship information between the multiple node objects so that the upper node can perform fault diagnosis on the multiple node objects based on the relationship information.
[0253] Optionally, in some embodiments, the diagnostic node may also receive fault diagnosis results. For example, the network element reports the relationship information of the multiple node objects to the network manager, and the network manager performs fault diagnosis based on the relationship information (and may also combine other information) and returns the fault diagnosis results to the network element.
[0254] Some examples of fault diagnosis scenarios are given below.
[0255] Scenario 1: The diagnostic node includes the network element, which performs fault diagnosis
[0256] As shown in FIG15A , after the network element executes S102 , it may execute the following steps:
[0257] S201. The network element reports a fault analysis diagram to the network management system.
[0258] In some embodiments, after generating the fault analysis graph, the network element reports the fault analysis graph to the network manager. Exemplarily, the network element fault recovery system of the network element reports the fault analysis graph to the network manager.
[0259] S202. The network element reports the diagnosis result to the network management.
[0260] Optionally, the diagnosis result includes the root cause of the fault. The diagnosis result may also include the topology (such as the path) related to the root cause of the fault.
[0261] Exemplarily, the network element fault recovery system reports the diagnosis result to the network management.
[0262] Optionally, in various embodiments of the present application, the fault analysis diagram and the diagnosis result are reported in the same message or different messages.
[0263] Optionally, the format of the reported fault analysis diagram and / or diagnosis results is as follows:
[0264] S203: The network management system presents a fault analysis diagram.
[0265] Exemplarily, the fault visualization module of the network management presents a fault analysis diagram.
[0266] S204: The network management system presents the diagnosis result.
[0267] Exemplarily, the fault visualization module of the network management presents the fault diagnosis result.
[0268] S205: The network manager reports the fault analysis diagram to the OSS.
[0269] Exemplarily, the fault visualization module of the network management reports the fault analysis diagram generated by the network element to the OSS.
[0270] S206: The network manager reports the diagnosis result to the OSS.
[0271] Exemplarily, the fault visualization module of the network management reports the fault diagnosis result generated by the network element to the OSS.
[0272] S207. OSS presents a fault analysis diagram.
[0273] Exemplarily, the fault visualization module of the OSS presents the fault analysis diagram.
[0274] S208. OSS presents the diagnosis result.
[0275] Exemplarily, the fault visualization module of the OSS presents the fault diagnosis result.
[0276] In this scenario, the network element generates a fault analysis diagram and fault diagnosis results. For example, the SMF can generate a fault analysis diagram as shown in Figure 14 and determine the diagnosis results based on the fault analysis diagram, such as determining that the root cause of the SMF damage is damage to POD8 on path 6.
[0277] Scenario 2: The diagnosis nodes include the telecom cloud nodes, and the telecom cloud performs fault diagnosis
[0278] As shown in FIG15B , after the telecom cloud executes S102 , the following steps may be performed:
[0279] S301. Telecom Cloud reports the fault analysis diagram to the network management.
[0280] In some embodiments, after generating the fault analysis diagram, the telecom cloud reports the fault analysis diagram to the network management system. Exemplarily, the telecom cloud fault recovery system of the telecom cloud reports the fault analysis diagram to the network management system.
[0281] S302. Telecom Cloud reports the diagnosis results to the network management.
[0282] Exemplarily, the telecom cloud fault recovery system reports the diagnosis results to the network management.
[0283] For the detailed implementation of S303-S308, please refer to the relevant description of Figure 15A (such as S202-S208).
[0284] Scenario 3: When the diagnosis nodes include the network management, network elements, and telecom cloud, the network management can perform fault diagnosis.
[0285] As shown in FIG15C , the cross-layer (or cross-network element) fault diagnosis performed by the network management may include the following steps:
[0286] S401. The network element reports original diagnostic data to the network management system.
[0287] In some embodiments, if the network element fault recovery system of the network element determines that it cannot complete the fault diagnosis alone, it may report the original data of the fault diagnosis to the network manager so that the network manager can perform the fault diagnosis.
[0288] The original diagnostic data includes the relationship information between multiple node objects and other information used to perform fault diagnosis.
[0289] Optionally, the NE may also report a fault analysis diagram to the NMS. For example, if the NE obtains a fault analysis diagram through preliminary analysis and the diagram contains multiple paths, but the NE cannot determine the root cause of the fault based on the multiple paths, the NE may report the fault analysis diagram to the NMS.
[0290] S402. The telecom cloud reports the original diagnosis data to the network management system.
[0291] In some embodiments, if the telecom cloud fault recovery system of the telecom cloud determines that it cannot complete the fault diagnosis alone, it can report the original fault diagnosis data to the network management system.
[0292] Exemplarily, the original diagnosis data reported by the network element includes at least one of the following pieces of information: the reliability model of the service layer; the reliability status of the service object. Exemplarily, the reliability model includes the instances of each service object and the relationship instances between each service object.
[0293] Exemplarily, the original diagnosis data reported by the telecom cloud includes at least one of the following pieces of information: the reliability model of the resource layer; the reliability status of the resource object. Exemplarily, the reliability model includes the instances of each resource object and the relationship instances between each resource object.
[0294] An example of the format of the original diagnosis data is as follows:
[0295] <NE id = "1" name = "XXX" version = "XX">
[0296] <Obj id = "1" name = "SMF network element">
[0297] <Status = "damaged">
[0298] <Obj id = "2" name = "Service Cluster">
[0299] <Status = "damaged">
[0300] <Obj id = "3" name = "Service POD">
[0301] <Status = "damaged">
[0302] …
[0303] <Relation id = "1" DstObjId = 1 SrcObjId = 2 Relation = "Composition">
[0304] <Relation id = "2" DstObjId = 2 SrcObjId = 3 Relation = "Aggregation">
[0305] …
[0306]
[0307] Optionally, the telecom cloud can also report the fault analysis diagram obtained through preliminary analysis to the network management.
[0308] S102: The network management performs fault diagnosis on the multiple node objects based on the relationship information between the multiple node objects.
[0309] Optionally, raw diagnostic data can be synchronized between network management systems.
[0310] Exemplarily, the network fault recovery system of the network management performs fault diagnosis.
[0311] Optionally, the network management performs fault diagnosis based on the above relationship information and the preliminary fault analysis diagram.
[0312] S403: The network management system presents a fault analysis diagram.
[0313] Exemplarily, the network fault recovery system of the network management reports the fault analysis diagram to the fault visualization module of the network management, and the fault visualization module presents the fault analysis diagram.
[0314] S404: The network management system presents the diagnosis result.
[0315] Exemplarily, the network fault recovery system of the network management reports the fault diagnosis result to the fault visualization module of the network management, and the fault visualization module presents the fault diagnosis result.
[0316] For the specific implementation of S405-S410, reference may be made to the description of the relevant steps (such as S202-S208) of the relevant embodiments.
[0317] Scenario 4: The diagnostic node includes OSS. If the network management cannot complete the fault diagnosis, the OSS will perform the fault diagnosis.
[0318] As shown in Figure 15D, OSS performs fault diagnosis, which may include the following steps:
[0319] S501: The network management reports original diagnostic data to the OSS.
[0320] The original diagnostic data includes relationship information between multiple node objects.
[0321] For example, if the network fault recovery system of the network management determines that it is unable to complete the fault diagnosis alone, it may report the original data of the fault diagnosis to the OSS so that the OSS performs the fault diagnosis.
[0322] Optionally, the network manager reports the fault analysis diagram obtained through preliminary analysis to the OSS.
[0323] S102: OSS performs fault diagnosis on multiple node objects based on relationship information between the multiple node objects.
[0324] Exemplarily, the network fault recovery system of the OSS performs fault diagnosis.
[0325] Optionally, the OSS performs fault diagnosis based on the above relationship information and the preliminary fault analysis diagram reported by the network manager.
[0326] S502: The OSS network fault recovery system sends a fault analysis diagram to the fault visualization module.
[0327] S503: The OSS network fault recovery system sends the diagnosis result to the fault visualization module.
[0328] For the specific implementation of S504-S505, reference may be made to the relevant steps (such as S207-S208) of the relevant embodiments.
[0329] Figure 16A shows an example of the diagnostic data reporting process when the fault visibility capability is deployed on the OSS. The diagnostic data includes fault analysis diagrams and / or diagnostic results. Figure 16B shows another example of the diagnostic data reporting process when the fault visibility capability is deployed on the network management system.
[0330] Figure 16C shows an example of the original diagnostic data reporting process of the network fault recovery system deployed on the OSS. The solution of this example supports cross-domain fault diagnosis. Exemplarily, the cross-domain fault diagnosis data includes: cross-layer fault diagnosis data and cross-network element fault diagnosis data.
[0331] FIG16D shows an example of a process for reporting original diagnostic data when a network fault recovery system is deployed in a network management system. This example solution supports cross-layer / cross-network element fault diagnosis.
[0332] Figure 16E shows an example process for synchronizing raw diagnostic data between network management systems. This process collects office-level and element-level KPIs and communication link status for each network element within the network. By leveraging the reliability status relationships, abnormal paths and network elements can be identified. In this process, MDAFs can synchronize the following data: network-level reliability models, such as object instances and relationship instances, and the real-time status of network-level objects.
[0333] Alternatively, corresponding to the above scoring method, the starting point of the anomaly propagation path can be selected and the root causes can be sorted based on the scores of the anomaly propagation paths corresponding to the multiple starting points. Alternatively, the convergence point of the anomaly propagation path can be selected and the root causes can be sorted based on the scores of the anomaly propagation paths associated with the multiple convergence points.
[0334] In some embodiments, after performing fault diagnosis, the diagnostic node may also perform fault recovery to resolve the fault problem and resume normal operation.
[0335] As a possible implementation, the diagnostic node can rank suspected root causes based on the scores of the multiple anomaly propagation paths described above. Based on this ranking, the diagnostic node can determine a replacement unit for fault self-recovery. This replacement unit can also be referred to as a recommended self-healing object. For example, it may be recommended to replace node object A with node object A', which has a redundant relationship with node object A.
[0336] It should be noted that the above-mentioned multiple embodiments can be combined and the combined scheme can be implemented. Optionally, some operations in the process of each method embodiment are optionally combined, and / or the order of some operations is optionally changed. In addition, the execution order between the steps of each process is only exemplary and does not constitute a limitation on the execution order between the steps. There can also be other execution orders between the steps. It is not intended to indicate that the execution order is the only order in which these operations can be performed. Ordinary technicians in this field will think of many ways to reorder the operations in this article. In addition, it should be pointed out that the process details involved in a certain embodiment of this article are also applicable to other embodiments in a similar manner, or different embodiments can be used in combination.
[0337] In addition, some steps in the method embodiment may be equivalently replaced with other possible steps. Alternatively, some steps in the method embodiment may be optional and may be deleted in certain usage scenarios. Alternatively, other possible steps may be added to the method embodiment. Alternatively, the execution entities (such as functional modules) of some steps in the method embodiment may be replaced with other execution entities.
[0338] Furthermore, the above method embodiments may be implemented separately or in combination.
[0339] Some other embodiments of the present application provide a device, which may be the above-mentioned diagnostic node, etc. The device may include: a memory and one or more processors. The memory and the processor are coupled. The memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the device can perform the various functions or steps in the above-mentioned method embodiments. The structure of the device can refer to the device (device) shown in Figure 17. As shown in Figure 17, the device includes at least one processor 501 and a memory 503. Optionally, the memory 503 may also be included in the processor 501.
[0340] The processor 501 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application.
[0341] The memory 503 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including a compact disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may exist independently and be connected to the processor via a communication line. The memory may also be integrated with the processor.
[0342] The memory 503 is used to store computer-executable instructions for implementing the solution of the present application, and the execution is controlled by the processor 501. The processor 501 is used to execute the computer-executable instructions stored in the memory 503, thereby implementing the methods provided in the following embodiments of the present application.
[0343] Optionally, the computer-executable instructions in the embodiments of the present application may also be referred to as application code, instructions, computer programs or other names, which are not specifically limited in the embodiments of the present application.
[0344] In a specific implementation, as an embodiment, the processor 501 may include one or more CPUs, such as CPU0 and CPU1 in FIG17 .
[0345] In a specific implementation, as an embodiment, a device may include multiple processors, such as processor 501 and processor 504 in Figure 17. Each of these processors may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. The processor herein may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0346] Optionally, the device may further include at least one communication interface 502. The communication interface 502 is used to communicate with other devices. In the embodiment of the present application, the communication interface can be a module, circuit, bus, interface, transceiver, or other device capable of implementing a communication function, used to communicate with other devices. Optionally, when the communication interface is a transceiver, the transceiver can be an independently provided transmitter that can be used to send information to other devices, or the transceiver can be an independently provided receiver that is used to receive information from other devices. The transceiver can also be a component that integrates the functions of sending and receiving information. The embodiment of the present application does not limit the specific implementation of the transceiver.
[0347] It is understood that the structure shown in FIG17 does not constitute a specific limitation on the device. In other embodiments of the present application, the device may include more or fewer components than shown, or some components may be combined or separated, or the components may be arranged differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0348] The core structure of the device can be represented as the structure shown in FIG18 , and the device includes: a processing module 2301 , a storage module 2303 and a display module 2304 .
[0349] Processing module 2301 may include at least one of a central processing unit (CPU), an application processor (AP), or a communication processor (CP). Processing module 2301 may perform operations or data processing related to the control and / or communication of at least one of the other components of the user device. Specifically, processing module 2301 may be used to control the content displayed on the main screen based on certain trigger conditions. Processing module 2301 may also be used to process input instructions or data and determine a display style based on the processed data.
[0350] Optionally, an input module 2302 may be included to receive user input commands or data and transmit the received commands or data to other modules of the device. Specifically, the input module 2302 may accept input methods such as touch, gestures, proximity to the screen, or voice input. For example, the input module may be the device's screen, receive user input operations, generate input signals based on the received input operations, and transmit the input signals to the processing module 2301.
[0351] The storage module 2303 may include a volatile memory and / or a non-volatile memory. The storage module is used to store at least one instruction or data related to other modules of the user equipment device.
[0352] The display module 2304 may include, for example, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a microelectromechanical system (MEMS) display, or an electronic paper display, and is used to display user-viewable content (e.g., text, images, videos, icons, symbols, etc.).
[0353] Optionally, a communication module 2305 is also included to support personal devices communicating with other personal devices (via a communication network). For example, the communication module can be connected to a network via wireless communication or wired communication to communicate with other personal devices or network servers. Wireless communication can use at least one of cellular communication protocols, such as Long Term Evolution (LTE), Advanced Long Term Evolution (LTE-A), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System (UMTS), Wireless Broadband (WiBro), or Global System for Mobile Communications (GSM). Wireless communication may include, for example, short-range communication. Short-range communication may include at least one of Wireless Fidelity (Wi-Fi), Bluetooth, Near Field Communication (NFC), Magnetic Stripe Transmission (MST), or GNSS.
[0354] It should be noted that each functional module of the device can execute one or more steps in the above method embodiment.
[0355] An embodiment of the present application also provides a chip system, which includes at least one processor and at least one interface circuit. The processor and the interface circuit can be interconnected via lines. For example, the interface circuit 1402 can be used to receive signals from other devices (such as the memory of the device). For another example, the interface circuit can be used to send signals to other devices (such as processors). Exemplarily, the interface circuit can read instructions stored in the memory and send the instructions to the processor. When the instruction is executed by the processor, the device can execute the various steps in the above embodiment. Of course, the chip system can also include other discrete devices, which is not specifically limited in the embodiment of the present application.
[0356] An embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed on the above-mentioned device, the device executes each function or step in the above-mentioned method embodiment.
[0357] The embodiment of the present application further provides a computer program product, which, when executed on a computer, enables the computer to execute the functions or steps executed by the mobile phone in the above method embodiment.
[0358] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0359] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0360] Units described as separate components may or may not be physically separate, and components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0361] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0362] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0363] The above content is only a specific embodiment of this application, but the scope of protection of this application is not limited to this. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A fault diagnosis method, characterized in that, Including: Obtaining relationship information among multiple node objects, where the relationship information includes the topological relationship among the multiple node objects and the reliability status of the multiple node objects; Performing fault diagnosis on the multiple node objects based on the relationship information among the multiple node objects.
2. The method according to claim 1, wherein The multiple node objects include a first node object and at least one second node object; The at least one second node object is deployed on the first node object, and the topological relationship between the first node object and the at least one second node object is a deployment relationship; Or, the first node object includes the at least one second node object, and the topological relationship between the first node object and the at least one second node object is a composition relationship; Or, the first node object includes the at least one second node object, the topological relationship among the at least one second node objects is a redundancy relationship, and the topological relationship between the first node object and the at least one second node object is an aggregation relationship; Or, there is a communication relationship between the first node object and the at least one second node object, and the topological relationship between the first node object and the at least one second node object is an upstream-downstream relationship.
3. The method according to claim 1 or 2, characterized in that, The reliability status of any node object among the multiple node objects includes normal or abnormal; the abnormal includes any one of the following states: early warning, damage, overload, fault.
4. The method according to any one of claims 1-3, characterized in that When the topological relationship between the first node object and the at least one second node object is a deployment relationship and the reliability status of the first node object is abnormal, the reliability status of the at least one second node object is abnormal; Or, when the topological relationship between the first node object and the at least one second node object is a composition relationship and there is a second node object with an abnormal reliability status among the at least one second node objects, the reliability status of the first node object is abnormal; Or, when the topological relationship between the first node object and the at least one second node object is an aggregation relationship and the reliability statuses of all the at least one second node objects are abnormal, the reliability status of the first node object is abnormal.
5. The method according to any one of claims 1-4, characterized in that, The reliability status of any node object among the multiple node objects is determined according to the metrics of the node object; the metrics include at least one of the following metrics: the traffic of the node object, the overload degree of the node object, the response time of the node object, the data error rate of the node object.
6. The method according to any one of claims 1-5, characterized in that, The multiple node objects include resource objects and / or service objects; the service object is a node object with service processing functions; the resource object is a node object that provides resources.
7. The method according to any one of claims 1-6, characterized in that When at least one of the metrics of the traffic, overload degree, and response time of the node object is abnormal and the data error rate of the node object is normal, the reliability status of the node object is an early warning; Or, when at least one of the metrics of the node object, such as traffic, overload level, and response time, is abnormal, the data error rate of the node object is abnormal, and the node object has not lost its node function, the reliability status of the node object is damaged; Or, when the overload level of the node object is abnormal and the node object has not lost its node function, the reliability status of the node object is overloaded; Or, when the traffic, overload level, response time, and data error rate of the node object are all abnormal and the node object has lost its node function, the reliability status of the node object is faulty.
8. The method according to any one of claims 1-7, characterized in that, It further includes: Sending the result of the fault diagnosis and / or the fault analysis diagram.
9. The method according to claim 8, wherein The fault analysis diagram includes the topology of the node object associated with the root cause of the fault.
10. The method according to claim 9, wherein The topology of the node object associated with the root cause of the fault is determined based on multiple paths, and the multiple paths are determined based on the relationship information between the multiple node objects; each path in the multiple paths includes at least one node object.
11. The method according to claim 10, wherein Sending the information of the multiple paths.
12. The method according to claim 5, characterized in that, It further includes: Receiving the information of the metrics of the multiple node objects and the diagnostic rules corresponding to the metrics of the multiple node objects, where the metrics of the multiple node objects and the diagnostic rules are used to determine the reliability status of the multiple node objects.
13. The method according to any one of claims 1-12, characterized in that, Obtaining the relationship information between multiple node objects includes: Receiving the relationship information between the multiple node objects.
14. The method according to any one of claims 1-13, characterized in that, It further includes: Sending the relationship information between the multiple node objects, where the relationship information is used for fault diagnosis of the multiple node objects.
15. A fault diagnosis method, characterized in that, It includes: Sending the relationship information between multiple node objects, where the relationship information includes the topological relationship between the multiple node objects and the reliability status of the multiple node objects; The relationship information is used for fault diagnosis of the multiple node objects; Receiving the result of the fault diagnosis.
16. The method according to claim 15, wherein The multiple node objects include a first node object and at least one second node object; The at least one second node object is deployed on the first node object, and the topological relationship between the first node object and the at least one second node object is a deployment relationship; Or, the first node object includes the at least one second node object, and the topological relationship between the first node object and the at least one second node object is a composition relationship; Or, the first node object includes the at least one second node object, the topological relationship between the at least one second node objects is a redundant relationship, and the topological relationship between the first node object and the at least one second node object is an aggregation relationship; Or, there is a communication relationship between the first node object and the at least one second node object, and the topological relationship between the first node object and the at least one second node object is an upstream-downstream relationship.
17. The method according to claim 15 or 16, characterized in that, The reliability status of any node object in the multiple node objects includes normal or abnormal; the abnormal includes any of the following states: warning, damage, overload, and fault.
18. The method according to any one of claims 15-17, characterized in that, The topological relationship between the first node object and the at least one second node object is a deployment relationship. When the reliability status of the first node object is abnormal, the reliability status of the at least one second node object is abnormal; Or, the topological relationship between the first node object and the at least one second node object is a composition relationship. When there is a second node object with an abnormal reliability status among the at least one second node objects, the reliability status of the first node object is abnormal; Or, the topological relationship between the first node object and the at least one second node object is an aggregation relationship. When the reliability statuses of all the at least one second node objects are abnormal, the reliability status of the first node object is abnormal.
19. The method according to any one of claims 15-18, characterized in that, The reliability status of any node object among the multiple node objects is determined according to the metrics of the node object; the metrics include at least one of the following metrics: the traffic of the node object, the overload level of the node object, the response time of the node object, and the data error rate of the node object.
20. The method according to any one of claims 15-19, characterized in that The multiple node objects include resource objects and / or service objects; the service object is a node object with service processing functions; the resource object is a node object that provides resources.
21. The method according to any one of claims 15-20, wherein When at least one of the metrics of the node object, such as traffic, overload level, and response time, is abnormal and the data error rate of the node object is normal, the reliability status of the node object is a warning; Or, when at least one of the metrics of the node object, such as traffic, overload level, and response time, is abnormal, the data error rate of the node object is abnormal, and the node object has not lost its node function, the reliability status of the node object is damaged; Or, when the overload level of the node object is abnormal and the node object has not lost its node function, the reliability status of the node object is overloaded; Or, when all of the traffic, overload level, response time, and data error rate of the node object are abnormal and the node object has lost its node function, the reliability status of the node object is a failure.
22. The method according to any one of claims 15 - 21, characterized in that It further includes: Sending a fault analysis diagram, which includes the topology of the node objects associated with the root cause of the fault.
23. The method according to claim 22, wherein The topology of the node objects associated with the root cause of the fault is determined according to multiple paths, and the multiple paths are determined according to the relationship information between the multiple node objects; each path among the multiple paths includes at least one node object.
24. The method according to claim 23, wherein Sending the information of the multiple paths.
25. The method according to any one of claims 15-24, characterized in that, It further includes: Receiving the information of the metrics of the multiple node objects and the diagnostic rules corresponding to the metrics of the multiple node objects, and the metrics of the multiple node objects and the diagnostic rules are used to determine the reliability status of the multiple node objects.
26. The method according to any one of claims 15 - 25, characterized in that, Obtaining the relationship information between multiple node objects includes: Receiving the relationship information between the multiple node objects.
27. A fault diagnosis system, characterized in that, Comprising a first diagnostic device for performing the method according to any one of claims 1-14 and a second diagnostic device for performing the method according to any one of claims 15-26.
28. A fault diagnosis method, characterized in that The method comprises: The first diagnostic device performs the method according to any one of claims 1-14; The second diagnostic device performs the method according to any one of claims 15-26.
29. A computer-readable storage medium, characterized in that, Comprising a program or instructions which, when executed, implement the method according to any one of claims 1 to 14, or implement the method according to any one of claims 15 to 26.
30. A fault diagnosis device, characterized in that, The device comprises a processor and a memory; The memory is used for storing computer execution instructions. When the device runs, the processor executes the computer execution instructions stored in the memory so that the device performs the method according to any one of claims 1-14; or performs the method according to any one of claims 15-26.
Citation Information
Patent Citations
Equipment information processing method and device, terminal equipment and storage medium
CN108737179A
Construction method of deep neural network model and fault diagnosis method and system
CN111342997A
Offshore floating type wind power reliability evaluation method and device, equipment and storage medium
CN117034744A
Fault node positioning method and device, electronic equipment and nonvolatile storage medium
CN119030860A
Topology Alarm Correlation
US20230239206A1
Cited By
Vehicle fault diagnosis method and device, electronic equipment and storage medium
CN120560231A
Intelligent operation and maintenance monitoring method and system for data center
CN120602308A
System and method for intelligent research and judgment and automatic response of weblog fused with large model
CN121239555A
Automatic fault tracing method, system and equipment based on topological coding and medium
CN121239563A
Visual operation and maintenance fault diagnosis and problem attribution analysis method for information system
CN121478537A