Fault processing method and apparatus, electronic device, and storage medium
By detecting the communication link model of edge nodes in distributed cloud scenarios and identifying faulty equipment, the problem of batch false alarms of equipment at high failure rates of distributed cloud nodes is solved, and the fault processing efficiency is improved and the operation and maintenance pressure is reduced.
Patent Information
- Application Number
- PCT/CN2024/137532
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-24
- Filing Date
- 2024-12-06
- Publication Date
- 2025-07-31
AI Technical Summary
In the prior art, the distributed cloud node fault processing process fails to effectively respond to high failure rate scenarios, resulting in batch false alarms of edge node equipment, increasing operation and maintenance pressure.
By determining the communication link model of edge nodes, detecting the operating status of equipment between the center node and edge node, identifying link failures, reducing batch alarms of equipment, and improving fault processing efficiency.
Quickly detect link failures, reduce batch alarms of equipment, reduce operation and maintenance pressure, and improve fault handling efficiency.
Smart Images

Figure CN2024137532_31072025_PF_FP_ABST
Abstract
Description
[Corrected 26.12.2024 in accordance with Rule 26] Fault handling method, device, electronic device and storage medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 24, 2024, with application number 202410101673X and application name “Fault Handling Method, Device, Electronic Device and Storage Medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of cloud computing technology, and more specifically, to a method, device, electronic device, and storage medium for fault handling in a distributed cloud scenario. Background Art
[0003] Distributed cloud extends the capabilities of central cloud, aiming to provide ubiquitous cloud computing services to meet the needs of local access for edge data processing, real-time computing, and other applications. Compared to central cloud, distributed cloud can be used anywhere, making it compact, powerful, and ultra-low-cost. To reduce costs, distributed cloud nodes can be deployed with as few as three devices. A significant amount of non-essential capabilities are remotely connected to the central cloud via the public network, significantly reducing the cooling and power requirements of Internet Data Center (IDC) computer rooms. This, however, inevitably leads to an increased failure rate for cloud nodes.
[0004] In related technologies, the fault handling process for distributed cloud nodes follows the general process of the central cloud: all faults are handled in real time. However, this general process is suitable for central cloud scenarios with large node scale and low failure rates, but is not suitable for distributed cloud scenarios with small node scale and high failure rates. Specifically, distributed cloud nodes reuse the central cloud's fault management system via public network links. Public network link failures can cause the entire edge node to lose connectivity, triggering batch alarms for all devices within the edge node. Such devices, while operating normally, may experience batch false alarms due to link interruptions, further increasing O&M pressure. Summary of the Invention
[0005] The embodiments of the present application provide a method, apparatus, device, and storage medium for fault handling, which can facilitate rapid discovery of device disconnection in edge nodes caused by link failures and improve fault handling efficiency.
[0006] In a first aspect, an embodiment of the present application provides a fault handling method, which is applied to a distributed cloud scenario, wherein the distributed cloud scenario includes a central node and at least one edge node, and the method includes:
[0007] Determining a first device in the at least one edge node that cannot operate normally;
[0008] Obtain a first communication link model of a first edge node to which the first device belongs, where the first communication link model includes a detection module in the central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device;
[0009] The operating status of each device in the first communication link model is detected to determine link fault information, where the link fault information includes information about the faulty device in the first communication link model.
[0010] In a second aspect, an embodiment of the present application provides a fault handling device, which is applied to a distributed cloud scenario, wherein the distributed cloud scenario includes a central node and at least one edge node, and the device includes:
[0011] a determining unit, configured to determine a first device in the at least one edge node that cannot operate normally;
[0012] an acquiring unit, configured to acquire a first communication link model of a first edge node to which the first device belongs, the first communication link model including a detection module in the central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device;
[0013] The detection unit is used to detect the operating status of each device in the first communication link model and determine link failure information, where the link failure information includes information about the failed device in the first communication link model.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0015] a processor adapted to implement computer instructions; and,
[0016] The memory stores computer instructions, where the computer instructions are suitable for being loaded by the processor and executing the method of the first aspect.
[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are read and executed by a processor of a computer device, the computer device executes the method of the first aspect above.
[0018] In a fifth aspect, embodiments of the present application provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method of the first aspect described above.
[0019] Through the above technical solution, when the first device in the edge node cannot operate normally, the embodiment of the present application obtains the first communication link model corresponding to the first edge node to which the edge node belongs, and the first communication link model includes a detection module in the central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device. Then, the operating status of each device in the first communication link model is detected to identify the faulty device in the first communication link model. Therefore, when an edge node failure occurs, the embodiment of the present application can not only detect the device in the edge node, but also detect each device on the communication link between the detection module in the central node and the device of the edge node, thereby facilitating the rapid discovery of device failures in the edge node caused by link failures and improving fault handling efficiency. Furthermore, by quickly discovering link failures, it can be helpful to reduce or avoid batch alarms of devices in the edge node and reduce operation and maintenance pressure. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] FIG1 is an optional schematic diagram of a network architecture for a distributed cloud scenario;
[0021] FIG2 is a schematic flow chart of a fault handling method provided in an embodiment of the present application;
[0022] FIG3 is a schematic diagram of a network architecture of a distributed cloud scenario provided by an embodiment of the present application;
[0023] FIG4 is a schematic flow chart of another fault handling method provided in an embodiment of the present application;
[0024] FIG5 is a schematic diagram of a fault detection process provided in an embodiment of the present application;
[0025] FIG6 is a schematic diagram of a fault handling process provided in an embodiment of the present application;
[0026] FIG7 is a schematic block diagram of a fault handling apparatus provided in an embodiment of the present application;
[0027] FIG8 is a schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0028] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.
[0029] It should be understood that in the embodiments of the present application, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B based solely on A; B can also be determined based on A and / or other information.
[0030] In the description of this application, unless otherwise specified, "at least one" means one or more, and "plurality" means two or more than two. In addition, "and / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0031] It should also be understood that the first, second, etc. descriptions appearing in the embodiments of the present application are only for illustration and distinction of the description objects, and there is no order. They do not represent any special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation on the embodiments of the present application.
[0032] It should also be understood that the specific features, structures, or characteristics associated with the embodiments in the specification are included in at least one embodiment of the present application. In addition, these specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0033] In addition, the terms "include" and "have" and any variations thereof are intended to cover a non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or elements is not necessarily limited to those steps or elements expressly listed, but may include other steps or elements not expressly listed or inherent to such process, method, product, or device.
[0034] FIG1 is an optional schematic diagram of a network architecture for a distributed cloud scenario.
[0035] As shown in Figure 1, a distributed cloud scenario includes central nodes and edge nodes. The central node, also known as the central cloud, includes at least one device, such as device 1 and device 2. The edge node, also known as the edge cloud, includes at least one device, such as device 1' and device 2'. There can be multiple edge nodes 120.
[0036] Distributed cloud can provide various cloud computing services for terminals. Terminals can also be called user terminals, including but not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc.
[0037] In the fault handling process in a distributed cloud scenario, the central node is primarily responsible for the entire fault management lifecycle; the edge node is responsible for responding to the central node's detection commands and reporting monitoring data. As shown in Figure 1, a complete fault management chain includes a detection module 110, a response module 120, and a fault handling module 130.
[0038] Detection module 110 is responsible for initiating device status checks. Detection module 110 is deployed at the central node. Optionally, during device status checks, detection module 110 can use services as units, with each service managing its own devices independently. For example, service 1 manages corresponding device 1, and service 2 manages corresponding device 2. Exemplarily, detection module 110 can have a built-in timer. When the system is operating, detection module 110 periodically initiates fault detection heartbeats from within the service to check the survivability of the target device. Detection module 110 can also scan monitoring data reported by devices to check whether the devices are functioning properly.
[0039] In a distributed scenario, to reduce O&M costs, edge node devices reuse the fault detection capabilities of the detection module 110 in the central node. For example, as shown in Figure 1, service 1 can also manage its corresponding device 1′, and service 2 can also manage its corresponding device 2′. For example, this can be achieved by periodically initiating fault detection heartbeats within the service to check the survivability of corresponding devices in the edge node and scanning monitoring data reported by devices in the edge node. If a device experiences a heartbeat failure or abnormal monitoring data, it can be considered a fault.
[0040] The response module 120 is responsible for responding to the detection heartbeat initiated by the service and reporting monitoring data to the corresponding service. The response module 120 is deployed on the device. For example, the response module 120 is deployed on device 1 and device 2 in the central node, and the response module 120 is deployed on device 1' and device 2' in the edge node. During the equipment fault monitoring process, the response module 120 actively reports the real-time status monitoring data of the equipment and passively responds to the detection heartbeat. Among them, the internal equipment of the central node and the detection module 110 can directly transmit data. Unlike the equipment in the central node, when the detection module 110 in the central node issues a detection command, the data needs to go through a complex communication link to reach the device in the edge node. As shown in Figure 1, the communication link can, for example, include a routing module, a sending module, a public network link in the central node, a receiving module in the edge node, a routing module, etc.
[0041] Fault handling module 130 is responsible for resolving faults and is deployed in the central node. As shown in Figure 1, the fault handling module can include an alarm module and a processing module. When a device fault is detected, the corresponding service module in detection module 110 immediately triggers the alarm module to issue an alert, and the relevant processing module then performs fault resolution.
[0042] In the related art, in the edge node fault handling process, the detection heartbeat and detection data between the detection module 110 and the response module 120 need to be transmitted through a complex communication link. This communication link is a public network control link, with a long transmission link and many forwarding devices, which adds uncertainty to the data transmission. If the link fails, the entire edge node will lose connection, unable to respond to the detection heartbeat, and unable to report status monitoring data. Furthermore, the failure of the link will also trigger batch alarms for devices in the edge node. In daily fault handling, it is necessary to first filter out the real alarms from a large number of invalid alarms. This problem of a single device failure causing batch alarms for all devices in the edge node will seriously reduce the efficiency of fault handling and interfere with normal node operation and maintenance.
[0043] In view of this, the embodiments of the present application provide a fault handling method, device, electronic device and storage medium for application in a distributed cloud scenario, which can help quickly discover the loss of connection of devices in edge nodes caused by link failures and improve fault handling efficiency.
[0044] Specifically, the distributed cloud scenario includes a central node and at least one edge node. First, the first device that cannot operate normally in the at least one edge node is determined; then the first communication link model of the first edge node to which the first device belongs is obtained, and the first communication link model includes a detection module in the central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device; then, the operating status of each device in the first communication link model is detected to determine the link fault information, which includes information about the faulty device in the first communication link model.
[0045] Therefore, when the first device in the edge node cannot operate normally, the embodiment of the present application obtains the first communication link model corresponding to the first edge node to which the edge node belongs, and the first communication link model includes a detection module in the central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device, and then detects the operating status of each device in the first communication link model to identify the faulty device in the first communication link model. Therefore, when an edge node failure occurs, the embodiment of the present application can not only detect the device in the edge node, but also detect each device on the communication link between the detection module in the central node and the device of the edge node, thereby facilitating the rapid discovery of device failures in the edge node caused by link failures and improving fault handling efficiency. Furthermore, by quickly discovering link failures, it can be helpful to reduce or avoid batch alarms of devices in the edge node and reduce operation and maintenance pressure.
[0046] The following describes the technical solutions of the embodiments of the present application in detail through some embodiments. The following embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0047] Figure 2 is a schematic flow chart of a fault handling method 200 for a distributed cloud scenario provided in an embodiment of the present application. Method 200 can be executed by any electronic device with data processing capabilities, and this application does not limit this.
[0048] Exemplarily, method 200 can be applied to the network architecture in a distributed cloud scenario as shown in FIG3 . As shown in FIG3 , the network architecture includes a central node and at least one edge node, as well as a detection module 110 , a response module 120 , a fault handling module 130 , and a link diagnostic module 310 . Detection module 110 , response module 120 , and fault handling module 140 are similar to the corresponding modules in FIG1 . For details, please refer to FIG1 , and will not be repeated here. Link diagnostic module 310 is deployed in the central node.
[0049] Optionally, the network architecture may further include at least one of a fault list 320, a link model 330, a link entry module 340, and a fault orchestration module 350. The fault list 320, the link model 330, the link entry module 340, and the fault orchestration module 350 may be deployed in a central node.
[0050] Hereinafter, the method 200 will be described in conjunction with the network architecture in FIG3 . As shown in FIG2 , the method 200 may include steps 210 to 230 .
[0051] 210. Determine a first device in at least one edge node that cannot operate normally.
[0052] For example, referring to FIG3 , each service (such as service 1 and service 2) in the detection module 110 can initiate fault detection on the response module 120 of the device (such as device 1′ and device 2′) in the edge node, detect the survivability of the corresponding device in the edge node, and scan the monitoring data reported by the device in the edge node. When a device in the edge node fails to detect the heartbeat or the monitoring data is abnormal, the detection module 110 determines that the device cannot operate normally or that the device has an abnormality. At this time, the first device may have a fault, or the device (such as a forwarding device) on the communication link between the response device of the first device and the detection module 110 may have a fault.
[0053] Here, a heartbeat detection failure or abnormal monitoring data may include not receiving the device's heartbeat or monitoring data, or receiving the heartbeat and monitoring data but the monitoring data itself is abnormal. It is understood that the first device determined to be malfunctioning in step 210 may or may not be faulty. For example, when a device on the communication link between the first device's answering device and the detection model 110 is faulty, the first device may not be faulty.
[0054] Optionally, the link diagnosis module 310 may obtain information about a first device that cannot operate normally in at least one edge node of the detection module 110 .
[0055] In a possible implementation, the detection module 110 may send information about the first device in the edge node that is not operating normally to the link diagnosis module in real time.
[0056] In another possible implementation, the detection module 110 may send information about the first device that is not operating normally in the edge node detected in real time to the fault list 320 for caching. The fault list 320 includes information about at least one device that is not operating normally.
[0057] In some embodiments, when the communication link between the response module 120 in the edge node and the detection module 110 in the central node fails, multiple devices in the edge node may be detected as unable to operate normally at the same time. At this time, the multiple devices will all be added to the fault list 320.
[0058] Accordingly, the link diagnostic module 310 can obtain the information of the first device from the fault list 320 in a preset manner, thereby determining the first device that is not operating normally in the edge node. For example, the link diagnostic module 310 can read the information of the first device that is not operating normally in the fault list 320 one by one in the preset manner. As an example, the preset manner can include timed acquisition or periodic acquisition, etc., which is not limited in this embodiment of the present application.
[0059] Specifically, for scenarios where edge nodes are small in scale and have a high failure rate, the impact of a failure is small, and most of the time it only affects the inside of the node. Based on this, when the detection module 110 detects a failure in an edge node, it adds information about the device that cannot operate normally to the fault list instead of directly triggering a fault alarm, thereby meeting the requirement that edge nodes with lower availability requirements do not need to be processed in real time.
[0060] 220. Obtain a first communication link model of a first edge node to which the first device belongs. The first communication link model includes a detection module in the central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device.
[0061] Exemplarily, referring to FIG3 , after the link diagnosis module 310 determines that the first device cannot operate normally, it may determine the first edge node to which the first device belongs, and further obtain a first communication link model of the first edge node.
[0062] Specifically, the first communication link model includes a communication network between a central node (such as a detection module) and an edge node (such as a response module). The first communication link model can be determined based on the basic information of each device in the communication link between the central node (such as a detection module) and the edge node (such as a response module), as well as the connection relationship between each device.
[0063] In some embodiments, the first communication link model includes a tree structure with the detection module as the root node, the second device in the first edge node as the leaf node, and the third device used for data transmission between the detection module and the second device as the intermediate node. In the tree structure, the root node and the intermediate node record the devices corresponding to their child nodes. The devices corresponding to the child nodes of a node are the subsequent devices of the node, and the leaf nodes have no subsequent devices. Optionally, the device corresponding to the root node can be called a root device, and the devices corresponding to the leaf nodes can be called leaf devices.
[0064] Specifically, in the communication link between the central node's detection module and the edge node's device, the next-hop downstream device of the root node or intermediate node is its subsequent device. By recording the devices corresponding to their child nodes, i.e., the subsequent devices, the communication link model can accurately and effectively represent the communication network between the central node and the edge node.
[0065] Optionally, the second device may include the above-mentioned first device.
[0066] Specifically, the second device is a device in the first edge node, and the first device belongs to the first edge node, that is, the first device is a device in the first edge node that is not functioning properly. When all devices in the first edge node are not functioning properly, the first device and the second device are the same device. When some devices in the first edge node are not functioning properly, the first device is a subset of the second device.
[0067] In some embodiments, when multiple devices in an edge node are simultaneously detected to be malfunctioning, the multiple devices are all added to the fault list 320. Furthermore, since the multiple devices belong to the same edge node, only the communication link network of the edge node needs to be acquired.
[0068] In some embodiments, a communication link model of at least one edge node may be obtained, and the communication link model of at least one edge node may be saved in a link model pool. The link model pool may include the communication link model of at least one edge node. Alternatively, the link model pool may include the communication link models of all edge nodes. In this case, based on the information of the first device, the first communication link model of the first edge node to which the first device belongs may be obtained from the link model pool.
[0069] 3 , the link entry module 340 is responsible for completing the construction of the communication link model of the edge node, such as obtaining the communication link model of at least one edge node in advance and saving the communication link model of at least one edge node to the link model pool 330. For example, the communication link model can be saved using the edge node identifier as a key.
[0070] In some embodiments, the link entry module 340 may also record the service level provided by the edge node. Different response levels of fault handling strategies may be provided for different service levels. Optionally, the service level may be stored in a link model pool.
[0071] It should be understood that in method 200, the description is based on the example of the failure of the first device in the first edge node to operate normally. The fault handling process for devices in other edge nodes in the distributed cloud scenario that fail to operate normally is similar to the processing process corresponding to the first device in the first edge node. For details, please refer to the relevant description, and this embodiment of the application will not be repeated here.
[0072] 230 , detecting the operating status of each device in the first communication link model, and determining link fault information, where the link fault information includes information of the faulty device in the first communication link model.
[0073] For example, referring to FIG3 , the link diagnostic module 310 can sequentially detect the operating status of each device in the communication link model of the edge node to which the device that is not functioning properly belongs, based on the operating status of each device in the edge node. Specifically, the link diagnostic module 310 sequentially detects the operating status of each device in the communication link model. Specifically, the link diagnostic module 310 sequentially detects the detection module in the central node, the devices at the edge node, and each device (e.g., a forwarding device) on the communication link between the detection module and the devices at the edge node until a faulty device is detected. Subsequently, the faulty link information can be determined based on the information about the faulty device.
[0074] In some embodiments, for multiple devices that cannot operate normally in the same edge node in the fault list 320, by detecting the operating status of each device in the communication link network of the edge node, batch alarms caused by the same device failure can be merged and processed to avoid alarm abuse.
[0075] Therefore, when a fault occurs, the embodiments of the present application can detect not only the devices in the edge node but also each device on the communication link between the detection module in the central node and the devices in the edge node. This facilitates the rapid discovery of device faults in the edge node caused by link failures, thereby improving fault handling efficiency. Furthermore, by quickly discovering link failures, it can help reduce or avoid batch alarms from devices in the edge node, reducing operation and maintenance pressure.
[0076] Furthermore, by adding the information of the detected device that is not operating normally to the fault list, and then reading the information of the device that is not operating normally from the fault list, fault detection is performed on each device in the communication link model of the edge node, instead of directly triggering a fault alarm, so that the fault processing entrance can be converged, and batch alarms caused by the same device failure can be merged and processed to avoid alarm abuse.
[0077] In some embodiments, when the first communication link model includes a tree structure with a detection module as a root node, a second device in the first edge node as a leaf node, and a third device for data transmission between the detection module and the second device as an intermediate node, the operating status of each node in the tree structure can be traversed sequentially starting from the detection module corresponding to the root node (i.e., the root device), thereby determining the faulty device in the first communication link model. Then, the faulty link information can be determined based on the information about the faulty device.
[0078] For example, the node currently being detected in the first communication link model may be referred to as the current node. The current node may be a root node, an intermediate node, or a leaf node. For the current node, it may be detected whether the device corresponding to the current node is operating normally. If the current node is not operating normally, the device corresponding to the current node is determined to be a faulty device.
[0079] Optionally, if the current node is running normally and has child nodes, the device corresponding to the child node of the current node (ie, the subsequent device) is detected to be normal. At this time, the next current node can be updated to the device corresponding to the child node.
[0080] Optionally, if the current node is operating normally and has no child nodes, the sibling node of the current node may be checked, that is, whether the device corresponding to the sibling node is operating normally. In this case, the next current node may be updated to the sibling device.
[0081] Optionally, if the current node is not functioning properly, the current node's sibling nodes may be checked to complete the traversal of the tree structure corresponding to the entire first communication link model, thereby avoiding missing any faulty devices that may exist in the sibling nodes. If no sibling devices exist for the current node, the traversal of the tree structure corresponding to the entire first communication link model is complete.
[0082] Therefore, the embodiment of the present application sequentially traverses each device in the first communication link model. When a device responds normally and the detection data is normal, the device is considered a normal device, and the subsequent devices of the device can be detected along the first communication link model. When a device does not respond or the detection data is abnormal, the device is marked as a faulty device, and the subsequent devices of the device can be ignored, and the sibling devices of the device can be traversed instead, thereby completing the traversal check of the first communication link model, which can facilitate the rapid discovery of link faults.
[0083] Furthermore, by detecting the survival status or detection data of each device in the communication link model, the embodiments of the present application can collect fault metadata, thereby obtaining information about the faulty device. Fault metadata can be used to identify the fault type, thereby improving the accuracy of fault identification.
[0084] Optionally, the information about the faulty device may further include time information of the fault, wherein the time information of the fault may be used to determine the duration of the fault.
[0085] In some embodiments, referring to FIG. 4 , the method 200 may further include the following steps 240 and 250 .
[0086] 240. Determine the number of second devices that cannot operate normally in the first communication link.
[0087] For example, after traversing and checking each device in the first communication link model, the number of unreachable leaf devices in the first communication link model can be counted. The unreachable leaf devices are devices in the first edge node that cannot operate normally. The second devices that cannot operate normally are a subset of the second devices in the above step 220. The non-operational operation can refer to the relevant description above.
[0088] 250. Determine a fault state model of the first edge node according to the faulty device, the number of the second devices that cannot operate normally, and the service level information of the first edge node.
[0089] Specifically, the fault status model may be determined based on information about the faulty device in the first communication link model, the number of unreachable leaf devices in the first communication link model, and service level information of the first edge node.
[0090] Optionally, the fault state model can be saved to a link state pool. For example, the fault state model can be deduplicated using the fault device identifier as a key and saved to the link state pool. For example, if the key in the link state pool already includes the identifier of a fault device, the fault state model corresponding to the fault device in the link state pool can be updated.
[0091] The embodiment of the present application saves the fault status model to the link status pool instead of directly triggering a fault alarm or performing fault modification. This can meet the requirement that faults of edge nodes with low availability requirements do not need to be processed in real time, which is beneficial to reducing operation and maintenance costs.
[0092] If all devices are found to be functioning properly after traversing the first communication link model, the system is functioning properly. If a faulty device is found after traversing the communication link model, the system is faulty and requires further processing. For example, the fault processing module repairs the fault based on the link status.
[0093] In related technologies, the standard fault handling process involves manually confirming the fault twice before selecting a fault handling module for troubleshooting. This approach requires 24 / 7 manual service and is suitable for central clouds with stable operating environments and low failure rates. However, it is not suitable for distributed cloud scenarios with numerous uncontrollable factors, high failure rates, and limited impact. Furthermore, as the amount of data in distributed cloud nodes increases, the cost of manual processing increases rapidly.
[0094] In view of this, the embodiment of the present application further determines the fault level of the first edge node based on the detection results of the first communication link model, and then handles the faulty device according to the fault level. By hierarchically handling faults, it is possible to immediately handle high-priority faults and downgrade low-priority faults, thereby effectively reducing the number of times manual real-time fault handling is required, improving overall fault handling efficiency and reducing maintenance costs.
[0095] In some embodiments, referring again to FIG. 4 , method 200 may further include steps 260 and 270 .
[0096] 260. Determine a fault level of the first edge node according to at least one of a weight of the faulty device, service level information of the first edge node, and a fault duration.
[0097] Exemplarily, referring to FIG3 , the fault orchestration module 350 may determine the fault level of the first edge node according to at least one of the weight of the faulty device, the service level information of the first edge node, and the fault duration.
[0098] Exemplarily, the fault orchestration module 350 can obtain link fault information from the link diagnosis module 310 and, based on the link fault information, determine at least one of the weight of the faulty device, the service level information of the first edge node, and the fault duration. For example, the fault orchestration module 350 can obtain a fault status model from the link status pool in a preset manner. Exemplarily, the preset manner can include scheduled or periodic acquisition, etc., which is not limited in this embodiment of the present application. Then, based on the fault status model, at least one of the weight of the faulty device, the service level information of the first edge node, and the fault duration can be obtained.
[0099] In some embodiments, a grading score can be determined based on the weight of the faulty device, the service level information of the first edge node, and the fault duration, and then the fault level can be determined based on the grading score. For example, the grading score S can be expressed as the following formula (1):
[0100] Where error_count represents the weight of the faulty device. The weight of each device can be set based on its importance in the communication link model. ∑error_count is the sum of the weights of all subsequent devices affected by the faulty device. level_vip represents the service level information, which is the pre-entered availability level of the edge node. The higher the node availability requirement, the higher the service level. time represents the fault duration coefficient, which is positively correlated with the fault duration. max_score is the maximum score. The larger the calculated graded score S, the higher the processing priority.
[0101] In formula (1), when the edge node's service level is greater than or equal to 3, the grading score is the highest score, max_score. When the edge node's service level is less than 3, the grading score is calculated based on the number and weight of faulty devices, the edge node's service level information, and the fault duration. In this case, the more devices affected by the fault, the higher the edge node's service level, and the longer the fault duration, the higher the grading score.
[0102] For example, the value of max_score may be 30, which is not limited in this application.
[0103] 270 , process the faulty device according to the fault level of the first edge node.
[0104] For example, referring to Figure 3 , the fault orchestration module 350 can determine a strategy for handling the faulty device based on the fault level of the first edge node. Specifically, different response processes can be executed for different fault levels. For example, high-priority faults can be handled immediately, while low-priority faults can be handled in a downgraded manner, such as by waiting for idle time.
[0105] For example, for the grading score obtained according to the above formula (1), when the score reaches max_score, the fault needs to be handled immediately; when the service level is greater than or equal to 3, all faults need to provide real-time response; when the service level is less than 3, it is handled according to the calculated grading score. For example, the more devices affected by the fault, the higher the service level of the edge node, or the longer the fault duration of the device, the higher the grading score and the higher the corresponding handling priority.
[0106] For example, for high-priority faults, a 7×24-hour response can be provided, and for faults with a smaller impact range (such as a small number of faulty devices), it is possible to wait for idle time to handle them. In some instances, for faults whose fault types can be accurately identified and for which automatic processing methods exist, automatic recovery processing can be prioritized to reduce manual participation. For fault types that cannot be automatically recovered, the standard processing flow can be entered and manual processing can be performed after manual secondary confirmation. Therefore, the embodiments of the present application can effectively reduce the number of times that manual real-time fault processing is required, improve overall fault handling efficiency, and reduce maintenance costs.
[0107] Figure 5 is a schematic diagram of a fault detection process provided by an embodiment of the present application. It should be understood that Figure 5 illustrates steps or operations related to the fault detection process, but these steps or operations are merely examples. The embodiments of the present application may also perform other operations or variations of the operations in Figure 5. In addition, the steps in Figure 5 may be performed in a different order than that presented in Figure 5, and it is possible that not all operations in Figure 5 need to be performed. As shown in Figure 5, the fault detection process includes steps 401 to 408.
[0108] 401, the device in the edge node cannot operate normally.
[0109] Specifically, a fault occurs in an edge node, causing a device to malfunction. Information about the malfunctioning device in the edge node can be obtained. The specific process for obtaining information about the malfunctioning device can be found in the description of step 210 in FIG. 2 , and will not be repeated here.
[0110] 402, load the communication link model. M = root device.
[0111] As shown in Figure 5, the communication link model of the edge node to which the device that cannot operate normally belongs can be obtained from the link model pool. Exemplarily, the corresponding communication link model can be loaded according to the ID of the edge node to which the device that cannot operate normally belongs.
[0112] In one possible implementation, when a malfunctioning device is detected on an edge node, a module can be automatically activated to load the communication link model for the edge node to which the malfunctioning device belongs. For example, the communication link model can be a tree structure with the detection module in the central node as the root node, with each device recording its own successor devices. If a device has no successor devices, it is considered a leaf device. Devices in edge nodes are leaf devices.
[0113] Specifically, the process of loading the communication link model may refer to the relevant description in step 220 in FIG. 2 , which will not be repeated here.
[0114] After loading the communication link model, you can configure the device corresponding to the root node in the communication link model (i.e., the root device) as the current detection device, i.e., the M device. Starting from the root device, the survival status and test data of each device along the communication link model are detected in sequence until the faulty device is found.
[0115] Specifically, the detection process may include steps 403 to 407 .
[0116] 403, detect M device.
[0117] Specifically, the survival status and monitoring data of the M device may be detected to determine whether the M device is a normal device.
[0118] 404 , determine whether the M device is normal and a subsequent device exists.
[0119] For example, when the device responds normally and the monitoring data is normal, the device is considered normal. At this point, it can be determined whether there is a subsequent device to the M device based on the loaded communication link model. If there is a subsequent device to the M device, step 406 can be executed.
[0120] When both the conditions of the M device being normal and the subsequent device existing are not met at the same time, for example, when the M device cannot operate normally, or when the M device has no subsequent device, step 405 may be executed.
[0121] 405. Determine whether a brother device exists.
[0122] If a brother device exists, the next step is to execute step 407. If a brother device does not exist, the traversal check of the communication link model is completed.
[0123] When device M fails to operate normally, subsequent devices cannot receive or forward data from device M, and therefore all subsequent devices cannot operate normally. In this case, there is no need to detect subsequent devices of device M. Instead, it is possible to determine whether device M has sibling devices, thereby traversing each device in the communication link model in sequence.
[0124] When device M is operating normally but has no successor devices, it is already a leaf node in the current communication network branch, and the communication network branch is operating normally. At this point, it can be determined whether device M has sibling devices to traverse each device in the communication link model in sequence.
[0125] 406, M = subsequent device.
[0126] Specifically, the subsequent device of the M device in step 403 may be configured as the current M device.
[0127] 407, M = brother device.
[0128] Specifically, the sibling device of the M device in step 403 may be configured as the current M device.
[0129] After executing step 406 or 407, step 403 may be executed to detect the M device.
[0130] In summary, by looping through steps 403 to 407, a depth-first approach can be implemented, starting from the root device and sequentially traversing each device in the tree structure to check whether each device is functioning properly. If a device responds normally and the test data is normal, it is considered a normal device, and the communication link model can be used to continue probing its subsequent devices. If a device does not respond or the test data is abnormal, it is marked as a faulty device, and its subsequent devices can be ignored, with the traversal of its sibling devices completing the communication link model.
[0131] When all devices are found to be normal after the communication link model traversal is completed, it indicates that the system is working normally. When a faulty device is found after the communication link model traversal is completed, it indicates that there is a fault in the system and further processing is required for the faulty device, such as executing the following step 408.
[0132] 408: Save the link status.
[0133] Exemplarily, a faulty device detected in a communication link model may be saved as a link state in a link state pool.
[0134] Optionally, after traversing the communication link model, the number of unreachable leaf devices in the communication link model may be counted. Optionally, the link status may also include the number of unreachable leaf devices in the communication link model. Optionally, the link status may also include service level information of the edge node.
[0135] As an implementation approach, information such as the faulty device, the number of unreachable leaf devices, and the service level of edge nodes in the communication link model can be encapsulated into a link state model, keyed by the faulty device's identifier, and stored in a link state pool. Accordingly, the fault handling module can retrieve each link state model from the link state pool and handle the fault based on the link state model.
[0136] Therefore, when a malfunctioning device exists in an edge node, the embodiments of the present application can converge fault handling entry points by testing each device in the edge node's communication link model, merging batches of alarms caused by the same device failure in the communication link model, thereby helping to avoid alarm abuse. Statistical results show that in the fault handling process provided by the embodiments of the present application, the average number of monthly alarms for each distributed edge cloud node has been reduced from 18 to 1.3, reducing invalid alarms by 93%.
[0137] Figure 6 is a schematic diagram of a fault handling process provided by an embodiment of the present application. It should be understood that Figure 6 illustrates the steps or operations related to the fault handling process, but these steps or operations are merely examples. The embodiments of the present application may also perform other operations or variations of the operations in Figure 6. In addition, the steps in Figure 6 may be performed in a different order than that presented in Figure 6, and it is possible that not all operations in Figure 6 need to be performed. As shown in Figure 6, the fault detection process includes steps 501 to 509.
[0138] 501, load the fault state model.
[0139] For example, the fault state model can be loaded from the link state pool periodically. As an implementation method, the module can have a built-in timer to trigger a full scheduling at a fixed time. Specifically, when the scheduling task starts, all unprocessed fault state models can be loaded from the "link state pool".
[0140] Optionally, the processing of each fault state model may occupy one execution thread, so that the scheduling of state fault models may be initiated in batches through a multi-threaded manner.
[0141] 502, Is the detection device restored?
[0142] To ensure the real-time nature of link diagnostic data and avoid fault orchestration anomalies caused by outdated data, link diagnostics can be re-initiated before fault orchestration to obtain the latest link status. If the fault is detected to have been resolved, step 503 is executed. For faults that have not yet been processed, step 504 is executed.
[0143] 503, fault eliminated.
[0144] Specifically, when it is detected that the fault has been recovered, the corresponding fault state model can be removed from the link state pool and the current scheduling is terminated.
[0145] 504, determine whether the fault is being processed?
[0146] For a fault that has not been processed, it can be further determined whether the fault is a fault that is being processed. If the fault is a fault that is being processed, step 508 can be executed. If the fault is not a fault that is being processed, step 505 can be executed.
[0147] 505, fault classification.
[0148] Specifically, because the failure rate in a distributed cloud scenario is much higher than that in a central cloud scenario, in order to improve the efficiency of fault handling and ensure high-quality handling of major faults, faults need to be handled in a hierarchical manner. For example, the fault level of an edge node can be determined based on at least one of the weight of the faulty device, the service level information of the edge node, and the fault duration. For example, a hierarchical score can be calculated according to the above formula (1), and the fault level can be further determined based on the hierarchical score.
[0149] 506, does it need immediate action?
[0150] When the fault needs to be handled immediately, step 507 may be executed. When the fault does not need to be handled immediately, step 509 may be executed.
[0151] For example, if the service level is greater than or equal to 3, a real-time response can be provided, that is, the fault needs to be handled immediately; if the service level is less than 3, corresponding processing can be performed based on the calculated classification score. For example, faults with a classification score less than a certain value can be handled later, such as waiting for idle time.
[0152] 507, fault response.
[0153] Specifically, for faults that require immediate attention, different responses can be implemented based on the different fault levels. For example, for high-priority faults, a 24 / 7 response can be provided; for faults with a small impact (such as a small number of faulty devices), downgraded handling can be performed. In the fault handling process, for faults whose fault types can be accurately identified and for which automatic handling methods exist, automatic recovery can be prioritized to reduce manual intervention. For faults that cannot be automatically recovered, the standard handling process can be entered, and the fault can be manually handled after a second manual confirmation.
[0154] 508, refresh processing progress.
[0155] Specifically, for a fault that is being processed, the processing progress can be refreshed.
[0156] 509, status updated.
[0157] Specifically, after fault handling is complete, or for faults that do not require immediate handling, a fault status update can be performed. For example, a fault that has been handled can have its corresponding fault state model removed from the link state pool; faults that do not require immediate handling can wait for the next scheduling.
[0158] Therefore, by performing on-demand fault downgrade, upgrade, and automatic recovery based on fault severity, the present embodiment ensures high-quality handling of major faults, improves overall processing efficiency, and effectively reduces the number of times manual real-time processing is required. Statistics show that the average number of manual real-time fault handling required per distributed edge node per month has decreased from 0.8 to 0.3, bringing manual operation and maintenance costs close to those of the central cloud.
[0159] The specific embodiments of the present application are described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the specific details in the above embodiments. Within the technical concept of the present application, a variety of simple modifications can be made to the technical solution of the present application, and these simple modifications all fall within the scope of protection of the present application. For example, the various specific technical features described in the above specific embodiments can be combined in any suitable manner unless there is any contradiction. In order to avoid unnecessary repetition, the present application will not further explain various possible combinations. For another example, the various different embodiments of the present application can also be arbitrarily combined, and as long as they do not violate the ideas of the present application, they should also be regarded as the contents disclosed in the present application.
[0160] It should also be understood that in the various method embodiments of the present application, the order of the sequence numbers of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. It should be understood that these sequence numbers can be interchanged where appropriate, so that the embodiments of the present application described can be implemented in an order other than those shown or described.
[0161] The method embodiment of the present application is described in detail above in conjunction with Figures 1 to 6 , and the device embodiment of the present application is described in detail below in conjunction with Figures 7 to 8 .
[0162] Figure 7 is a schematic block diagram of a fault handling apparatus 10 according to an embodiment of the present application. Apparatus 10 can be used in a distributed cloud scenario, which includes a central node and at least one edge node. As shown in Figure 7 , apparatus 10 may include a determination unit 11, an acquisition unit 12, and a detection unit 13.
[0163] A determining unit 11 is configured to determine a first device in the at least one edge node that cannot operate normally;
[0164] an acquiring unit 12, configured to acquire a first communication link model of a first edge node to which the first device belongs, the first communication link model including a detection module in the central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device;
[0165] The detection unit 13 is configured to detect the operating status of each device in the first communication link model and determine link fault information, where the link fault information includes information about the faulty device in the first communication link model.
[0166] In some embodiments, the first communication link model includes a tree structure with the detection module as the root node, the second device as the leaf node, and the third device as the intermediate node; wherein the root node and the intermediate node in the tree structure record the devices corresponding to their child nodes.
[0167] In some embodiments, the detection unit 13 is specifically configured to:
[0168] Starting from the detection module corresponding to the root node, traversing the operating status of each node in the tree structure in sequence, and determining the faulty device in the first communication link model;
[0169] The link failure information is determined according to the information of the faulty device.
[0170] In some embodiments, the detection unit 13 is specifically configured to:
[0171] Detecting whether a device corresponding to a current node in the first communication link is operating normally; wherein the current node includes the root node, the intermediate node, or the leaf node;
[0172] If the current node cannot operate normally, the device corresponding to the current node is determined as the faulty device.
[0173] In some embodiments, the detection unit 13 is further configured to:
[0174] If the current node has a sibling node, then it is detected whether the device corresponding to the sibling node is normal.
[0175] In some embodiments, the detection unit 13 is further configured to:
[0176] If the current node operates normally and the current node has child nodes, whether devices corresponding to the child nodes of the current node are normal is detected.
[0177] In some embodiments, the determining unit 11 is specifically configured to:
[0178] The information of the first device that cannot operate normally in the at least one edge node is obtained from the detection module; wherein the detection module is used to obtain the operating status of each device in the central node and the at least one edge node.
[0179] In some embodiments, the apparatus 10 further includes a cache unit configured to:
[0180] The information of the first device that cannot operate normally in the at least one edge node is obtained from a fault list, wherein the fault list is used to cache the information of the at least one device that cannot operate normally from the detection module.
[0181] In some embodiments, the acquiring unit 12 is specifically configured to:
[0182] According to the information of the first device, the first communication link model of the first edge node to which the first device belongs is obtained from a link model pool; wherein the link model pool includes the communication link model of the at least one edge node.
[0183] In some embodiments, the apparatus 10 further includes a processing unit configured to:
[0184] determining a fault level of the first edge node according to at least one of a weight of the faulty device, service level information of the first edge node, and a fault duration;
[0185] Processing the faulty device according to the fault level of the first edge node.
[0186] In some embodiments, the determining unit 11 is further configured to:
[0187] determining a number of the second devices that are not operating normally in the first communication link;
[0188] determining a fault state model of the first edge node according to the link fault information, the number of the second devices that cannot operate normally, and the service level information of the first edge node;
[0189] At least one of the weight of the faulty device, the service level information, and the fault duration is determined according to the fault status model.
[0190] It should be understood that the device embodiment and the method embodiment can correspond to each other, and similar descriptions can refer to the method embodiment. To avoid repetition, no further details are given here. Specifically, when the fault handling device 10 in this embodiment can correspond to the method 200 for performing fault handling in an embodiment of the present application, the aforementioned and other operations and / or functions of each module in the device 10 are respectively for implementing the corresponding processes in each method in Figure 2. For the sake of brevity, no further details are given here.
[0191] The above describes the apparatus and system of the embodiment of the present application from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in hardware form, can be implemented by instructions in software form, and can also be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software form instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor for execution, or can be completed by a combination of hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps in the above method embodiment in conjunction with its hardware.
[0192] FIG8 is a schematic block diagram of an electronic device 30 provided in an embodiment of the present application.
[0193] As shown in FIG8 , the electronic device 30 may include:
[0194] The memory 31 and the processor 32 are configured to store computer programs and transmit the program code to the processor 32. In other words, the processor 32 can call and run the computer program from the memory 31 to implement the method in the embodiment of the present application.
[0195] For example, the processor 32 may be configured to execute the steps of each execution subject in the above method 200 according to the instructions in the computer program, including:
[0196] determining a first device in at least one edge node that cannot operate normally;
[0197] Obtain a first communication link model of a first edge node to which the first device belongs, where the first communication link model includes a detection module in a central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device;
[0198] The operating status of each device in the first communication link model is detected to determine link fault information, where the link fault information includes information about the faulty device in the first communication link model.
[0199] In some embodiments of the present application, the processor 32 may include but is not limited to:
[0200] General-purpose processor, Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.
[0201] In some embodiments of the present application, the memory 31 includes but is not limited to:
[0202] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).
[0203] In some embodiments of the present application, the computer program may be divided into one or more modules, which are stored in the memory 31 and executed by the processor 32 to implement the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device 30.
[0204] Optionally, as shown in FIG8 , the electronic device 30 may further include:
[0205] The communication interface 33 can be connected to the processor 32 or the memory 31 .
[0206] The processor 32 may control the communication interface 33 to communicate with other devices. Specifically, the processor 32 may send information or data to other devices or receive information or data sent by other devices. For example, the communication interface 33 may include a transmitter and a receiver. The communication interface 33 may further include one or more antennas.
[0207] It should be understood that the various components in the electronic device 30 are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.
[0208] According to one aspect of the present application, a communication device is provided, including a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory, so that the encoder executes the method of the above method embodiment.
[0209] According to one aspect of the present application, a computer storage medium is provided, on which a computer program is stored. When the computer program is executed by a computer, the computer is enabled to perform the method of the above-mentioned method embodiment. Alternatively, the present application also provides a computer program product containing instructions. When the computer is executed by the instructions, the computer is enabled to perform the method of the above-mentioned method embodiment.
[0210] According to another aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method of the above-described method embodiment.
[0211] In other words, when implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0212] It is understandable that in the specific implementation of this application, when the above embodiments of this application are applied to specific products or technologies and involve relevant data such as user information, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards.
[0213] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0214] In the several embodiments provided in this application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0215] Modules described as separate components may or may not be physically separate, and components displayed as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected based on actual needs to achieve the purpose of the present embodiment. For example, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module.
[0216] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for fault handling, characterized in that, The method is applied to a distributed cloud scenario, which includes a central node and at least one edge node. The method includes: Determine a first device in the at least one edge node that cannot operate normally; Obtain a first communication link model of a first edge node to which the first device belongs. The first communication link model includes a detection module in the central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device; Detect the operating status of each device in the first communication link model, and determine link failure information, where the link failure information includes information about the faulty device in the first communication link model.
2. The method according to claim 1, wherein The first communication link model includes a tree structure with the detection module as the root node, the second device as the leaf node, and the third device as the intermediate node; wherein, in the tree structure, the root node and the intermediate node record the devices corresponding to their child nodes.
3. The method according to claim 2, wherein The detecting the operating status of each device in the first communication link model and determining the link failure information includes: Starting from the detection module corresponding to the root node, sequentially traverse the operating status of each node in the tree structure, and determine the faulty device in the first communication link model; Determine the link failure information according to the information about the faulty device.
4. The method according to claim 3, characterized in that, The starting from the detection module corresponding to the root node, sequentially traversing the operating status of each node in the tree structure, and determining the faulty device in the first communication link model includes: Detect whether the device corresponding to the current node in the first communication link operates normally; wherein, the current node includes the root node, the intermediate node or the leaf node; If the current node cannot operate normally, determine the device corresponding to the current node as the faulty device.
5. The method according to claim 4, characterized in that, It further includes: If the current node has sibling nodes, detect whether the devices corresponding to the sibling nodes are normal.
6. The method according to claim 4, characterized in that, It further includes: If the current node operates normally and the current node has child nodes, detect whether the devices corresponding to the child nodes of the current node are normal.
7. The method according to any one of claims 1-6, characterized in that, The determining the first device in the at least one edge node that cannot operate normally includes: Obtain information about the first device in the at least one edge node that cannot operate normally from the detection module; wherein, the detection module is used to obtain the operating status of each device in the central node and the at least one edge node.
8. The method according to any one of claims 1-7, characterized in that It further includes: Obtain information about the first device in the at least one edge node that cannot operate normally from the fault list; wherein, the fault list is used to cache information about at least one device that cannot operate normally from the detection module.
9. The method according to claim 8, characterized in that, The obtaining the first communication link model of the first edge node to which the first device belongs includes: According to the information about the first device, obtain the first communication link model of the first edge node to which the first device belongs from the link model pool; wherein, the link model pool includes communication link models of the at least one edge node.
10. The method according to any one of claims 1-9, characterized in that, It further includes: Determine the failure level of the first edge node according to at least one of the weight of the faulty device, the service level information of the first edge node, and the failure duration; Process the faulty device according to the failure level of the first edge node.
11. The method according to claim 10, wherein It further includes: Determine the number of the second devices that cannot operate normally in the first communication link; Determine the failure state model of the first edge node according to the link failure information, the number of the second devices that cannot operate normally, and the service level information of the first edge node; Determine at least one of the weight of the faulty device, the service level information, and the failure duration according to the failure state model.
12. A device for fault handling, characterized in that, The device is applied to a distributed cloud scenario, and the distributed cloud scenario includes a central node and at least one edge node. The device includes: A determination unit, configured to determine a first device that cannot operate normally among the at least one edge node; An acquisition unit, configured to acquire a first communication link model of a first edge node to which the first device belongs. The first communication link model includes a detection module in the central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device; A detection unit, configured to detect the operating state of each device in the first communication link model and determine link failure information. The link failure information includes information about the faulty device in the first communication link model.
13. An electronic device, characterized in that, It includes a processor and a memory. Instructions are stored in the memory. When the processor runs the instructions, the processor executes the method according to any one of claims 1-11.
14. A computer storage medium, characterized in that, It includes instructions that, when running on a computer, cause the computer to execute the method according to any one of claims 1-11.
15. A computer program product, characterized in that, It includes computer program code that, when run by an electronic device, causes the electronic device to execute the method according to any one of claims 1-11.
Citation Information
Patent Citations
Network fault diagnosis method and device
CN114866398A
Equipment anomaly detection system and method based on cloud edge cooperation mode
CN115098330A
Power equipment fault detection method, device and equipment
CN115267409A
Fault processing method and device, electronic equipment and storage medium
CN118869527A
Cloud technology-based fault handling method, cloud management platform and related device
WO2024001299A1
Cited By
Path device state detection method and device
CN120896873A