Fault processing method and apparatus, electronic device, and storage medium
By detecting faulty devices in the communication link model in a distributed cloud scenario, the problem of batch alarms of edge node devices under high failure rate is solved, improving fault handling efficiency and reducing operation and maintenance pressure.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2024-12-06
- Publication Date
- 2026-06-04
Smart Images

Figure CN2024137532_04062026_PF_FP_ABST
Abstract
Description
[Amended according to Rule 26, 26.12.2024] Methods, apparatus, electronic equipment and storage media for fault handling
[0001] This application claims priority to Chinese Patent Application No. 202410101673X, filed on January 24, 2024, entitled “Method, Apparatus, Electronic Device and Storage Medium for Fault Handling”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of cloud computing technology, and specifically to a method, apparatus, electronic device, and storage medium for fault handling in a distributed cloud scenario. Background Technology
[0003] Distributed cloud is an extension of centralized cloud capabilities, aiming to provide ubiquitous cloud computing services and meet the needs of edge data processing and real-time computing for proximity access. Compared to centralized cloud, distributed cloud can be used anywhere, which determines its compact, efficient, and ultra-low-cost characteristics. To reduce costs, distributed cloud nodes support a minimum of 3 devices, and many non-essential capabilities are used remotely via the public network to connect to the central cloud. This significantly reduces the requirements for heat dissipation and power supply in Internet Data Center (IDC) server rooms, inevitably leading to a higher failure rate for cloud nodes.
[0004] In related technologies, the fault handling process for distributed cloud nodes follows the general process of central clouds: all faults are handled in real time. However, the general process is suitable for central cloud scenarios with large node scale and low failure rate, but not for distributed cloud scenarios with small node scale and high failure rate. Specifically, distributed cloud nodes reuse the fault management system of the central cloud through public network links. A failure of the public network link will cause the entire edge node to lose connection, and all devices within the edge node will trigger batch alarms. This batch of false alarms caused by link interruption, even though the devices are working normally, further exacerbates the operation and maintenance pressure. Summary of the Invention
[0005] This application provides a method, apparatus, device, and storage medium for fault handling, which can facilitate the rapid detection of device disconnection in edge nodes caused by link failures and improve fault handling efficiency.
[0006] In a first aspect, embodiments of this application provide a fault handling method, the method being applied to a distributed cloud scenario, the distributed cloud scenario including a central node and at least one edge node, the method comprising:
[0007] Identify the first device among the at least one edge node that is not functioning properly;
[0008] Obtain the first communication link model of the first edge node to which the first device belongs. The first communication link model includes the detection module in the central node, the second device in the first edge node, and the third device for data transmission between the detection module and the second device.
[0009] The operating status of each device in the first communication link model is detected to determine link fault information, which includes information about faulty devices in the first communication link model.
[0010] Secondly, embodiments of this application provide a fault handling apparatus, which is applied to a distributed cloud scenario, the distributed cloud scenario including a central node and at least one edge node, the apparatus comprising:
[0011] A determining unit is configured to determine a first device that is not functioning properly among the at least one edge node;
[0012] The acquisition unit is used to acquire a first communication link model of a first edge node to which the first device belongs. The first communication link model includes a detection module in the central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device.
[0013] The detection unit is used to detect the operating status of each device in the first communication link model and determine link fault information, which includes information about faulty devices in the first communication link model.
[0014] Thirdly, embodiments of this application provide an electronic device, including:
[0015] Processor, adapted to implement computer instructions; and,
[0016] A memory that stores computer instructions adapted for loading by a processor and executing the method described in the first aspect above.
[0017] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer instructions that, when read and executed by a processor of a computer device, cause the computer device to perform the method described in the first aspect.
[0018] Fifthly, embodiments of this application provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method described in the first aspect.
[0019] Through the above technical solution, when a first device in an edge node fails to operate normally, this embodiment of the application obtains a first communication link model corresponding to the first edge node to which the edge node belongs. This first communication link model includes a detection module in the central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device. Then, the operating status of each device in the first communication link model is detected, and the faulty device in the first communication link model is identified. Therefore, this embodiment of the application can detect not only the devices in the edge node when an edge node failure occurs, but also each device on the communication link between the detection module in the central node and the devices in the edge node. This facilitates the rapid detection of device failures in the edge node caused by link failures, improving fault handling efficiency. Furthermore, by quickly detecting link failures, it is beneficial to reduce or avoid batch alarms from devices in the edge node, reducing operational and maintenance pressure. Attached Figure Description
[0020] Figure 1 is a schematic diagram of an optional network architecture in a distributed cloud scenario;
[0021] Figure 2 is a schematic flowchart of a fault handling method provided in an embodiment of this application;
[0022] Figure 3 is a schematic diagram of the network architecture of a distributed cloud scenario provided in an embodiment of this application;
[0023] Figure 4 is a schematic flowchart of another fault handling method provided in an embodiment of this application;
[0024] Figure 5 is a schematic diagram of a fault detection process provided in an embodiment of this application;
[0025] Figure 6 is a schematic diagram of a fault handling process provided in an embodiment of this application;
[0026] Figure 7 is a schematic block diagram of a fault handling device provided in an embodiment of this application;
[0027] Figure 8 is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0028] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0029] It should be understood that in the embodiments of this application, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B solely based on A; B can also be determined based on A and / or other information.
[0030] In the description of this application, unless otherwise stated, "at least one" means one or more, and "multiple" means two or more. Additionally, "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0031] It should also be understood that the descriptions of "first", "second", etc. appearing in the embodiments of this application are only for illustration and to distinguish the objects being described, and there is no order to them. They do not indicate any special limitation on the number of devices in the embodiments of this application, and cannot constitute any limitation on the embodiments of this application.
[0032] It should also be understood that specific features, structures, or characteristics relating to embodiments in the specification are included in at least one embodiment of this application. Furthermore, these specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0033] Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion, such that a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such processes, methods, products, or devices.
[0034] Figure 1 is a schematic diagram of an optional network architecture in a distributed cloud scenario.
[0035] As shown in Figure 1, a distributed cloud scenario includes central nodes and edge nodes. The central node, also known as the central cloud, includes at least one device, such as device 1, device 2, etc. The edge node, also known as the edge cloud, includes at least one device, such as device 1', device 2', etc. There can be multiple edge nodes 120.
[0036] Distributed cloud can provide various cloud computing services to terminals. Terminals, also known as user terminals, include, but are not limited to, mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle terminals, and aircraft.
[0037] In the fault handling process of a distributed cloud scenario, the central node is mainly responsible for the entire lifecycle of fault management; the edge nodes are responsible for responding to the central node's probe commands and reporting monitoring data. As shown in Figure 1, a complete fault management link includes a detection module 110, a response module 120, and a fault handling module 130.
[0038] The detection module 110 is responsible for initiating device status checks. The detection module 110 is deployed at the central node. Optionally, during device status checks, the detection module 110 can operate on a business-by-business basis, with each business managing its own devices independently. For example, business 1 manages its corresponding device 1, and business 2 manages its corresponding device 2. For instance, the detection module 110 can have a built-in timer; when the system is running, the detection module 110 periodically initiates fault detection heartbeats from within the business to check the liveness of the target device. The detection module 110 can also scan the monitoring data reported by the devices to check if the devices are functioning normally.
[0039] In a distributed scenario, to reduce operational costs, edge node devices reuse the fault detection capabilities of the detection module 110 in the central node. For example, as shown in Figure 1, service 1 can also manage corresponding device 1', and service 2 can also manage corresponding device 2'. This includes periodically initiating fault detection heartbeats within the service to check the liveness of corresponding devices in the edge nodes, and scanning the monitoring data reported by the devices in the edge nodes. When a device fails to detect a heartbeat or displays abnormal monitoring data, it can be considered that the device has malfunctioned.
[0040] The response module 120 is responsible for responding to probe heartbeats initiated by the service and reporting monitoring data to the corresponding service. The response module 120 is deployed on the device. For example, response modules 120 are deployed on device 1 and device 2 in the central node, and on device 1' and device 2' in the edge node. During device fault monitoring, the response module 120 actively reports the real-time status monitoring data of the device and passively responds to probe heartbeats. The internal devices of the central node can directly transmit data with the detection module 110. Unlike the devices within the central node, when the detection module 110 in the central node issues a detection command, the data needs to pass through a complex communication link to reach the device in the edge node. As shown in Figure 1, this communication link may include, for example, a routing module, a sending module, and a public network link in the central node, and a receiving module and a routing module in the edge node.
[0041] The fault handling module 130 is responsible for fault repair and is deployed at the central node. As shown in Figure 1, the fault handling module may include an alarm module and a processing module. When a fault is detected in a device, the corresponding business module in the detection module 110 immediately triggers the alarm module to issue an alarm, and then the relevant processing module performs fault handling.
[0042] In related technologies, during edge node fault handling, the detection heartbeat and detection data between the detection module 110 and the response module 120 need to be transmitted through a complex communication link. This communication link is a public network control link, with a long transmission link and many forwarding devices, increasing the uncertainty of data transmission. If the link fails, the entire edge node will lose connection, unable to respond to detection heartbeats or report status monitoring data. Furthermore, link failure can also trigger batch alarms from devices in the edge node. In daily fault handling, it is necessary to first filter out the real alarms from a large number of invalid alarms. This problem of a single device failure causing batch alarms from all devices in the edge node will seriously reduce fault handling efficiency and interfere with normal node operation and maintenance.
[0043] In view of this, embodiments of this application provide a fault handling method, apparatus, electronic device, and storage medium applicable to distributed cloud scenarios, which can facilitate the rapid detection of device disconnection in edge nodes caused by link failures and improve fault handling efficiency.
[0044] Specifically, the distributed cloud scenario includes a central node and at least one edge node. First, a first device that cannot operate normally in the at least one edge node is identified. Then, a first communication link model of the first edge node to which the first device belongs is obtained. The first communication link model includes a detection module in the central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device. After that, the operating status of each device in the first communication link model is detected to determine link fault information, which includes information about the faulty device in the first communication link model.
[0045] Therefore, in this embodiment, when the first device in an edge node fails to operate normally, a first communication link model corresponding to the first edge node to which the edge node belongs is obtained. This first communication link model includes a detection module in the central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device. Then, the operating status of each device in the first communication link model is detected, and the faulty device in the first communication link model is identified. Therefore, this embodiment can detect not only the devices in the edge node when an edge node failure occurs, but also each device on the communication link between the detection module in the central node and the devices in the edge node. This facilitates rapid detection of device failures in the edge node caused by link failures, improving fault handling efficiency. Furthermore, by quickly detecting link failures, it is possible to reduce or avoid batch alarms from devices in the edge node, reducing operational and maintenance pressure.
[0046] The technical solutions of the embodiments of this application will be described in detail below through some examples. The following embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0047] Figure 2 is a schematic flowchart of a fault handling method 200 applied in a distributed cloud scenario according to an embodiment of this application. Method 200 can be executed by any electronic device with data processing capabilities, and this application does not limit it.
[0048] For example, method 200 can be applied to the network architecture in the distributed cloud scenario shown in Figure 3. As shown in Figure 3, the network architecture includes a central node and at least one edge node, as well as a detection module 110, a response module 120, a fault handling module 130, and a link diagnosis module 310. The detection module 110, response module 120, and fault handling module 140 are similar to the corresponding modules in Figure 1, and can be referred to the relevant descriptions in Figure 1, which will not be repeated here. The link diagnosis module 310 is deployed in the central node.
[0049] Optionally, the network architecture may also include at least one of the following: a fault list 320, a link model 330, a link entry module 340, and a fault orchestration module 350. The fault list 320, link model 330, link entry module 340, and fault orchestration module 350 may be deployed in the central node.
[0050] The method 200 will now be described with reference to the network architecture shown in Figure 3. As shown in Figure 2, the method 200 may include steps 210 to 230.
[0051] 210, Identify a first device in at least one edge node that is not functioning properly.
[0052] For example, referring to Figure 3, each service (such as service 1, service 2) in the detection module 110 can initiate fault detection on the response module 120 of the devices (such as device 1', device 2') in the edge node, detect the liveness of the corresponding devices in the edge node, and scan the monitoring data reported by the devices in the edge node. When a device in the edge node fails to perform a heartbeat detection or has abnormal monitoring data, the detection module 110 determines that the device cannot operate normally or that the device is abnormal. At this time, the first device may be faulty, or the device (such as a forwarding device) on the communication link between the response device of the first device and the detection module 110 may be faulty.
[0053] The failure of heartbeat detection or abnormal monitoring data may include not receiving the device's heartbeat or monitoring data, or receiving heartbeat and monitoring data but the monitoring data itself is abnormal. It is understood that the first device identified as malfunctioning in step 210 may or may not be faulty. For example, if there is a fault in the communication link between the first device's response device and the detection model 110, the first device may not be faulty.
[0054] Optionally, the link diagnostic module 310 can obtain information from at least one first device in the edge node that is not functioning properly, which is detected by the detection module 110.
[0055] In one possible implementation, the detection module 110 can send information about the first malfunctioning device among the detected edge nodes to the link diagnostic module in real time.
[0056] In another possible implementation, the detection module 110 can send information about the first malfunctioning device among the detected edge nodes to the fault list 320 for caching in real time. The fault list 320 includes information about at least one malfunctioning device.
[0057] In some embodiments, when the communication link between the response module 120 in the edge node and the detection module 110 in the central node fails, multiple devices in the edge node may be detected as not operating normally at the same time, and these multiple devices will be added to the fault list 320.
[0058] Accordingly, the link diagnostic module 310 can obtain information about the first device from the fault list 320 according to a preset method, thereby identifying the first device in the edge node that cannot operate normally. For example, the link diagnostic module 310 can read the information of the first devices that cannot operate normally in the fault list 320 one by one according to this preset method. As an example, this preset method may include timed acquisition or periodic acquisition, etc., and this embodiment of the application does not limit this.
[0059] Specifically, for scenarios with small edge node size and high failure rate, the impact of a failure is relatively small, mostly affecting only the internal structure of the node. Based on this, when the detection module 110 detects a failure in an edge node, it adds the information of the malfunctioning device to the fault list instead of directly triggering a fault alarm, thus satisfying the requirement that failures of edge nodes with low availability requirements do not need to be processed in real time.
[0060] 220. Obtain the first communication link model of the first edge node to which the first device belongs. The first communication link model includes a detection module in the central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device.
[0061] For example, continuing to refer to Figure 3, after the link diagnosis module 310 determines the first device that cannot operate normally, it can determine the first edge node to which the first device belongs, and further obtain the first communication link model of the first edge node.
[0062] Specifically, the first communication link model includes a communication network between a central node (such as a detection module) and an edge node (such as a response module). The first communication link model can be determined based on the basic information of each device in the communication link between the central node (such as a detection module) and the edge node (such as a response module), as well as the connection relationship between each device.
[0063] In some embodiments, the first communication link model includes a tree structure with the detection module as the root node, the second device in the first edge node as the leaf node, and the third device used for data transmission between the detection module and the second device as the intermediate node. The root node and intermediate nodes in the tree structure record the devices corresponding to their child nodes. The device corresponding to a node's child node is the subsequent device of that node; leaf nodes have no subsequent devices. Optionally, the device corresponding to the root node can be called the root device, and the device corresponding to the leaf node can be called the leaf device.
[0064] Specifically, in the communication link between the detection module of the central node and the devices of the edge nodes, the next-hop downstream device of the root node or intermediate node is its subsequent device. By recording the devices corresponding to its child nodes, i.e., the subsequent devices, by the root node and intermediate nodes, the communication link model can accurately and effectively represent the communication network between the central node and the edge nodes.
[0065] Optionally, the second device may include the first device described above.
[0066] Specifically, the second device is a device within the first edge node, and the first device belongs to the first edge node; that is, the first device is a device within the first edge node that is not functioning properly. When all devices in the first edge node are not functioning properly, the first device and the second device are the same device. When some devices in the first edge node are not functioning properly, the first device is a subset of the second device.
[0067] In some embodiments, when multiple devices within an edge node are simultaneously detected as malfunctioning, all of these devices are added to the fault list 320. Furthermore, since these multiple devices belong to the same edge node, only the communication link network of that edge node needs to be obtained.
[0068] In some embodiments, a communication link model of at least one edge node may be obtained; and the communication link model of at least one edge node may be saved to a link model pool. The link model pool may include the communication link models of at least one edge node. Optionally, the link model pool may include the communication link models of all edge nodes. In this case, the first communication link model of the first edge node to which the first device belongs can be obtained from the link model pool based on the information of the first device.
[0069] For example, continuing to refer to Figure 3, the link entry module 340 is responsible for constructing the communication link model of the edge node. For example, it obtains the communication link model of at least one edge node in advance and saves the communication link model of at least one edge node to the link model pool 330. For example, the communication link model can be saved using the edge node identifier as the key.
[0070] In some embodiments, the link entry module 340 can also enter the service level provided by the edge node. Different fault handling strategies with different response levels can be provided for different service levels. Optionally, the service levels can be stored in a link model pool.
[0071] It should be understood that method 200 is described using the example of the first device in the first edge node failing to operate normally. The fault handling process for devices in other edge nodes in a distributed cloud scenario that fail to operate normally is similar to the handling process for the first device in the first edge node, and can be found in the relevant descriptions. This application embodiment will not repeat the details.
[0072] 230. Detect the operating status of each device in the first communication link model and determine the link fault information, which includes the information of the faulty device in the first communication link model.
[0073] For example, continuing to refer to Figure 3, the link diagnosis module 310 can, based on the operating status of each device in the edge node, sequentially detect the operating status of each device in the communication link model of the edge node to which the malfunctioning device belongs. That is, it sequentially detects the detection module in the central node, the devices in the edge nodes, and each device (such as forwarding devices) on the communication link between the detection module and the devices in the edge nodes, until the faulty device is detected. Then, the faulty link information can be determined based on the information of the faulty device.
[0074] In some embodiments, for multiple non-operating devices within the same edge node in the fault list 320, by detecting the operating status of each device in the communication link network of the edge node, batch alarms caused by the same device fault can be merged and processed to avoid alarm abuse.
[0075] Therefore, this embodiment of the application can detect not only the devices in the edge nodes when a fault occurs, but also the devices on the communication link between the detection module in the central node and the devices in the edge nodes. This facilitates the rapid detection of device faults in the edge nodes caused by link failures, improving fault handling efficiency. Furthermore, by quickly detecting link failures, it is possible to reduce or avoid batch alarms from devices in the edge nodes, reducing operational and maintenance pressure.
[0076] Furthermore, by adding the information of detected malfunctioning devices to the fault list, and then reading the information of malfunctioning devices from the fault list, fault detection is performed on each device in the communication link model of the edge node, instead of directly triggering fault alarms. This allows for the convergence of fault handling entry points, merging of batch alarms caused by the same device fault, and avoiding alarm abuse.
[0077] In some embodiments, when the first communication link model includes a tree structure with the detection module as the root node, the second device in the first edge node as the leaf node, and the third device used for data transmission between the detection module and the second device as the intermediate node, the operating status of each node in the tree structure can be traversed sequentially, starting from the detection module (i.e., the root device) corresponding to the root node, thereby identifying the faulty device in the first communication link model. Then, the faulty link information can be determined based on the information of the faulty device.
[0078] For example, the node currently being detected in the first communication link model can be referred to as the current node. The current node can be a root node, an intermediate node, or a leaf node. For the current node, it can be checked whether its corresponding device is operating normally. If the current node is not operating normally, the device corresponding to the current node is identified as a faulty device.
[0079] Optionally, if the current node is running normally and has child nodes, then check whether the device corresponding to the child node of the current node (i.e., the subsequent device) is normal. In this case, the next current node can be updated to the device corresponding to that child node.
[0080] Optionally, if the current node is running normally and has no child nodes, its sibling nodes can be checked, specifically whether the device corresponding to that sibling node is functioning correctly. In this case, the next current node can be updated to be that sibling device.
[0081] Optionally, if the current node is not functioning properly, its sibling nodes can also be checked to complete a traversal of the tree structure corresponding to the entire first communication link model, avoiding overlooking any faulty devices among the sibling nodes. When the current node has no sibling devices, the traversal of the tree structure corresponding to the entire first communication link model is complete.
[0082] Therefore, this embodiment of the application sequentially traverses each device in the first communication link model. When a device responds normally and the detection data is normal, the device is considered normal, and the subsequent devices of that device can be probed along the first communication link model. When a device does not respond or the detection data is abnormal, the device is marked as a faulty device, and the subsequent devices of that device can be ignored. Instead, the sibling devices of that device are traversed, thereby completing the traversal and inspection of the first communication link model, which is beneficial for quickly discovering link faults.
[0083] Furthermore, this embodiment of the application can collect fault metadata by detecting the survival status or detection data of each device in the communication link model, thereby obtaining information about the faulty device. The fault metadata can be used to identify the fault type, thereby improving the accuracy of fault identification.
[0084] Optionally, the information about the faulty device may also include the time of the fault. The time of the fault can be used to determine the duration of the fault.
[0085] In some embodiments, referring to FIG4, method 200 may further include steps 240 and 250.
[0086] 240, Determine the number of second devices that are not functioning properly in the first communication link.
[0087] For example, after traversing and checking all devices in the first communication link model, the number of unreachable leaf devices in the first communication link model can be counted. These unreachable leaf devices are devices in the first edge node that cannot operate normally. The second set of devices that cannot operate normally is a subset of the second devices in step 220 above. The definition of devices that cannot operate normally can be found in the relevant description above.
[0088] 250. Based on the number of faulty devices, the number of second devices that cannot operate normally, and the service level information of the first edge node, determine the fault state model of the first edge node.
[0089] Specifically, the fault state model can be determined based on information about faulty devices in the first communication link model, the number of unreachable leaf devices in the first communication link model, and the service level information of the first edge node.
[0090] Optionally, the fault state model can also be saved to the link state pool. For example, the fault state model can be deduplicated and saved to the link state pool using the identifier of the faulty device as the key. For instance, if the key in the link state pool already includes the identifier of a faulty device, the fault state model corresponding to that faulty device in the link state pool can be updated.
[0091] This application embodiment saves the fault state model to the link state pool instead of directly triggering fault alarms or modifying faults. This allows for the requirement that faults in edge nodes with low availability requirements do not need to be processed in real time, which helps reduce operation and maintenance costs.
[0092] If all devices are found to be functioning normally after traversing the first communication link model, the system is operating normally. If a faulty device is found after traversing the communication link model, it indicates a system fault, requiring further processing. For example, the fault handling module can perform fault repair based on the link status.
[0093] In related technologies, the standard fault handling process involves manually confirming the fault a second time before selecting a fault handling module for processing. This approach requires 24 / 7 human support and is suitable for centralized clouds with stable operating environments and low failure rates, but not for distributed cloud scenarios with many uncontrollable factors, high failure rates, and limited fault impact. Furthermore, as the amount of data on distributed cloud nodes increases, the cost of manual processing rises rapidly.
[0094] In view of this, the embodiments of this application further determine the fault level of the first edge node based on the detection results of the first communication link model, and then process the faulty device according to the fault level. By classifying and processing faults, it is beneficial to handle high-priority faults immediately and downgrade low-priority faults, thereby effectively reducing the number of times that need to be manually handled in real time, improving the overall fault handling efficiency, and reducing maintenance costs.
[0095] In some embodiments, continuing to refer to FIG4, method 200 may further include steps 260 and 270.
[0096] 260. Determine the fault level of the first edge node based on at least one of the following: the weight of the faulty device, the service level information of the first edge node, and the fault duration.
[0097] For example, continuing to refer to Figure 3, the fault orchestration module 350 can determine the fault level of the first edge node based on at least one of the weight of the faulty device, the service level information of the first edge node, and the fault duration.
[0098] For example, the fault orchestration module 350 can obtain link fault information from the link diagnosis module 310, and determine at least one of the following based on the link fault information: the weight of the faulty device, the service level information of the first edge node, and the fault duration. For instance, the fault orchestration module 350 can obtain a fault state model from the link state pool according to a preset method. For example, this preset method can include timed acquisition or periodic acquisition, etc., and this embodiment does not limit this. Then, based on the fault state model, at least one of the following can be obtained: the weight of the faulty device, the service level information of the first edge node, and the fault duration.
[0099] In some embodiments, a grading score can be determined based on the weight of the faulty device, the service level information of the first edge node, and the fault duration, and then the fault level can be determined based on the grading score. For example, the grading score S can be represented by the following formula (1):
[0100] Here, `error_count` represents the weight of the faulty device, and the weight of each device can be set according to its importance in the communication link model; `∑error_count` is the sum of the weights of all subsequent devices affected by the faulty device; `level_vip` represents the service level information, which is the availability level of the pre-entered edge node. The higher the node availability requirement, the higher the service level; `time` represents the fault duration coefficient, which is positively correlated with the fault duration; and `max_score` is the maximum score. The larger the calculated grade score `S`, the higher the processing priority.
[0101] In formula (1), when the service level of an edge node is greater than or equal to 3, the grading score is the highest score, max_score. When the service level of an edge node is less than 3, the grading score is calculated based on the number and weight of the faulty devices, the service level information of the edge node, and the fault duration. In this case, the more devices affected by the fault, the higher the service level of the edge node, or the longer the fault duration, the higher the grading score.
[0102] For example, the value of max_score can be 30, but this application does not limit it.
[0103] 270. Based on the fault level of the first edge node, handle the faulty equipment.
[0104] For example, continuing to refer to Figure 3, the fault orchestration module 350 can determine a strategy for handling faulty devices based on the fault level of the first edge node. Specifically, different response processes can be executed for different fault levels. For example, high-priority faults can be handled immediately, while low-priority faults can be downgraded, such as waiting for idle time to process them.
[0105] For example, for the grading score obtained according to the above formula (1), when the score reaches max_score, the fault needs to be handled immediately; when the service level is greater than or equal to 3, all faults need to provide a real-time response; when the service level is less than 3, it is handled according to the calculated grading score. For example, the more devices affected by the fault, the higher the service level of the edge node, or the longer the fault duration, the higher the grading score of the device, and the higher the corresponding processing priority.
[0106] For example, high-priority faults can be responded to 24 / 7, while faults with a smaller impact (e.g., fewer faulty devices) can be addressed during off-peak hours. In some instances, for faults whose types can be accurately identified and for which automatic handling methods exist, automatic recovery can be prioritized to reduce manual intervention. For fault types that cannot be automatically recovered, a standard processing flow can be implemented, requiring manual processing after secondary confirmation. Therefore, the embodiments of this application can effectively reduce the number of times faults need to be manually handled in real time, improve overall fault handling efficiency, and reduce maintenance costs.
[0107] Figure 5 is a schematic diagram of a fault detection process provided in an embodiment of this application. It should be understood that Figure 5 illustrates steps or operations related to the fault detection process, but these steps or operations are merely examples. Other operations or variations of the operations in Figure 5 can also be performed in this embodiment. Furthermore, the steps in Figure 5 may be performed in a different order than those presented in Figure 5, and it is not necessary to perform all the operations in Figure 5. As shown in Figure 5, the fault detection process includes steps 401 to 408.
[0108] 401, the device in the edge node is not functioning properly.
[0109] Specifically, a fault occurs at an edge node, causing some devices to malfunction. In this case, information about the malfunctioning devices at the edge node can be obtained. The process for obtaining this information can be found in step 210 of Figure 2, and will not be repeated here.
[0110] 402, Load communication link model. M = root device.
[0111] As shown in Figure 5, the communication link model of the edge node to which the malfunctioning device belongs can be obtained from the link model pool. For example, the corresponding communication link model can be loaded based on the edge node ID to which the malfunctioning device belongs.
[0112] One possible implementation involves automatically activating a module to load the communication link model of the edge node to which the malfunctioning device belongs when a malfunctioning device is detected in an edge node. For example, the communication link model can be a tree structure with the detection module in the central node as the root node, where each device records its subsequent devices. When a device has no subsequent devices, it is considered a leaf device. Devices in the edge nodes are leaf devices.
[0113] Specifically, the process of loading the communication link model can be referred to in the relevant description in step 220 of Figure 2, which will not be repeated here.
[0114] After loading the communication link model, the device corresponding to the root node in the communication link model (i.e., the root device) can be configured as the current probe device, i.e., device M. Then, starting from the root device, the liveness status and detection data of each device are probed sequentially along the communication link model until a faulty device is found.
[0115] Specifically, the detection process can be carried out in steps 403 to 407.
[0116] 403, Detection device M.
[0117] Specifically, it can detect the survival status and monitoring data of device M to determine whether device M is a normal device.
[0118] 404 indicates a problem. Determine if device M is functioning correctly and if there are subsequent devices.
[0119] For example, when the device responds normally and the monitoring data is normal, the device is considered a normal device. At this time, it can be determined whether there are any subsequent devices for device M based on the loaded communication link model. If there are subsequent devices for device M, step 406 can be executed.
[0120] When both conditions of M device being normal and the existence of subsequent devices cannot be met simultaneously, such as when M device is not operating normally or when M device does not have subsequent devices, step 405 can be executed.
[0121] 405, determine if a sibling device exists.
[0122] If a sibling device exists, proceed to step 407. If no sibling device exists, complete the traversal check of the communication link model.
[0123] When device M malfunctions, its subsequent devices will also malfunction because they cannot receive and forward data from it. Therefore, there is no need to probe the subsequent devices of device M. Instead, it is possible to determine if device M has sibling devices, allowing for sequential traversal of each device in the communication link model.
[0124] When device M is running normally, but has no subsequent devices, device M is already a leaf node in the current communication network branch, and all nodes in that branch are functioning correctly. At this point, it can be determined whether device M has sibling devices, allowing for sequential traversal of each device in the communication link model.
[0125] 406, M = Subsequent equipment.
[0126] Specifically, the subsequent device of device M in step 403 can be configured as the current device M.
[0127] 407, M = sibling device.
[0128] Specifically, the sibling device of the M device in step 403 can be configured as the current M device.
[0129] After completing steps 406 or 407, step 403 can be executed to perform the detection of device M.
[0130] In summary, by iteratively executing steps 403 to 407, a depth-first approach can be used to traverse the tree structure sequentially, starting from the root device, and check whether each device is functioning correctly. When a device responds normally and the detection data is normal, the device is considered normal, and the process can continue to probe its subsequent devices along the communication link model. When a device does not respond or its detection data is abnormal, the device is marked as faulty, and its subsequent devices can be ignored. Instead, the process can proceed to traverse the device's sibling devices, thus completing the traversal and check of the communication link model.
[0131] If all devices are found to be functioning normally after traversing the communication link model, the system is working properly. If a faulty device is found after traversing the communication link model, it indicates that the system has a fault and further processing is required for the faulty device, such as by executing step 408 below.
[0132] 408, Save link status.
[0133] For example, faulty devices detected in the communication link model can be saved as link states in the link state pool.
[0134] Optionally, after traversing the communication link model, the number of unreachable leaf devices in the communication link model can also be counted. Optionally, the link state can also include the number of unreachable leaf devices in the communication link model. Optionally, the link state can also include the service level information of edge nodes.
[0135] One feasible approach is to encapsulate information such as the number of faulty devices, unreachable leaf devices, and service level information of edge nodes in the communication link model into a link state model, using the identifier of the faulty device as the key, and save it in the link state pool. Correspondingly, the fault handling module can retrieve each link state model from the link state pool and handle the fault according to that model.
[0136] Therefore, in this embodiment of the application, when there are malfunctioning devices in the edge nodes, by detecting each device in the communication link model of the edge nodes, the fault handling entry point can be converged, and batch alarms caused by the same device failure in the communication link model can be merged, which helps to avoid alarm abuse. Statistical results show that in the fault handling process provided by this embodiment of the application, the average number of alarms per distributed edge cloud node per month is reduced from 18 to 1.3, reducing invalid alarms by 93% per month.
[0137] Figure 6 is a schematic diagram of a fault handling process provided in an embodiment of this application. It should be understood that Figure 6 illustrates steps or operations related to the fault handling process, but these steps or operations are merely examples. Other operations or variations of the operations in Figure 6 can also be performed in this embodiment. Furthermore, the steps in Figure 6 may be performed in a different order than those presented in Figure 6, and it is not necessary to perform all the operations in Figure 6. As shown in Figure 6, the fault detection process includes steps 501 to 509.
[0138] 501, Loading fault state model.
[0139] For example, fault state models can be loaded from the link state pool periodically. As one implementation approach, the module can have a built-in timer to trigger a full scheduling operation at regular intervals. Specifically, when the scheduling task starts, all unprocessed fault state models can be loaded from the "link state pool".
[0140] Optionally, the processing of each fault state model can occupy an execution thread, so that the scheduling of state fault models can be initiated in batches through multi-threading.
[0141] 502, has the testing equipment been restored?
[0142] To ensure the real-time nature of link diagnostic data and avoid fault orchestration anomalies due to outdated data, link diagnostics can be initiated again before starting fault orchestration to obtain the latest link status. If the detection indicates that the fault has been recovered, proceed to step 503. For faults that have not yet been processed, proceed to step 504.
[0143] 503, fault rejection.
[0144] Specifically, when the fault is detected and found to have been recovered, the corresponding fault state model can be removed from the link state pool and the current scheduling can be terminated.
[0145] 504 - How do I determine if a fault is in progress?
[0146] For faults that have not yet been processed, it can be further determined whether the fault is currently being processed. For faults that are currently being processed, step 508 can be executed. For faults that are not currently being processed, step 505 can be executed.
[0147] 505, fault classification.
[0148] Specifically, because the failure rate in distributed cloud scenarios is much higher than that in centralized cloud scenarios, in order to improve the efficiency of fault handling and ensure the high-priority handling of major faults, it is necessary to classify faults. For example, the fault level of an edge node can be determined based on at least one of the following: the weight of the faulty device, the service level information of the edge node, and the fault duration. For instance, the classification score can be calculated according to the above formula (1), and the fault level can be further determined based on the classification score.
[0149] 506, should I handle this immediately?
[0150] If the fault requires immediate attention, step 507 can be performed. If the fault does not require immediate attention, step 509 can be performed.
[0151] For example, for service levels greater than or equal to 3, a real-time response can be provided, meaning immediate fault handling is required; for service levels less than 3, corresponding handling can be performed based on the calculated grading score. For instance, faults with grading scores below a certain value may not be handled immediately, such as waiting for idle time before processing.
[0152] 507, fault response.
[0153] Specifically, for faults requiring immediate attention, different responses can be implemented based on different fault levels. For example, high-priority faults can be addressed with 24 / 7 support; faults with a small impact (e.g., a small number of faulty devices) can be downgraded. In the fault handling process, faults whose types can be accurately identified and for which automated handling methods exist can be prioritized for automatic recovery, minimizing manual intervention. For faults that cannot be automatically recovered, a standard handling procedure can be initiated, requiring manual verification before manual processing.
[0154] 508, update processing progress.
[0155] Specifically, for faults that are being processed, the processing progress can be refreshed.
[0156] 509, Status Updated.
[0157] Specifically, after a fault is handled, or for faults that do not require immediate handling, a state fault update can be performed. For example, a fault that has been handled can be removed from the link state pool; faults that do not require immediate handling can wait for the next scheduling.
[0158] Therefore, this application embodiment, by performing fault downgrading, upgrading, and automatic recovery processing as needed based on the fault level, can help ensure high-priority handling of major faults, improve overall processing efficiency, and effectively reduce the number of times manual real-time handling is required. Statistical results show that the average number of times manual real-time fault handling is required per distributed edge node per month has decreased from 0.8 times to 0.3 times, and the manual operation and maintenance cost is now close to that of the central cloud.
[0159] The specific embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this application, various simple modifications can be made to the technical solutions of this application, and these simple modifications all fall within the protection scope of this application. For example, the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this application will not describe the various possible combinations separately. Furthermore, various different embodiments of this application can also be arbitrarily combined, as long as they do not violate the spirit of this application, they should also be considered as the content disclosed in this application.
[0160] It should also be understood that, in the various method embodiments of this application, the sequence numbers of the above processes do not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. It should be understood that these sequence numbers can be interchanged where appropriate so that the embodiments of this application described can be implemented in a sequence other than those illustrated or described.
[0161] The method embodiments of this application have been described in detail above with reference to Figures 1 to 6, and the device embodiments of this application have been described in detail below with reference to Figures 7 and 8.
[0162] Figure 7 is a schematic block diagram of a fault handling apparatus 10 according to an embodiment of this application. The apparatus 10 can be used in a distributed cloud scenario, which includes a central node and at least one edge node. As shown in Figure 7, the apparatus 10 may include a determining unit 11, an acquiring unit 12, and a detecting unit 13.
[0163] Determining unit 11 is used to determine the first device that cannot operate normally among the at least one edge node;
[0164] The acquisition unit 12 is used to acquire the first communication link model of the first edge node to which the first device belongs. The first communication link model includes the detection module in the central node, the second device in the first edge node, and the third device for data transmission between the detection module and the second device.
[0165] The detection unit 13 is used to detect the operating status of each device in the first communication link model and determine link fault information, which includes information about faulty devices in the first communication link model.
[0166] In some embodiments, the first communication link model includes a tree structure with the detection module as the root node, the second device as the leaf node, and the third device as the intermediate node; wherein the root node and intermediate node in the tree structure record the devices corresponding to their child nodes.
[0167] In some embodiments, the detection unit 13 is specifically used for:
[0168] Starting from the detection module corresponding to the root node, the operating status of each node in the tree structure is traversed sequentially to determine the faulty device in the first communication link model;
[0169] Based on the information of the faulty device, the link fault information is determined.
[0170] In some embodiments, the detection unit 13 is specifically used for:
[0171] Detect whether the device corresponding to the current node in the first communication link is operating normally; wherein, the current node includes the root node, the intermediate node, or the leaf node;
[0172] If the current node cannot operate normally, the device corresponding to the current node is identified as the faulty device.
[0173] In some embodiments, the detection unit 13 is further configured to:
[0174] If the current node has sibling nodes, then check whether the device corresponding to the sibling node is functioning properly.
[0175] In some embodiments, the detection unit 13 is further configured to:
[0176] If the current node is running normally and the current node has child nodes, then check whether the device corresponding to the child node of the current node is normal.
[0177] In some embodiments, the determining unit 11 is specifically used for:
[0178] The detection module obtains information about the first device that is not functioning properly in the at least one edge node; wherein, the detection module is used to obtain the operating status of each device in the central node and the at least one edge node.
[0179] In some embodiments, the device 10 further includes a cache unit for:
[0180] Information on the first device that is not functioning properly in the at least one edge node is obtained from the fault list; wherein the fault list is used to cache information on at least one device that is not functioning properly from the detection module.
[0181] In some embodiments, the acquisition unit 12 is specifically used for:
[0182] Based on the information of the first device, the first communication link model of the first edge node to which the first device belongs is obtained from the link model pool; wherein, the link model pool includes the communication link models of the at least one edge node.
[0183] In some embodiments, the apparatus 10 further includes a processing unit for:
[0184] The fault level of the first edge node is determined based on at least one of the weight of the faulty device, the service level information of the first edge node, and the fault duration.
[0185] The faulty device is processed according to the fault level of the first edge node.
[0186] In some embodiments, the determining unit 11 is further configured to:
[0187] Determine the number of second devices that are not functioning properly in the first communication link;
[0188] Based on the link failure information, the number of the second devices that cannot operate normally, and the service level information of the first edge node, determine the failure state model of the first edge node;
[0189] Based on the fault state model, at least one of the following is determined: the weight of the faulty device, the service level information, and the fault duration.
[0190] It should be understood that the device embodiments and method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, they will not be repeated here. Specifically, when the fault handling device 10 in this embodiment can correspond to the fault handling method 200 of the present application embodiment, the foregoing and other operations and / or functions of each module in the device 10 are respectively to implement the corresponding processes in the various methods in FIG2. For the sake of brevity, they will not be repeated here.
[0191] The apparatus and system of this application embodiments have been described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that these functional modules can be implemented in hardware, in software instructions, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in this application can be completed by integrated logic circuits in the processor's hardware and / or by software instructions. The steps of the methods disclosed in this application embodiments can be directly embodied as being executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. Optionally, the software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps in the above method embodiments.
[0192] Figure 8 is a schematic block diagram of the electronic device 30 provided in an embodiment of this application.
[0193] As shown in Figure 8, the electronic device 30 may include:
[0194] The system includes a memory 31 and a processor 32. The memory 31 stores computer programs and transfers the program code to the processor 32. In other words, the processor 32 can retrieve and run the computer programs from the memory 31 to implement the methods described in the embodiments of this application.
[0195] For example, the processor 32 can be used to execute the steps of each execution entity in the method 200 described above according to the instructions in the computer program, including:
[0196] Identify the first device in at least one edge node that is not functioning properly;
[0197] Obtain a first communication link model of the first edge node to which the first device belongs. The first communication link model includes a detection module in the central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device.
[0198] The operating status of each device in the first communication link model is detected to determine link fault information, which includes information about faulty devices in the first communication link model.
[0199] In some embodiments of this application, the processor 32 may include, but is not limited to:
[0200] General-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0201] In some embodiments of this application, the memory 31 includes, but is not limited to:
[0202] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0203] In some embodiments of this application, the computer program may be divided into one or more modules, which are stored in the memory 31 and executed by the processor 32 to perform the method provided in this application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the electronic device 30.
[0204] Optionally, as shown in Figure 8, the electronic device 30 may further include:
[0205] Communication interface 33, which can be connected to processor 32 or memory 31.
[0206] The processor 32 can control the communication interface 33 to communicate with other devices; specifically, it can send information or data to other devices or receive information or data sent by other devices. For example, the communication interface 33 may include a transmitter and a receiver. The communication interface 33 may further include antennas, and the number of antennas may be one or more.
[0207] It should be understood that the various components in the electronic device 30 are connected through a bus system, which includes a data bus, a power bus, a control bus, and a status signal bus.
[0208] According to one aspect of this application, a communication device is provided, including a processor and a memory for storing a computer program, the processor for calling and running the computer program stored in the memory, causing the encoder to perform the method of the above-described method embodiment.
[0209] According to one aspect of this application, a computer storage medium is provided that stores a computer program thereon, which, when executed by a computer, enables the computer to perform the methods of the above-described method embodiments. Alternatively, embodiments of this application also provide a computer program product containing instructions that, when executed by a computer, cause the computer to perform the methods of the above-described method embodiments.
[0210] According to another aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method described in the above-described method embodiments.
[0211] In other words, when implemented using software, it can be implemented wholly or partially in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0212] It is understood that in the specific implementation of this application, when the above embodiments of this application are applied to specific products or technologies and involve user information and other related data, user permission or consent is required, and the collection, use and processing of related data must comply with relevant laws, regulations and standards.
[0213] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0214] In the several embodiments provided in this application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or modules may be electrical, mechanical, or other forms.
[0215] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. For example, the functional modules in the various embodiments of this application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0216] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for troubleshooting, characterized in that, The method is applied to a distributed cloud scenario, which includes a central node and at least one edge node. The method includes: Identify the first device among the at least one edge node that is not functioning properly; Obtain the first communication link model of the first edge node to which the first device belongs. The first communication link model includes the detection module in the central node, the second device in the first edge node, and the third device for data transmission between the detection module and the second device. The operating status of each device in the first communication link model is detected to determine link fault information, which includes information about faulty devices in the first communication link model.
2. The method according to claim 1, characterized in that, The first communication link model includes a tree structure with the detection module as the root node, the second device as the leaf node, and the third device as the intermediate node; wherein, the root node and the intermediate node in the tree structure record the devices corresponding to their child nodes.
3. The method according to claim 2, characterized in that, The step of detecting the operating status of each device in the first communication link model and determining link fault information includes: Starting from the detection module corresponding to the root node, the operating status of each node in the tree structure is traversed sequentially to determine the faulty device in the first communication link model; Based on the information of the faulty device, the link fault information is determined.
4. The method according to claim 3, characterized in that, The step of starting from the detection module corresponding to the root node, sequentially traversing the operating status of each node in the tree structure, and determining the faulty device in the first communication link model includes: Detect whether the device corresponding to the current node in the first communication link is operating normally; wherein, the current node includes the root node, the intermediate node, or the leaf node; If the current node cannot operate normally, the device corresponding to the current node is identified as the faulty device.
5. The method according to claim 4, characterized in that, Also includes: If the current node has sibling nodes, then check whether the device corresponding to the sibling node is functioning properly.
6. The method according to claim 4, characterized in that, Also includes: If the current node is running normally and the current node has child nodes, then check whether the device corresponding to the child node of the current node is normal.
7. The method according to any one of claims 1-6, characterized in that, The determination of the first device that is not functioning properly among the at least one edge node includes: The detection module obtains information about the first device that is not functioning properly in the at least one edge node; wherein, the detection module is used to obtain the operating status of each device in the central node and the at least one edge node.
8. The method according to any one of claims 1-7, characterized in that, Also includes: Information on the first device that is not functioning properly in the at least one edge node is obtained from the fault list; wherein the fault list is used to cache information on at least one device that is not functioning properly from the detection module.
9. The method according to claim 8, characterized in that, The step of obtaining the first communication link model of the first edge node to which the first device belongs includes: Based on the information of the first device, the first communication link model of the first edge node to which the first device belongs is obtained from the link model pool; wherein, the link model pool includes the communication link models of the at least one edge node.
10. The method according to any one of claims 1-9, characterized in that, Also includes: The fault level of the first edge node is determined based on at least one of the weight of the faulty device, the service level information of the first edge node, and the fault duration. The faulty device is processed according to the fault level of the first edge node.
11. The method according to claim 10, characterized in that, Also includes: Determine the number of second devices that are not functioning properly in the first communication link; Based on the link failure information, the number of second devices that cannot operate normally, and the service level information of the first edge node, determine the failure state model of the first edge node; Based on the fault state model, at least one of the following is determined: the weight of the faulty device, the service level information, and the fault duration.
12. A fault handling apparatus, characterized in that, The device is applied to a distributed cloud scenario, which includes a central node and at least one edge node, and the device includes: A determining unit is configured to determine a first device that is not functioning properly among the at least one edge node; The acquisition unit is used to acquire a first communication link model of a first edge node to which the first device belongs. The first communication link model includes a detection module in the central node, a second device in the first edge node, and a third device for data transmission between the detection module and the second device. The detection unit is used to detect the operating status of each device in the first communication link model and determine link fault information, which includes information about faulty devices in the first communication link model.
13. An electronic device, characterized in that, The method includes a processor and a memory, wherein the memory stores instructions, and when the processor executes the instructions, it causes the processor to perform the method according to any one of claims 1-11.
14. A computer storage medium, characterized in that, Includes instructions that, when run on a computer, cause the computer to perform the method of any one of claims 1-11.
15. A computer program product, characterized in that, It includes computer program code that, when executed by an electronic device, causes the electronic device to perform the method according to any one of claims 1-11.