Node management method and device, storage medium and electronic equipment

By automatically detecting faults in node clusters and taking over network devices and child nodes, the problem of low management efficiency under limited node resources is solved, achieving rapid response and efficient fault recovery, and improving the stability and disaster recovery capabilities of network management.

CN121690981APending Publication Date: 2026-03-17CHINA SATELLITE NETWORK INNOVATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-12
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

In complex network environments with limited node resources, existing technologies cannot guarantee real-time management and control between nodes, resulting in low node management efficiency.

Method used

When a fault is detected by the target parent node, the system automatically takes over the network devices and child nodes connected to the target node and sends node change messages to update the parent node information, thereby realizing a node management method and device that ensures the continuity and efficiency of management relationships.

Benefits of technology

It reduces node failure response time and management recovery complexity, improves network management stability and disaster recovery capabilities, and enhances the efficient management and fault recovery capabilities of nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121690981A_ABST
    Figure CN121690981A_ABST
Patent Text Reader

Abstract

The invention discloses a node management method and device, a storage medium and electronic equipment. Specifically, the method executed by a target father node comprises the steps that when it is detected that a target node breaks down, network equipment connected with the target node and a target child node are taken over, the target node, the target child node and the target father node all belong to a node cluster, the node cluster comprises N nodes, the Nth node in the N nodes is the father node of the first node, and the Nth node is the father node of the first node; the ith node is a father node of the (i + 1) th node, the target father node is a father node of the target node, and the target node is a father node of the target child node; and then sending a node change message to instruct the target child node to change the parent node of the target child node. Therefore, the node can efficiently detect the fault of the adjacent node and quickly take over the fault, so that the technical problem of low node management efficiency under the condition of limited node resources is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and more specifically, to a node management method and apparatus, a storage medium, and an electronic device. Background Technology

[0002] In complex network environments, high node reliability and efficient fault recovery mechanisms are crucial for ensuring network service continuity and quality. For example, for resource-constrained edge nodes, a failure not only affects the node's own functionality but also leads to the loss of effective management of its managed devices, thus creating a chain reaction on the stability and performance of the entire network. Existing node management solutions often fail to guarantee real-time management and control between nodes, resulting in low node management efficiency.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This application provides a node management method and apparatus, storage medium and electronic device to at least solve the technical problem of low node management efficiency when node resources are limited.

[0005] According to one aspect of the embodiments of this application, a node management method is provided, applied to a target parent node, comprising: in response to a failure of the target node, taking over the network devices and target child nodes connected to the target node, wherein the target node, the target child node, and the target parent node all belong to a node cluster, the node cluster includes N nodes, the Nth node of the N nodes is the parent node of the 1st node, the i-th node is the parent node of the (i+1)-th node, i is less than N, and i and N are both positive integers, the target parent node is the parent node of the target node, and the target node is the parent node of the target child node; sending a node change message to instruct the target child node to change its own parent node.

[0006] According to one aspect of the embodiments of this application, a node management device is also provided, applied to a target parent node, comprising: a takeover module, configured to take over the network devices and target child nodes connected to the target node in response to a failure of the target node, wherein the target node, the target child node, and the target parent node all belong to a node cluster, the node cluster includes N nodes, the Nth node of the N nodes is the parent node of the 1st node, the i-th node is the parent node of the (i+1)-th node, i is less than N, and i and N are both positive integers, the target parent node is the parent node of the target node, and the target node is the parent node of the target child node; and a transmission module, configured to send a node change message to instruct the target child node to change its own parent node.

[0007] According to another aspect of the embodiments of this application, a node management method is provided, applied to a target node, comprising: determining a target parent node based on its own node number when normal operation is restored, wherein the target parent node is used to take over the network devices and target child nodes connected to the target node when the target node fails, the target node, the target child node, and the target parent node all belong to a node cluster, the node cluster includes N nodes, the Nth node of the N nodes is the parent node of the 1st node, the i-th node is the parent node of the (i+1)-th node, i is less than N, i and N are both positive integers, the target parent node is the parent node of the target node, and the target node is the parent node of the target child node; sending a fault recovery message to the target parent node, wherein the fault recovery message indicates that the target node has resumed normal operation; and sending a node change message to the target child node to instruct the target child node to change its own parent node.

[0008] According to another aspect of the embodiments of this application, a node management device is provided, applied to a target node, comprising: a determination module, configured to determine a target parent node based on its own node number when normal operation is restored, wherein the target parent node is configured to take over the network devices and target child nodes connected to the target node when the target node fails, the target node, the target child node, and the target parent node all belong to a node cluster, the node cluster includes N nodes, the Nth node of the N nodes is the parent node of the 1st node, the i-th node is the parent node of the (i+1)-th node, i is less than N, i and N are both positive integers, the target parent node is the parent node of the target node, and the target node is the parent node of the target child node; a transmission module, configured to send a fault recovery message to the target parent node, wherein the fault recovery message indicates that the target node has resumed normal operation; and send a node change message to the target child node to instruct the target child node to change its own parent node.

[0009] According to another aspect of the embodiments of this application, a node management method is provided, applied to a target child node, comprising: in response to a target node malfunction, receiving a node change message and changing its own parent node from the target node to the target parent node, wherein the target parent node is used to take over the network devices connected to the target node and the target child node when the target node malfunctions, the target node, the target child node and the target parent node all belong to a node cluster, the node cluster includes N nodes, the Nth node of the N nodes is the parent node of the 1st node, the i-th node is the parent node of the (i+1)-th node, i is less than N, i and N are both positive integers, the target parent node is the parent node of the target node, and the target node is the parent node of the target child node; in response to the target node resuming normal operation, receiving the node change message and changing its own parent node from the target parent node to the target node.

[0010] According to another aspect of the embodiments of this application, a node management device is provided, applied to a target child node, comprising: a change module, configured to receive a node change message in response to a target node failure, and change its own parent node from the target node to the target parent node, wherein the target parent node is configured to take over the network devices connected to the target node and the target child node when the target node fails, the target node, the target child node, and the target parent node all belong to a node cluster, the node cluster includes N nodes, the Nth node of the N nodes is the parent node of the 1st node, the i-th node is the parent node of the (i+1)-th node, i is less than N, and i and N are both positive integers, the target parent node is the parent node of the target node, and the target node is the parent node of the target child node; and a transmission module, configured to receive the node change message in response to the target node resuming normal operation, and change its own parent node from the target parent node to the target node.

[0011] According to another aspect of the embodiments of this application, a node management system is provided, comprising: a target parent node, which, in response to a failure of the target node, takes over the network devices and target child nodes connected to the target node, wherein the target node, the target child nodes, and the target parent node all belong to a node cluster, the node cluster comprising N nodes, wherein the Nth node is the parent node of the 1st node, the i-th node is the parent node of the (i+1)-th node, i is less than N, and i and N are both positive integers, the target parent node is the parent node of the target node, and the target node is the parent node of the target child node; and sending a node change message to indicate a change in the target child node. The target node, upon resuming normal operation, determines its target parent node based on its own node number; sends a fault recovery message to the target parent node, indicating that the target node has resumed normal operation; sends a node change message to the target child node to instruct the target child node to change its own parent node; the target child node, in response to a fault in the target node, receives the node change message and changes its own parent node from the target node to the target parent node; in response to the target node resuming normal operation, receives the node change message and changes its own parent node from the target parent node to the target node.

[0012] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer-readable storage medium, and the computer program is configured to execute the above-described node management method at runtime.

[0013] According to another aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the node management method described above.

[0014] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-described node management method through the computer program.

[0015] In this embodiment, the node cluster first includes three nodes: a target node, a target parent node (which is the parent node of the target node), and a target child node (which is the child node of the target node). Then, during node operation, if the target node fails, the target parent node takes over the network devices connected to the target node and the target child nodes originally managed by the target node. The target parent node then sends a node change message to instruct the target child nodes to update their parent node information and make the target parent node their new parent node. This enables the nodes to detect the failures of adjacent nodes more efficiently and quickly take over management responsibilities, achieving the technical effect of reducing node failure response time and management recovery complexity. This solves the technical problem of low node management efficiency when node resources are limited. Attached Figure Description

[0016] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0017] Figure 1 This is a schematic diagram of an application environment for an optional node management method according to an embodiment of this application;

[0018] Figure 2 This is a schematic diagram of an optional node failure occurrence process according to an embodiment of this application;

[0019] Figure 3 This is a schematic diagram of an optional node fault recovery process according to an embodiment of this application;

[0020] Figure 4 This is a flowchart illustrating an optional node management method according to an embodiment of this application;

[0021] Figure 5 This is a schematic diagram of an optional node relationship according to an embodiment of this application;

[0022] Figure 6 This is a schematic diagram of an optional node takeover process according to an embodiment of this application;

[0023] Figure 7 This is a schematic diagram illustrating an optional change in node relationships according to an embodiment of this application;

[0024] Figure 8 This is a schematic diagram of an optional node recovery process according to an embodiment of this application;

[0025] Figure 9 This is another optional node relationship change diagram according to an embodiment of this application;

[0026] Figure 10 This is a schematic diagram of an optional node management device according to an embodiment of this application;

[0027] Figure 11 This is a schematic diagram of the structure of an optional node management product according to an embodiment of this application. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] The present application will be described below with reference to embodiments:

[0031] According to one aspect of the embodiments of this application, a node management method is provided.

[0032] Optionally, in this embodiment, the above-described node management method can be applied to multiple fields such as the Internet of Things, cloud computing, and edge computing, such as centralized + distributed hybrid networks based on edge management nodes, and large-scale equipment monitoring systems. Specifically, the node can be a server, or other electronic devices such as terminal devices. The server connects to the terminal device through a network and can be used to provide services to the terminal device or applications installed on the terminal device to implement the above-described node management method. A database can also be set up on the server or independently of the server to provide data storage services for implementing the above-described node management method. Specifically:

[0033] The aforementioned servers and terminal devices can be any node in a distributed system, such as a blockchain system. This blockchain system can be formed by connecting multiple nodes through network communication. The nodes can form any type of network, and any type of computing device, such as any electronic device, can become a node in this distributed system by joining the network formed between the nodes.

[0034] For example, taking a distributed network as an example, the distributed network includes a network management center layer, an edge management node layer, and a network device layer, such as... Figure 1 As shown, the network management center layer is used to coordinate the edge management nodes in the edge management node layer. The edge management nodes are directly connected to various network devices in the network device layer, and are responsible for issuing commands and collecting data from the network management center layer, and then sending the data back to the network management center layer.

[0035] Furthermore, taking edge management node 201 in the edge management node layer as the target parent node as an example, edge management node 201 can periodically or timedly execute the above node management method, including the following steps:

[0036] S1, as Figure 2 As shown, edge management node 201 detects whether another edge management node 202 (the aforementioned target node, i.e., a child node of edge management node 201) connected to it has failed through heartbeat monitoring or other node status detection mechanisms. If a failure is detected in edge management node 202, edge management node 201 enters S2 to perform a failure response operation; if no failure is detected, it continues to wait for the next time window to execute the above node management method.

[0037] S2, as Figure 2 As shown, edge management node 201 takes over the network devices managed by edge management node 202 and edge management node 203 (the aforementioned target sub-node, i.e., the sub-node of edge management node 202). That is, after a failure occurs, edge management node 201 will assume the management responsibilities originally performed by edge management node 202. At the same time, edge management node 201 will list edge management node 203 as its own sub-node to ensure the continuity and effectiveness of the management relationship.

[0038] S3, as Figure 3 As shown, edge management node 201 sends node change messages in the edge management node layer to adjust data transmission and command reception paths, thereby maintaining stable network operation, including but not limited to:

[0039] S3-1, After the failure of edge management node 202, edge management node 201 sends a node change message to edge management node 203 to instruct edge management node 203 to update its parent node to edge management node 201;

[0040] S3-2, After the edge management node 202 resumes normal operation, the edge management node 201 will receive the fault recovery message sent by the edge management node 202. At this time, the edge management node 201 can send a node change message to the edge management node 202 and return the network device and the edge management node 203 taken over in step S2 to the edge management node 202.

[0041] S3-3, Edge Management Node 201 then sends a node change message to Edge Management Node 203 to instruct Edge Management Node 203 to update its parent node to Edge Management Node 202.

[0042] It should be noted that the node change messages in steps S3-1, S3-2, and S3-3 are different. The node change message in step S3-1 is used to instruct edge management node 203 to update its parent node from the faulty edge management node 202 to edge management node 201. The node change message in step S3-2 is used to respond to the fault recovery message issued by edge management node 202 that has resumed normal operation. The node change message in step S3-3 is used to instruct edge management node 203 to update its parent node from edge management node 201 to edge management node 202 that has resumed normal operation.

[0043] Alternatively, as an alternative implementation method, such as Figure 4 As shown, the node management method proposed in this application is executed on the target parent node in any multi-node network, specifically including but not limited to the following steps:

[0044] S402, in response to a failure of the target node, take over the network devices and target child nodes connected to the target node. The target node, target child nodes, and target parent node all belong to a node cluster. The node cluster includes N nodes. Among the N nodes, the Nth node is the parent node of the 1st node, the i-th node is the parent node of the (i+1)-th node, i is less than N, and i and N are both positive integers. The target parent node is the parent node of the target node, and the target node is the parent node of the target child node.

[0045] S404, Send a node change message to instruct the target child node to change its own parent node.

[0046] Specifically, the above-mentioned node management method refers to the process where, when a fault is detected in a node cluster, its parent node automatically takes over the management responsibility of the faulty node, as well as the subsequent node relationship change notification process.

[0047] Optionally, in the embodiments of this application, the target node, target child node, and target parent node may be, but are not limited to, edge computing nodes, IoT gateway devices, distributed network managers, or other terminal devices with network management and device monitoring capabilities.

[0048] Optionally, in this embodiment of the application, there may be multiple nodes, which are then divided into several node clusters. One node cluster includes N nodes, forming a closed-loop management chain. The target parent node is the direct superior of the target node, and the target node is the direct superior of the target child node.

[0049] For example, a node cluster is illustrated as follows: Figure 5 As shown:

[0050] The node cluster contains a cyclical hierarchical management system, in which each node is either the parent or child node of its neighboring nodes, ensuring the continuity and closed-loop characteristics of the management relationship. The node number identifies which node it is, and each node can store its own node information locally, including its parent and child nodes.

[0051] For example, each node can maintain a local information table in local storage, as shown in Table 1:

[0052] Table 1

[0053]

[0054] For example, the intra-cluster relationship of node clusters is determined in ascending order of node number. Each node has a local data table to describe the node cluster. The node with the smaller number is the parent node, and the node with the largest number is the parent node of the node with the smallest number. This intra-cluster relationship can be represented as a circular link, that is, the parent node of the Nth node is the 1st node, and the child node of the 1st node is the 2nd node, and so on until a closed loop is formed.

[0055] In addition, each node maintains a unique global information table, which contains information about all network devices managed by the cluster, the largest node number, and other data.

[0056] It should be noted that this global information table can store only the list of devices managed by the edge management node, and the global information table can be statically configured.

[0057] In other words, the management device scope and ring link connection order of each node are determined during node initialization. At this time, the relationship between nodes and the allocation of devices are clear and will not change with the dynamic addition or removal of nodes.

[0058] For example, the global information table is shown in Table 2:

[0059] Table 2

[0060]

[0061] Furthermore, upon detecting a failure in the target node, the target parent node will immediately take over the network devices and target child nodes under the responsibility of the target node, ensuring that the continuity of network management is not affected.

[0062] Furthermore, the transfer of management responsibility and changes in relationships of nodes can be applied to various network environments and scenarios, including but not limited to the Internet of Things, cloud computing, and edge computing. This application does not specifically limit these aspects.

[0063] In an exemplary embodiment, the above node management method includes, but is not limited to, a fault detection process, a takeover process, and a recovery process. The fault detection process determines whether the target node is functioning correctly; the takeover process allows the target parent node to take over the network devices and target child nodes managed by the faulty target node; and the recovery process allows the target node to re-take over the network devices and target child nodes it manages after it recovers.

[0064] In the node management process, a multi-node network can be divided into multiple node clusters, each containing several nodes. A node from one of these clusters is designated as the target parent node, and the node management method described above is executed on that target parent node, including but not limited to:

[0065] S1, the target parent node detects that its own child node (i.e., the target node) has failed;

[0066] It should be noted that each node cluster will perform independent heartbeat packet detection between each group of parent and child nodes, and then the corresponding parent node will determine whether its own child node is a faulty node.

[0067] S2, the target parent node manages the nodes and network devices, including but not limited to managing the target node's child nodes (i.e., the target child nodes) and the network devices connected to the target node.

[0068] S3, the target parent node directly sends a node change message to the target child node, so that the target child node can update its own parent node information in a timely manner and change its own parent node from the target node to the target parent node;

[0069] S4. After the target node resumes normal operation, the target parent node can send the above-mentioned node change message to the target node, so that the target node can continue to take over the network devices and target child nodes that were connected when no failure occurred. It can also send the above-mentioned node change message to the target child node, so that the target child node can update its own parent node information in a timely manner and change its own parent node from the target parent node to the target node.

[0070] Understandably, the node management process described above does not cause unnecessary message broadcasting interference to other normally operating nodes. Instead, it utilizes a disaster recovery mechanism solely through parent and child nodes. This avoids the need for reporting to a central node when a node fails, requiring the central node to adjust multiple nodes in a multi-node network to take over the tasks of the failed node. This effectively reduces node failure response time and the complexity of management and recovery. Furthermore, by flexibly switching management permissions between parent and child nodes, it prevents resource-limited nodes from becoming bottlenecks during failures, enhancing node stability and disaster recovery capabilities. This ultimately solves the technical problem of low node management efficiency when node resources are limited.

[0071] As an optional approach, after taking over the network devices and target sub-nodes connected to the target node in response to a target node failure, the method further includes: receiving a fault recovery message, wherein the fault recovery message indicates that the target node has resumed normal operation; and sending a first change message, wherein the node change message includes the first change message, and the first change message is used to update at least one of the node information of the target sub-node and the node information of the target node.

[0072] Optionally, in the embodiments of this application, the fault recovery message refers to a signal or indication that the target node has been restored to normal operation, including but not limited to the heartbeat signal returning to normal, node status report, recovery confirmation instruction, or recovery operation notification.

[0073] Optionally, in the embodiments of this application, the first change message refers to a signal or instruction used to instruct the target child node to update its current parent node information, including but not limited to node relationship change notification, new parent node number, connection parameter update or management permission adjustment instruction, etc.

[0074] Optionally, in the embodiments of this application, node information can be understood as data describing the current state of the edge management node and its management responsibilities, including but not limited to node number, parent node number, list of child node numbers, list of managed network devices, node operating status (such as normal, fault) and node connection information (such as network address, port number), etc.

[0075] For example, taking the target parent node as the executing entity, when the target parent node receives the fault recovery message sent by the target node, the target parent node executes the recovery process according to the message content, updates its own stored local information table and global information table to reflect the change in the parent-child relationship of the nodes in the node cluster. The target parent node can also send a first change message to the target child node and / or the target node to instruct the target child node and / or the target node to update the node information.

[0076] In an exemplary embodiment, taking the scenario of edge management node recovery as an example, the target node is node 3, the target parent node is node 1, and the target child nodes are nodes 4 and 5:

[0077] S1, after node 3 recovers from the failure, it sends a failure recovery message to node 1;

[0078] S2, after receiving the message, node 1 confirms that node 3 has recovered and executes the recovery process, including but not limited to sending a first change message to any of nodes 3, 4, and 5, controlling the node that receives the first change message to update its own node information, and completing the recovery process.

[0079] Through the embodiments of this application, an automatic node fault detection and recovery mechanism based on node clusters is adopted to achieve the technical effects of adaptive recovery and automatic transfer of management rights after edge management node failure, thereby achieving the goal of deploying a simple and efficient disaster recovery solution on resource-constrained edge management nodes.

[0080] As an optional approach, sending the first change message includes at least one of the following: sending the first change message to the target child node to instruct the target child node to change its parent node from the target parent node to the target node that has resumed normal operation; or sending the first change message to the target node to instruct the target node to take over the network devices that were connected before the failure occurred.

[0081] For example, after receiving the fault recovery message from the target node, the target parent node may perform the recovery process by generating a first change message and sending the first change message to the target child node to instruct the target child node to update its parent node information to the target node.

[0082] For example, after receiving the fault recovery message from the target node, the target parent node may further perform the recovery process by generating a first change message and sending the first change message to the target child node to instruct the target node to take over the management equipment it had before the fault occurred.

[0083] In an exemplary embodiment, taking the scenario of edge management node recovery and handover of management rights as an example, the target node is node 3, the target parent node is node 1, and the target child nodes are nodes 4 and 5:

[0084] S1, after the fault is recovered, node 3 sends a fault recovery message to its parent node, i.e., node 1.

[0085] S2, Node 1 receives the fault recovery message, confirms that Node 3 has resumed normal operation, generates the first change message, and sends the first change message to Node 4 and Node 5 respectively, instructing them to change the parent node information from the current Node 1 to Node 3;

[0086] S3, Node 1 can also send a first change message to Node 3, instructing Node 3 to take over its network devices and child nodes again from the point of failure;

[0087] S4, Nodes 3, 4 and 5 complete the node information update, restoring the normal node parent-child relationship and device management.

[0088] It should be noted that the execution order of S2 and S3 is not limited in the embodiments of this application.

[0089] Through the embodiments of this application, the target parent node adopts an automatic recovery process based on node clusters, which realizes the technical effect of automatic relationship adjustment and network device takeover after the subordinate child nodes resume normal operation, thereby achieving the purpose of quickly restoring network management functions and ensuring the continuous management of network devices.

[0090] As an optional approach, in response to a target node failure, the method further includes: taking over the network devices connected to the target node and the target child nodes in response to the target node failure; updating the node information of the target parent node; sending a second change message to the target child node to instruct the target child node to change its parent node from the target node to the target parent node; the updating of the node information of the target parent node includes at least one of the following: updating the network device information connected to the target parent node; updating the child node information of the target parent node.

[0091] Optionally, in this embodiment of the application, the second change message refers to a message used by the target parent node to notify the target child node to change its parent node information after the target parent node takes over the management, including but not limited to parent node number update instructions, new connection parameters, or relationship adjustment notifications.

[0092] It should be noted that when the target parent node takes over management, it can send the second change message to the target child node immediately, or it can send it after completing the initial takeover operation. The specific timing depends on the network architecture design and actual needs, and this application does not limit this.

[0093] For example, in response to a failure of the target node, the target parent node dynamically takes over the network devices and child nodes connected to it, updates the local node information to reflect the new management relationship, and then sends a second change message to the target child node, instructing the target child node to change its parent node information to the target parent node.

[0094] In one exemplary embodiment, taking a scenario of disaster recovery for multiple edge management nodes as an example:

[0095] S1, when the target parent node, such as node 1, detects that its target node, such as node 2, has failed, it actively takes over the network devices managed by node 2 and the target child nodes (child nodes of node 2), such as node 3.

[0096] S2, Node 1 updates its own node information to reflect the parent-child relationship of the node after takeover;

[0097] S3, Node 1 sends the second change message to Node 3;

[0098] S4. Based on the second change message, change the parent node from the faulty node 2 to node 1, and complete the change of the management relationship.

[0099] It should be noted that if the target child node also has a fault, the target parent node will take over the network devices connected to the target child node, in addition to taking over the network devices connected to the target child node.

[0100] Through the embodiments of this application, seamless takeover of faulty target nodes by normal target parent nodes and automatic adjustment of parent-child relationships between nodes in a node cluster are achieved, thereby improving the stability and reliability of network management architecture under resource-constrained conditions, ensuring continuous management of network devices, and improving network management efficiency.

[0101] As an optional approach, before taking over the network device connected to the target node, the method further includes at least one of the following: determining that the target node has failed if the heartbeat packet sent by the target node meets a preset fault condition; obtaining the operating status of the target node if the heartbeat packet sent by the target node meets the preset fault condition, and determining that the target node has failed in response to the operating status of the target node indicating an abnormal node process.

[0102] Optionally, in the embodiments of this application, the preset fault conditions refer to a series of rules or thresholds for determining whether a node is faulty, including but not limited to the number of times a heartbeat packet is not received consecutively, the heartbeat packet interval time exceeding a reasonable range, and error information in the node's running status report.

[0103] It should be noted that the specific conditions for determining the failure of the edge management node can be adjusted according to different network environments and device characteristics, and this application does not limit them.

[0104] For example, heartbeat detection is performed between each pair of parent and child nodes in the node cluster, such as:

[0105] A node cluster consists of a first node, a second node, and a third node. The first node is a child node of the second node, the second node is a child node of the third node, and the third node is a child node of the first node. In this case, the first node will only send heartbeat packets to its own parent node (i.e., the second node), the second node will only send heartbeat packets to its own parent node (i.e., the third node), and the third node will only send heartbeat packets to its own parent node (i.e., the first node). Then, each node will determine whether the heartbeat packets it receives from its child nodes meet the preset fault conditions.

[0106] In other words, during node status detection, heartbeat packets are only sent to nodes with a parent-child relationship, avoiding the sending of redundant heartbeat packets to other nodes without a parent-child relationship. This makes the fault detection process more accurate and efficient, reduces unnecessary network communication overhead, and improves the overall performance of the system. Thus, while ensuring rapid fault response, it optimizes the utilization of network resources.

[0107] In addition, node status checks can be performed periodically or periodically, and this application does not limit this.

[0108] Furthermore, the preset fault conditions corresponding to different nodes can be the same or different.

[0109] For example, if different nodes have the same preset fault conditions, and the first node does not receive more than 5 heartbeat packets within a cycle, then the third node is determined to have failed; if the second node also does not receive more than 5 heartbeat packets within a cycle, then the first node is determined to have failed.

[0110] For example, different nodes have different preset fault conditions. If the first node does not receive more than 5 heartbeat packets within a cycle, it is determined that the third node has failed. If the second node detects the abnormal flag of the first node's process in the heartbeat packets received within a cycle, it is determined that the first node has failed.

[0111] Furthermore, in addition to preliminary node fault detection based on preset fault conditions, more detailed comprehensive judgments can be made. For example, if the target parent node determines that the target node does not meet the preset fault conditions when it has not received multiple consecutive heartbeat packets, the target node will not be directly identified as a faulty node. Instead, the fault detection algorithm will continue to be executed to further obtain the running status of the processes in the node. This algorithm includes, but is not limited to:

[0112] Analyze the last known state of the target node, attempt to re-establish communication links with the target node, obtain the target node's system logs or monitoring data, and check the target node's availability through indirect means (such as obtaining information from other edge management nodes or network devices).

[0113] To determine the running status of processes in the node, and based on the obtained process running status, to further confirm whether the target node is truly faulty. Once confirmed, the takeover process will be formally initiated.

[0114] In one exemplary embodiment, the scenario of fault detection and disaster recovery switching of the edge management node is taken as an example:

[0115] S1, the active target parent node, such as node 1, periodically receives heartbeat packets from the target node, such as node 2;

[0116] S2, During the monitoring process, node 1 found that it failed to receive heartbeat packets from node 2 multiple times in a row, and node 2 met the preset fault conditions.

[0117] S3, Node 1 further obtains the running status of the process in Node 2 and finds that the process marking of Node 2 is abnormal;

[0118] S4. After confirming that the process flag of node 2 is abnormal, node 2 is determined to be faulty. Node 1 begins to take over the network devices connected to node 2 and updates its own node information.

[0119] Through the embodiments of this application, a fault detection mechanism based on heartbeat packet monitoring and node operation status analysis is adopted to achieve the purpose of timely detection and efficient response to node faults.

[0120] As an optional approach, the above method further includes: obtaining a set of edge management nodes that serve the central node; dividing the set of edge management nodes into several node clusters; configuring a unique node number for each node in the node cluster, wherein the node number is used to determine the node information of the corresponding node in the node cluster and the initial state of the node.

[0121] Optionally, in this embodiment of the application, the set of edge management nodes served by the central node refers to a collection of multiple sets of edge management nodes that are uniformly managed or configured by the central node, including but not limited to the running status data, network device connection information, heartbeat packet sending records, etc. of each edge management node.

[0122] Optionally, in the embodiments of this application, a node cluster refers to a group of edge management nodes divided into multiple independent management units according to a preset strategy. The edge management nodes within each unit can take over each other's management tasks to achieve disaster recovery and restoration.

[0123] It should be noted that the node clustering strategy can be flexibly adjusted according to factors such as network topology, distribution of edge management nodes, resource constraints, and load balancing requirements. For example, clustering can be based on geographical location, device type, or network traffic. This application does not limit this.

[0124] For example, first, a set of all edge management nodes is obtained. Then, based on node location, resource status, and management requirements, these edge management nodes are divided into multiple node clusters. Next, a unique node number is configured for each node in each node cluster. This number reflects the relative position of the node in its cluster, is used to determine the parent-child relationship and management order between nodes, as well as the subsequent execution process for fault detection and takeover management.

[0125] In one exemplary embodiment, a multi-layer network management architecture is used as an example:

[0126] S1, the central node obtains the set of all edge management nodes and divides them into M node clusters according to the node resource allocation;

[0127] S2, For each node cluster, the central node configures a unique node number for each edge management node in it. This number is continuous and regular within the cluster, and the parent and child nodes of the node can be inferred from the number.

[0128] S3, the central node sets the intra-cluster relationship of each node cluster according to the node number, and sets the node information of each node in the initial state, including its parent node and child node numbers.

[0129] In S2 above, for example, in a node cluster, the central node can specify that the node number starts from the minimum value and increases in fixed steps until the maximum number of nodes in the cluster is reached, forming a closed loop. This allows any node to quickly determine its adjacent parent and child node numbers and their corresponding entities through simple mathematical operations or by querying a pre-established node relationship table, thereby achieving rapid location and efficient communication in the fault detection and takeover process.

[0130] Assume a node cluster contains five edge management nodes, each with a node number, and these numbers follow a pattern within the cluster (the next or previous node number can be deduced from a given node number, or a global information table records the numbering relationship of all nodes). This application's embodiment uses an incremental numbering system:

[0131] First, the intra-cluster relationships of nodes are determined in ascending order of node numbers. Each node has a local data table describing these relationships, with the node with the smallest number being the parent node, and the node with the largest number being the parent node of the node with the smallest number. In addition, each node maintains a global information table containing information about all network devices managed by its cluster, the node with the largest number, and other data.

[0132] Through the embodiments of this application, edge management nodes are divided into multiple clusters to avoid the impact of an excessive number of nodes within a cluster on the efficiency of node fault detection and recovery. The cluster relationship determination method based on node number not only simplifies the identification and interaction between edge management nodes, but also provides convenience for the fault detection process. That is, edge management nodes can quickly identify and determine the neighboring nodes of the faulty node, thereby triggering the takeover and management process, effectively improving the speed and efficiency of disaster recovery switching.

[0133] Optionally, as an optional implementation, the node management method proposed in this application can be executed on the target node in any multi-node network, specifically including but not limited to the following steps:

[0134] S1, when resuming normal operation, determine the target parent node based on its own node number. The target parent node is used to take over the network devices and target child nodes connected to the target node when the target node fails. The target node, the target child node, and the target parent node all belong to a node cluster. The node cluster includes N nodes. Among the N nodes, the Nth node is the parent node of the 1st node, the ith node is the parent node of the (i+1)th node, i is less than N, and i and N are both positive integers. The target parent node is the parent node of the target node, and the target node is the parent node of the target child node.

[0135] S2, send a fault recovery message to the aforementioned target parent node, wherein the aforementioned fault recovery message indicates that the aforementioned target node has resumed normal operation;

[0136] S3, send a node change message to the target child node to instruct the target child node to change its own parent node.

[0137] As an optional approach, determining the target parent node based on the node number of the target node includes: increasing the value of the node number of the target node by a preset step size to obtain the target value; and determining the node whose node number is the target value as the target parent node.

[0138] Taking any node in a node cluster as the target node, the above node management method is executed on the target node, including but not limited to:

[0139] When a target node fails, the target parent node will manage it, including but not limited to managing the target node's child nodes (i.e., target child nodes) and the network devices connected to the target node, and directly sending node change messages to the target child nodes so that the target child nodes can update their parent node information in a timely manner.

[0140] Furthermore, after the target node recovers and resumes normal operation, since the target node is uncertain whether its parent node is still active (i.e., whether it can operate normally), it will determine the parent node based on its own node number. If the parent node is also faulty, it will recursively search for its parent node within the cluster. After finding its parent node, it will notify the parent node of its recovery information.

[0141] In other words, if the target node finds that it cannot communicate normally with its parent node before the failure, the target node does not need to update its parent node information and will still use that parent node as the target parent node. Conversely, if the target node finds that it cannot communicate normally with its parent node before the failure and determines that the parent node has also failed, the target node will continue to determine whether the parent node of the failed parent node has failed. If it has not failed, the target node will determine that parent node as its own parent node (the target parent node).

[0142] The target node then sends a fault recovery message to the target parent node. Subsequently, if the target parent node receives the fault recovery message from the target node, it will send a node change message to the target node.

[0143] Next, after receiving the node change message from the target node, the target node will continue to take over the network devices and target child nodes that were connected before the failure occurred.

[0144] Optionally, as an optional implementation, the node management method proposed in this application can be executed on the target child node in any multi-node network, specifically including but not limited to the following steps:

[0145] S1, in response to a target node failure, receives a node change message and changes its own parent node from the target node to the target parent node. The target parent node is used to take over the network devices connected to the target node and the target child nodes when the target node fails. The target node, the target child nodes, and the target parent node all belong to a node cluster. The node cluster includes N nodes. Among the N nodes, the Nth node is the parent node of the 1st node, the i-th node is the parent node of the (i+1)-th node, i is less than N, and i and N are both positive integers. The target parent node is the parent node of the target node, and the target node is the parent node of the target child node.

[0146] S2, in response to the above target node resuming normal operation, receives the above node change message and changes its own parent node from the above target parent node to the above target node.

[0147] Using any node as the target child node, execute the node management method described above on that target child node, including but not limited to:

[0148] In an exemplary embodiment, taking an edge management node cluster under a three-layer network management architecture as an example, assume that the cluster consists of 5 nodes, numbered sequentially from node 1 to node 5, where node 1 is the parent node of node 2, and so on to form a closed loop. The node management method proposed in this application is executed on the target child node in any multi-node network.

[0149] S1, in response to a failure in target node 2, node 3 receives a node change message from node 1 and changes its parent node from node 2 to node 1. At this time, node 1, as the parent node of node 2, takes over the network devices connected to node 2 and node 3. Node 3 updates its local information table, changing the parent node number to node 1 to reflect the new parent-child relationship.

[0150] S2. Subsequently, when node 2 resumes normal operation, node 3 receives the node change message from node 1 again and changes its parent node back to node 2. Node 1 transfers the information of the network devices and child nodes it has taken over to manage to node 2, while node 2 updates its own information table, including restoring the list of network devices it manages and controls and the connections of child nodes.

[0151] According to another aspect of the embodiments of this application, the embodiments of this application also provide a node management system, including:

[0152] The target parent node, in response to a failure of the target node, takes over the network devices and target child nodes connected to the target node. The target node, target child nodes, and target parent node all belong to a node cluster containing N nodes. The Nth node is the parent of the first node, the i-th node is the parent of the (i+1)-th node, i is less than N, and both i and N are positive integers. The target parent node is the parent of the target node, and the target node is the parent of the target child node. A node change message is sent to instruct the target child node to change its own parent node.

[0153] When a target node resumes normal operation, it determines its target parent node based on its own node number; it sends a fault recovery message to the target parent node, indicating that the target node has resumed normal operation; and it sends a node change message to the target child nodes to instruct the child nodes to change their own parent node.

[0154] When the target node fails, the target child node receives a node change message and changes its own parent node from the target node to the target parent node; when the target node resumes normal operation, it receives a node change message and changes its own parent node from the target parent node to the target node.

[0155] Let's take an example where the target node, target parent node, and target child node are all edge management nodes:

[0156] S1, Fault Detection Process: When all edge management nodes are running normally, each node will determine whether its child nodes are alive by sending and receiving heartbeat packets. If it does not receive heartbeat packets from its own child nodes (the target nodes mentioned above) for several consecutive cycles, it can use other means (including but not limited to determining whether the child node process is running normally) to make a comprehensive judgment. If it is determined that the target node is in an abnormal state and cannot continue to manage the connected network devices, the target parent node will enter the takeover management process.

[0157] S2, Takeover Management Process: When the target parent node determines that its target node has failed, it actively takes over the management of the devices managed by the target node. Simultaneously, it sends a node change message to the target child nodes to instruct them to modify their local information tables. This indicates that a takeover management process has occurred, and the parent node needs to be changed. Taking a failure in node 2 as an example, the specific process is as follows: Figure 6 As shown, this includes, but is not limited to, node 1 representing the target parent node, node 2 representing the target node, and node 3 representing the target child node:

[0158] First, Node 1 determines that Node 2 has failed and, based on the global information table, attempts to manage the network devices managed by Node 2.

[0159] Next, if node 1 successfully manages the network devices managed by node 2, it notifies node 3 that its parent node has changed to node 1. The relationships between nodes in the initial node cluster are then changed, and the management takeover is successful. The diagram illustrating the changes in node relationships at this point is shown below. Figure 7 As shown.

[0160] Similarly, if multiple nodes fail, the above procedure can be followed to complete the replacement of the pipeline.

[0161] S3, Recovery Process: When the target node recovers, it first attempts to connect to its parent node. If the parent node also fails, it recursively searches for the target parent node within the cluster. Once the target parent node is found, it notifies the target parent node of its recovery information. The target parent node will return the node that was originally managed by the target child node. It can also send a node change message to the target child node to instruct the target child node to change its current parent node from the target parent node to the target node that has recovered and is running normally, thus completing the recovery process.

[0162] It should be noted that the parent node can periodically or periodically perform the S1 fault detection process on its own child nodes. Therefore, multiple node failures may occur at a certain time. For example, Figure 7 Taking the node relationship diagram as an example, Node 2 discovers that its child node, Node 3, has failed, and Node 2 then takes over the network devices that were originally connected to Node 3. At this time, the parent and child nodes of Nodes 2 and 3 remain unchanged, only the information of the connected network devices changes. Next, Node 1 discovers that its child node, Node 2, has failed, and Node 1 then takes over the network devices that were originally connected to Nodes 2 and Node 3, and identifies Node 3, which was originally a child node of Node 2, as its own child node. At this time, since Nodes 2 and 3 are already in a failed state, they cannot update their own parent, child, and network device information. Node 1 will update its own node information, which can be recorded in the local information table, including but not limited to Node 1 adding Node 3 to the child node list and taking over the network devices that were originally connected to Nodes 2 and Node 3.

[0163] It should also be noted that the above global information table, as the node information initialized for each node, will not update its specific content in real time as the node status changes. For example, if a node changes from normal operation to a fault state and then back to normal operation, the static information in the global information table, such as the node number and device list, will remain unchanged. This ensures the stability of node operation and facilitates each node to quickly locate and adjust its management responsibilities based on the preset global information during fault detection and recovery.

[0164] For example, taking the recovery of node 3 (the target node) after the failure of nodes 2, 3, and 4 as an example, the recovery process is as follows: Figure 8 As shown, including but not limited to:

[0165] After node 3 recovers, it recursively searches for its parent node and finds that its latest parent node is node 1. Node 3 notifies node 1 that it has recovered and requests to be reinstated. Node 1 removes node 4 and node 5 from its list of child nodes based on its list of managed nodes. Then, either node 1 or the recovered node 3 notifies node 4 and node 5 to change their parent node to node 3.

[0166] Furthermore, since node 4 is also in a faulty state, node 3 will not only manage the network devices connected to node 3, but will also directly manage the network devices connected to node 4 according to the S2 takeover management process. After node 3 is successfully managed, the recovery process ends normally. The changes in the relationships within the edge management node cluster before and after the recovery process are illustrated in the diagram below. Figure 9 As shown.

[0167] Specifically, such as Figure 9As shown, before node 3 recovered from the fault, only nodes 1 and 5 were available in the node cluster, so they formed a new node cluster with the following parent-child relationship: node 1 is the parent of node 5 and node 5 is the parent of node 1. After node 3 recovered from the fault, nodes 1, 3, and 5 were available in the node cluster, so these three nodes formed a new node cluster with the following parent-child relationship: node 1 is the parent of node 3 and node 3 is the parent of node 5.

[0168] In summary, the node management system in this application embodiment can be applied to a multi-node network. If the target parent node detects that the target node has failed, it will directly send a second change message to the target node's child node, i.e., the target child node, to control the target child node to change its own parent node to the target node.

[0169] Next, after the target node resumes normal operation, the target parent node or the target node will send a first change message to the target child node, instructing the target child node to restore the parent node information to the target node, and for the target node to regain control of the network devices and child nodes it originally managed. This process does not require the intervention of a central node, achieving automation and efficiency in fault takeover and recovery.

[0170] This effectively avoids the delays, resource consumption, and centralization risks associated with coordinating fault takeover operations of other nodes through a central node in existing technologies. It makes the fault detection, takeover, and recovery process faster, lighter, and more decentralized, thereby enhancing the stability and reliability of the network management architecture.

[0171] Considering that existing disaster recovery solutions are more focused on data disaster recovery and cannot ensure the continuous management of network devices, this application simplifies the disaster recovery mechanism through direct communication and cooperation between edge nodes, improves the continuity and efficiency of network device management, and effectively enhances the robustness and response speed of overall network management in scenarios where edge node resources are limited and widely distributed. It can ensure the continuity of network device management and improve the reliability of network management. Moreover, the node management method provided in this application relies on fewer resources, is easier to implement, and does not require a single node to maintain multiple master-slave backups, making it suitable for edge management node scenarios with insufficient resources.

[0172] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0173] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0174] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0175] Based on this understanding, the technical solution of this application, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / random access memory (RAM), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0176] According to another aspect of the embodiments of this application, a node management apparatus for implementing the above-described node management method is also provided. This node management apparatus can be used to implement the node management method provided in the above embodiments, and details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0177] Taking the execution of the above node management method on the target parent node as an example, the above node management device can be deployed on the target parent node. Figure 10 This is a structural block diagram of an optional node management device according to an embodiment of this application, such as... Figure 10 As shown, the node management device includes:

[0178] The takeover module 1002 is used to take over the network devices and target child nodes connected to the target node in response to a failure of the target node. The target node, target child nodes, and target parent node all belong to a node cluster. The node cluster includes N nodes. Among the N nodes, the Nth node is the parent node of the 1st node, the i-th node is the parent node of the (i+1)-th node, i is less than N, and i and N are both positive integers. The target parent node is the parent node of the target node, and the target node is the parent node of the target child node.

[0179] The transmission module 1004 is used to send node change messages to instruct the target child node to change its own parent node.

[0180] As an optional solution, the apparatus is further configured to: in response to a failure of the target node, after taking over the network devices and target sub-nodes connected to the target node, receive a fault recovery message, wherein the fault recovery message indicates that the target node has resumed normal operation; and send a first change message, wherein the node change message includes a first change message, the first change message being used to update at least one of the node information of the target sub-node and the node information of the target node.

[0181] As an optional approach, the device is used to send a first change message in at least one of the following ways: sending a first change message to a target child node to instruct the target child node to change its own parent node from the target parent node to the target node that has resumed normal operation; or sending a first change message to a target node to instruct the target node to take over the network devices that were connected before the failure occurred.

[0182] As an optional solution, the device is also used to: take over the network devices connected to the target node and the target child node in response to a failure of the target node before receiving the fault recovery message; update the node information of the target parent node; and send a second change message to the target child node to instruct the target child node to change its own parent node from the target node to the target parent node.

[0183] As an optional approach, the device is used to update the node information of the target parent node in at least one of the following ways: updating the network device information connected to the target parent node; updating the child node information of the target parent node.

[0184] As an optional approach, before taking over the network devices connected to the target node, the device is further configured to: determine that the target node has failed if the heartbeat packet sent by the target node meets the preset fault conditions; obtain the operating status of the target node if the heartbeat packet sent by the target node meets the preset fault conditions, and determine that the target node has failed in response to the operating status of the target node indicating that the node process is abnormal.

[0185] As an optional solution, the device is also used to: acquire a set of edge management nodes serving the central node; divide the set of edge management nodes into several node clusters; and configure a unique node number for each node in the node cluster, wherein the node number is used to determine the node information of the corresponding node in the node cluster and the initial state of the node.

[0186] Taking the execution of the above node management method on the target node as an example, the above node management device can be deployed on the target node, and the node management device includes:

[0187] The determination module is used to determine the target parent node based on its own node number when normal operation is restored. The target parent node is used to take over the network devices and target child nodes connected to the target node when the target node fails. The target node, target child nodes and target parent node all belong to a node cluster. The node cluster includes N nodes. Among the N nodes, the Nth node is the parent node of the 1st node, the i-th node is the parent node of the (i+1)-th node, i is less than N, and i and N are both positive integers. The target parent node is the parent node of the target node, and the target node is the parent node of the target child node.

[0188] The transmission module is used to send a fault recovery message to the target parent node, whereby the fault recovery message indicates that the target node has resumed normal operation; and to send a node change message to the target child node, instructing the target child node to change its own parent node.

[0189] As an optional approach, the device determines the target parent node based on the node number of the target node in the following manner: the value of the node number of the target node is increased by a preset step size to obtain the target value; the node whose node number is the target value is determined as the target parent node.

[0190] Taking the execution of the above node management method on a target child node as an example, the above node management device can be deployed on the target child node, and the node management device includes:

[0191] The change module is used to respond to a target node failure by receiving a node change message and changing its own parent node from the target node to the target parent node. The target parent node is used to take over the network devices and target child nodes connected to the target node when the target node fails. The target node, target child nodes, and target parent node all belong to a node cluster. The node cluster includes N nodes. The Nth node is the parent node of the 1st node, the ith node is the parent node of the (i+1)th node, i is less than N, and i and N are both positive integers. The target parent node is the parent node of the target node, and the target node is the parent node of the target child node.

[0192] The transmission module is used to respond to the target node resuming normal operation by receiving node change messages and changing its own parent node from the target parent node to the target node.

[0193] Regarding the apparatus in the above embodiments, the terms "module" or "unit" refer to a computer program or part of a computer program with a predetermined function, which works together with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit. The specific manner in which each module performs its operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.

[0194] According to another aspect of the embodiments of this application, an electronic device is provided.

[0195] The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor is configured to perform the steps in any of the above method embodiments via the computer program. In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor. Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.

[0196] According to one aspect of this application, a computer program product is also provided, which includes a computer program.

[0197] The computer program product includes a computer program / instructions containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via communication section 1109, and / or installed from removable media 1111. When the computer program is executed by central processing unit 1101, it performs various functions provided in the embodiments of this application. The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0198] Figure 11 A schematic block diagram of a computer system architecture for implementing embodiments of the present application is shown. Figure 11As shown, the computer system 1100 includes a Central Processing Unit (CPU) 1101, which can perform various appropriate actions and processes based on programs stored in ROM 1102 or programs loaded into RAM 1103 from storage section 1108. Random access memory 1103 also stores various programs and data required for system operation. The CPU 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. Input / output (I / O) interface 1105 is also connected to bus 1104.

[0199] The following components are connected to I / O interface 1105: an input section 1106 including a keyboard, mouse, etc.; an output section 1107 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1108 including a hard disk, etc.; and a communication section 1109 including a network interface card such as a local area network card, modem, etc. The communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to I / O interface 1105 as needed. Removable media 1111, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1110 as needed so that computer programs read from them can be installed into storage section 1108 as needed.

[0200] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0201] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer programs / instructions. For example, embodiments of this application include a computer program / instruction comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs various functions defined in the system of this application. In such embodiments, the computer program / instruction can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program / instruction is executed by a central processing unit, it performs the aforementioned node management method.

[0202] According to one aspect of this application, a computer-readable storage medium is also provided.

[0203] The processor of the aforementioned electronic device can read the computer instructions from a computer-readable storage medium, and the processor executes the computer instructions to cause the electronic device to perform the node management methods provided in the various optional implementations of node management in the aforementioned node cluster.

[0204] Optionally, in this embodiment, the computer-readable storage medium described above may be configured to store methods for performing the embodiments of this application.

[0205] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0206] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0207] In the several embodiments provided in this application, it should be understood that the disclosed application can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.

[0208] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0209] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0210] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A node management method characterized by comprising: The method is applied to a target parent node, and comprises: In response to a target node being faulty, taking over network devices connected by the target node and a target child node, wherein the target node, the target child node and the target parent node belong to a node cluster, the node cluster comprises N nodes, an Nth node in the N nodes is a parent node of a 1st node, an ith node is a parent node of an (i+1)th node, i is less than N, i and N are positive integers, the target parent node is a parent node of the target node, and the target node is a parent node of the target child node; Sending a node change message to indicate that the target child node changes a parent node thereof.

2. The method of claim 1, wherein, After the taking over of the network devices connected by the target node and the target child node, the method further comprises: Receiving a fault recovery message, wherein the fault recovery message indicates that the target node is in normal operation; Sending a first change message, wherein the node change message comprises the first change message, and the first change message is used to update at least one of node information of the target child node and node information of the target node.

3. The method of claim 2, wherein, The sending of the first change message comprises at least one of: Sending the first change message to the target child node to indicate that the target child node changes a parent node thereof from the target parent node to the target node in normal operation; Sending the first change message to the target node to indicate that the target node takes over network devices connected before the target node is faulty.

4. The method of claim 1, wherein, After the response to the target node being faulty, the method further comprises: Taking over the network devices connected by the target node and the target child node; Updating node information of the target parent node; Sending a second change message to the target child node to indicate that the target child node changes a parent node thereof from the target node to the target parent node.

5. The method of claim 4, wherein, The updating of the node information of the target parent node comprises at least one of: Updating network device information connected by the target parent node; Updating child node information of the target parent node.

6. The method of claim 1, wherein, Before the taking over of the network devices connected by the target node, the method further comprises at least one of: In a case where a heartbeat packet sent by the target node satisfies a preset fault condition, determining that the target node is faulty; In a case where a heartbeat packet sent by the target node satisfies a preset fault condition, acquiring an operation state of the target node, and in response to the operation state of the target node indicating that a node process is abnormal, determining that the target node is faulty.

7. The method of claim 1, wherein, The method further comprises: Acquiring a group of edge management nodes serving a center node; Dividing the group of edge management nodes into a plurality of node clusters; Configuring a unique node number for each node in the node cluster, wherein the node number is used to determine position of a corresponding node in the node cluster and node information of an initial state.

8. A node management method characterized by, The method is applied to a target node, and comprises: When the target node resumes normal operation, a target parent node is determined according to a node number of the target node, wherein the target parent node is used to take over network devices connected by the target node and a target child node when the target node fails, the target node, the target child node and the target parent node all belong to a node cluster, the node cluster includes N nodes, an Nth node in the N nodes is a parent node of a 1st node, an ith node is a parent node of an (i+1)th node, i is less than N, i and N are positive integers, the target parent node is a parent node of the target node, and the target node is a parent node of the target child node; a failure recovery message is sent to the target parent node, wherein the failure recovery message indicates that the target node resumes normal operation; a node change message is sent to the target child node to instruct the target child node to change a parent node thereof.

9. The method of claim 8, wherein, The target parent node is determined according to the node number of the target node, and includes: a preset step is added to the node number of the target node to obtain a target value; a node with the target value of the node number is determined as the target parent node.

10. A node management method characterized by comprising: Applied to the target child node, and including: in response to a failure of the target node, a node change message is received, and a parent node thereof is changed from the target node to a target parent node, wherein the target parent node is used to take over network devices connected by the target node and the target child node when the target node fails, the target node, the target child node and the target parent node all belong to a node cluster, the node cluster includes N nodes, an Nth node in the N nodes is a parent node of a 1st node, an ith node is a parent node of an (i+1)th node, i is less than N, i and N are positive integers, the target parent node is a parent node of the target node, and the target node is a parent node of the target child node; in response to the target node resuming normal operation, the node change message is received, and the parent node thereof is changed from the target parent node to the target node.

11. A node management system, characterized by including: a target parent node, in response to a failure of a target node, takes over network devices connected by the target node and a target child node, wherein the target node, the target child node and the target parent node all belong to a node cluster, the node cluster includes N nodes, an Nth node in the N nodes is a parent node of a 1st node, an ith node is a parent node of an (i+1)th node, i is less than N, i and N are positive integers, the target parent node is a parent node of the target node, and the target node is a parent node of the target child node; a node change message is sent to instruct the target child node to change a parent node thereof; the target node, when resuming normal operation, determines the target parent node according to a node number of the target node; a failure recovery message is sent to the target parent node, wherein the failure recovery message indicates that the target node resumes normal operation; the node change message is sent to the target child node to instruct the target child node to change a parent node thereof. The target sub-node, in response to the target node being out of service, receives the node change message, changes its parent node from the target node to a target parent node; and in response to the target node being back to normal operation, receives the node change message, changes its parent node from the target parent node to the target node.

12. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instructions, when executed by the processor, implement the steps of the method of any one of claims 1-7, 8-9, 10.

13. A computer-readable storage medium, characterized in that, The computer program / instructions, when executed by the processor, implement the steps of the method of any one of claims 1-7, 8-9, 10.

14. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The computer program / instructions, when executed by the processor, implement the steps of the method of any one of claims 1-7, 8-9, 10.