Network fault recovery method, device, and computer readable medium

WO2025118846A9PCT designated stage expired Publication Date: 2025-07-17ZTE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/126077
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-06
Filing Date
2024-10-21
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

The existing wireless Mesh network cannot be automatically restored when a failure occurs. Users need to manually restart the device to solve the problem. The main node device is selected to back up only based on the received signal strength and does not fully consider the relevant information of the child node device, which may lead to a degradation of network performance.

Method used

The primary node device receives the attribute information sent by the child node device, including capability information, load information and connection information, determines the backup master node device, and sends a notification message to it. The backup master node device switches to the master node device when the master node fails, re-organizes the network to realize automatic recovery of network failures.

Benefits of technology

The automatic recovery of wireless Mesh network when a failure occurs is realized. The selection of backup master node equipment takes into account the capabilities and network load of child node equipment, and improves network performance and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024126077_17072025_PF_FP_ABST
    Figure CN2024126077_17072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a network fault recovery method. The method comprises: receiving child node attribute information sent by child node devices connected to a master node device, wherein the child node attribute information comprises at least one of the following: capability information of the child node devices, load information, and connection information between the child node devices and a terminal device; and determining a backup master node device on the basis of the child node attribute information, and sending a backup master node notification message to the backup master node device, wherein the backup master node device is a child node device of a first network and is used for replacing the master node device when the first network fails and is reestablished. The present disclosure further provides a master node device, a child node device, and a computer readable medium.
Need to check novelty before this filing date? Find Prior Art

Description

Network fault recovery method, device and computer readable medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application No. 202311666211.4 filed with the China Patent Office on December 6, 2023, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to, but is not limited to, the field of computer technology. Background Art

[0004] Wireless Mesh network is a new type of wireless broadband access network. With its multi-hop interconnection and mesh topology characteristics, wireless Mesh network is suitable for various wireless access networks such as broadband home networks, community networks, enterprise networks and metropolitan area networks. In related technologies, Mesh network includes a main node device and multiple sub-node devices. All sub-node devices synchronize the status of their downstream devices to the main node device at regular intervals. The main node device sets parameters for the sub-node devices according to demand and controls the sub-node devices. Usually, in the process of wireless Mesh networking and when problems are encountered in the use of Mesh network, users are required to manually reset the main node device and sub-node devices. Since most users are not professional technicians, wireless Mesh networking causes them great difficulties.

[0005] Currently, there is technology that can complete mode switching and automatically join the Mesh network without manually setting up the master and sub-nodes during the networking process. This optimization makes it convenient for users to establish a Mesh network, but it still cannot solve the problem of Mesh network failure in use. Users can only rely on manually restarting the device to solve it.

[0006] Summary of the Invention

[0007] The present disclosure provides a network failure recovery method, device, and computer-readable medium.

[0008] In a first aspect, an embodiment of the present disclosure provides a network fault recovery method, which is applied to a master node device in a first network, the method comprising: receiving sub-node attribute information sent by each sub-node device connected to the master node device, the sub-node attribute information comprising at least one of the following: capability information, load information, and connection information between the sub-node device and the terminal device; determining a backup master node device based on the sub-node attribute information, and sending a backup master node notification message to the backup master node device; the backup master node device is a sub-node device of the first network, and is used to replace the master node device when the first network fails and is re-established.

[0009] On the other hand, an embodiment of the present disclosure provides a network fault recovery method, which is applied to a sub-node device in a first network, the method comprising: sending first sub-node attribute information of the sub-node device to a main node device to which the sub-node device is connected; upon receiving second sub-node attribute information sent by a downstream sub-node device of the sub-node device, sending the second sub-node attribute information to the main node device; the first sub-node attribute information and the second sub-node attribute information include at least one of the following: capability information, load information, and connection information between the sub-node device and the terminal device, and the first sub-node attribute information and the second sub-node attribute information are used to determine a backup main node device.

[0010] On the other hand, an embodiment of the present disclosure also provides a master node device, comprising: one or more processors; a storage device on which one or more programs are stored; when the one or more programs are executed by the one or more processors, the one or more processors implement the network fault recovery method as described above; one or more I / O interfaces, connected between the processor and the storage device, configured to implement information interaction between the processor and the storage device.

[0011] On the other hand, an embodiment of the present disclosure also provides a sub-node device, including: one or more processors; a storage device on which one or more programs are stored; when the one or more programs are executed by the one or more processors, the one or more processors implement the network fault recovery method as described above; one or more I / O interfaces, connected between the one or more processors and the storage device, configured to implement information interaction between the one or more processors and the storage device.

[0012] On the other hand, an embodiment of the present disclosure further provides a computer-readable medium having a computer program stored thereon, wherein the computer program implements the network failure recovery method as described above when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] FIG1 is a schematic diagram of a network failure recovery process with a master node device as the execution subject provided by an embodiment of the present disclosure;

[0014] FIG2 is a schematic diagram of a network failure recovery process with a master node device as the execution subject provided by an embodiment of the present disclosure;

[0015] FIG3 is a schematic diagram of a network failure recovery process with a master node device as the execution subject provided by an embodiment of the present disclosure;

[0016] FIG4 is a schematic diagram of various mode switching of a sub-node device provided by an embodiment of the present disclosure;

[0017] FIG5 is a schematic diagram of a network failure recovery process with a sub-node device as the execution subject provided by an embodiment of the present disclosure;

[0018] FIG6 is a schematic diagram of a network failure recovery process with a sub-node device as the execution subject provided by an embodiment of the present disclosure;

[0019] FIG7 is a schematic diagram of a network failure recovery process with a sub-node device as the execution subject provided by an embodiment of the present disclosure;

[0020] FIG8 is a schematic diagram of a network failure recovery process with a sub-node device as the execution subject provided by an embodiment of the present disclosure;

[0021] FIG9 is a schematic diagram of a network failure recovery process with a subnode device as the execution subject provided by an embodiment of the present disclosure;

[0022] FIG10a is a schematic diagram of a Mesh network topology for initial networking provided by an embodiment of the present disclosure;

[0023] FIG10b is a schematic diagram of the Mesh network topology after selecting a backup master node device according to an embodiment of the present disclosure;

[0024] FIG10c is a schematic diagram of the topology after rapid network reorganization according to an embodiment of the present disclosure;

[0025] FIG11a is a schematic diagram of a module of a master node device provided by an embodiment of the present disclosure;

[0026] FIG11b is a schematic diagram of a module of a sub-node device provided in an embodiment of the present disclosure;

[0027] FIG12 is a schematic diagram of the structure of a master node device and a sub-node device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] Example embodiments will be described more fully hereinafter with reference to the accompanying drawings, but the example embodiments may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the scope of this disclosure to those skilled in the art.

[0029] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0030] The terms used herein are used only to describe specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It will also be understood that when the terms "comprising" and / or "made of" are used in this specification, the presence of the features, wholes, steps, operations, elements, and / or components is specified, but the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups thereof is not excluded.

[0031] The embodiments described herein may be described with reference to plan views and / or cross-sectional views, with the aid of idealized schematic diagrams of the present disclosure. Thus, the example illustrations may be modified based on manufacturing techniques and / or tolerances. Therefore, the embodiments are not limited to the embodiments shown in the accompanying drawings, but include modifications of the configurations formed based on the manufacturing process. Therefore, the regions illustrated in the accompanying drawings are schematic in nature, and the shapes of the regions shown in the drawings illustrate specific shapes of the regions of the elements, but are not intended to be limiting.

[0032] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly defined as such herein.

[0033] A Mesh network is composed of a master node device and multiple sub-node devices. The uplink connection of the master node device can provide WAN (Wide Area Network) access to itself and the sub-node devices. If the WAN connection of the master node device or the Mesh network connection cannot be restored due to special software or hardware reasons, the entire Mesh network will be paralyzed. In the related art, the above problem is solved by selecting a backup master node device to switch the master node device. The basis for selecting the backup master node device is a preset value, and the sub-node device with the largest received signal strength or greater than the preset value among all the sub-node devices is selected as the backup master node device. However, selecting the backup master node device based solely on the received signal strength does not fully consider the relevant information of the candidate sub-node device itself. It is possible that the candidate master node device selected based on the received signal strength will lower the performance of the entire Mesh network.

[0034] To at least address the aforementioned issues, embodiments of the present disclosure provide a network fault recovery method, which is applied to a master node device in a first network. In embodiments of the present disclosure, the first network is a mesh network. As shown in FIG1 , the method may include the following steps S11 and S12.

[0035] In step S11, subnode attribute information sent by each subnode device connected to the master node device is received. The subnode attribute information includes at least one of the following: capability information of the subnode device, load information, and connection information between the subnode device and the terminal device.

[0036] In a mesh network, the master node device connects to each first-level sub-node device, which in turn connects to its downstream sub-node device (i.e., the second-level sub-node device). Each sub-node device reports its sub-node attribute information to its upstream node device. If its upstream node device is another sub-node device, the other sub-node device further reports the information to its upstream sub-node device until the information reaches the master node device.

[0037] Each sub-node device can carry sub-node attribute information in the extended information of the 1905 protocol. The sub-node attribute information is the information of the sub-node device itself, and can include one of the following or any combination: capability information of the sub-node device, load information, and connection information between the sub-node device and the terminal device. Capability information can include the theoretical supported bandwidth of the sub-node device and the supported network standard information, for example, whether it supports 5G network, etc.; load information can include the number of data packets sent and received and the transmission rate of the terminal device connected to the sub-node device, such as TX and RX data transmission and reception information; the connection information between the sub-node device and the terminal device can include the number of terminal devices connected to the sub-node device, the connection duration, and the traffic used by the terminal device on the sub-node device. The connection information between the sub-node device and the terminal device can be a connection list. The connection list can reflect the user's usage preferences. For example, the terminal device accesses the router in the living room the most times and has the longest connection time, indicating that the user's frequent activities are in the living room.

[0038] In order to select a backup master node device, each sub-node device and the master node device will periodically exchange information, and the information includes at least one of the following: MAC address, RSSI (Received Signal Strength Indication) information, connection information between the sub-node device and the terminal device, WAN / LAN capability information, and actual data transmission and reception information of each sub-node device. In the embodiment of the present disclosure, the information may also include at least one of the following: WAN side bandwidth, theoretical and actual rates; LAN side (including wired and wireless) theoretical bandwidth, actual rate, and WiFi capability information. The information can be sent via a socket or carried in an extended field of the 1905 protocol.

[0039] In step S12, a backup master node device is determined based on the sub-node attribute information, and a backup master node notification message is sent to the backup master node device; the backup master node device is a sub-node device of the first network, and is used to replace the master node device when the first network fails and is re-established.

[0040] The master node device can calculate a current optimal sub-node device based on the sub-node attribute information reported by each sub-node and in combination with the Mesh topology diagram, and use the current optimal sub-node device as the backup master node device. In some embodiments, the backup master node device can be one of the sub-node devices. Once the Mesh network failure cannot be recovered, the master node device exits the Mesh network, and the backup master node device switches to the master node device for rapid re-networking. In some embodiments, multiple backup master node devices can also be determined, each backup master node device is a sub-node device, and a backup master node device queue is obtained by sorting by priority. The sub-node device with the best sub-node attribute information is placed at the front of the queue. In the case that the Mesh network failure cannot be recovered, the master node device will give priority to switching to the sub-node device with the highest priority. If the sub-node device is unable to perform the function, the next sub-node device in the queue will be selected as the master node device.

[0041] In some embodiments, the backup master node device can be determined based on the subnode attribute information and the priority or weight of the preset subnode attribute information. The priority or weight can be set based on user needs. In this way, the determined backup master node device better meets user needs and further improves the user service experience.

[0042] After determining the backup master node device, the master node device can send a backup master node notification message to the backup master node device to notify the backup master node device that it has been selected as the backup master node. After receiving the backup master node notification message, the backup master node device can broadcast in the Mesh network to announce that it has been selected as the backup master node device.

[0043] The network fault recovery method provided by the embodiment of the present disclosure comprises the following steps: a master node device receives sub-node attribute information sent by each sub-node device directly connected to it, determines a backup master node based on the sub-node attribute information and notifies the backup master node, so that the backup master node can replace the master node device when the network failure is re-established; wherein the sub-node attribute information includes at least one of the following: capability information, load information, and connection information between the sub-node device and the terminal device; the embodiment of the present disclosure determines the backup master node when the network is normal, so that when the network fails, the backup master node device can automatically switch to the master node device for re-establishing the network, thereby realizing automatic recovery of the network failure; the selection of the backup master node device will take into account at least one of the capability of the sub-node device itself, the load of the entire network, and the habitual preference of the terminal device to connect to the sub-node device, so that the network performance after recovery is higher, and can provide users with a better data service experience.

[0044] In some embodiments, as shown in Figure 2, after determining the backup master node device based on the sub-node attribute information and sending a backup master node notification message to the backup master node device (i.e., step S12), the method further includes the following steps: Step S13, when the backup master node device is different from the backup master node device determined last time, sending a sub-node recovery message to the backup master node device determined last time, the sub-node recovery message is used to indicate that the backup master node device determined last time is restored to a sub-node device; wherein, the backup master node device determined last time is determined by the master node device based on the sub-node attribute information received last time.

[0045] Each child node device can transmit its own child node attribute information according to a preset first period, or transmit its own child node attribute information when the child node device's child node attribute information changes. For each instance of child node attribute information reported by each child node device, the master node device will determine the current backup master node device. If the currently determined backup master node device is different from the previously determined backup master node device, the previously determined backup master node device will be restored as the child node device.

[0046] In addition to accessing the first network, the master node device also has upstream access to the second network. In some embodiments, the first network is a Mesh network and the second network is the Internet. In the Mesh network, the master node device and the sub-node device are connected through a LAN (Local Area Network) or Wifi, and the master node device accesses the Internet through a WAN.

[0047] Because a Mesh network involves multiple sub-node devices, users may not be able to determine which sub-node device has a problem, resulting in a poor user experience. To address this issue, the disclosed embodiments add a first network connection detection function to the master node device. Therefore, in some embodiments, the method may further include the following steps: detecting the connection status of the first network and / or the connection status of the second network.

[0048] The network failure recovery process is described in detail below with reference to FIG. 3 . As shown in FIG. 3 , the network failure recovery process includes the following steps S21 to S28 .

[0049] In step S21, the connection status of the first network and the connection status of the second network are detected. If the first network connection is normal and the second network connection is abnormal, step S22 is executed; if the first network connection is abnormal, step S26 is executed.

[0050] If the master node device does not receive the Mesh message sent by its downstream child node device within the timeout period A, it will be marked and stored. If the master node device does not receive the Mesh message sent by all downstream child nodes within the timeout period B, it is considered that there is a problem with the Mesh network connection between the master node device and each child node device, that is, the first network connection is abnormal. In this case, execute step S26.

[0051] If the master node device receives Mesh messages sent by all its downstream child node devices within the timeout period A, it is considered that the first network connection is normal.

[0052] In the embodiment of the present disclosure, the data transmission and reception situation can be monitored by a monitoring thread to detect whether the connection between the main node device and the second network is normal. When necessary, the main node can also detect whether the connection with the second network is normal by actively initiating a ping of the second network. In some embodiments, the second network connection is abnormal, including: the second network connection is detected abnormally a preset number of times, the preset number of times can be three times. For example, when the WAN connection abnormality is detected for the first time, a second detection is performed. If the connection is still abnormal, a third detection is performed. If the WAN connection is still abnormal in this detection, the second network connection is considered abnormal. Continuously detecting the second network connection status for multiple times can improve the accuracy of network connection status detection, avoid frequent networking, and improve network stability.

[0053] If the first network connection is normal and the second network connection is abnormal, step S22 is executed.

[0054] In step S22, a second network abnormality notification message is sent to each sub-node device connected to the master node device.

[0055] If the Mesh network connection is normal and the Internet connection is abnormal, the master node device sends a second network abnormality notification message to the downstream child node devices connected to it, so that each child node device can switch from normal working mode to waiting mode, and each child node device starts timing for a preset time locally.

[0056] In step S23, the connection of the second network is repaired.

[0057] In order to avoid frequent networking and improve network stability, in the event of an abnormal WAN connection, the master node device can also repair the connection failure of the second network. For example, the WAN connection can be repaired by soft restart (restarting the protocol stack).

[0058] In step S24, it is determined whether the second network connection is restored to normal within the first preset time period. If the second network connection is restored to normal within the first preset time period, step S25 is executed; otherwise, the process ends.

[0059] In step S25, a second network recovery notification message is sent to each sub-node device connected to the master node device.

[0060] If the WAN connection is successfully repaired within the first preset time, that is, the WAN connection returns to normal, the master node device sends a second network recovery notification message to the downstream child node devices connected to it, so that each child node device can switch from waiting mode to normal working mode and clear the timing of the first preset time.

[0061] In step S26, the connection of the first network is repaired.

[0062] If the Mesh network connection is abnormal, the master node device can also attempt to repair the Mesh network connection. For example, the master node device determines whether the Mesh network connection abnormality is caused by a backhaul problem or a Mesh master process problem. If the problem is a backhaul problem, it will be repaired by restarting the master process. If the problem is a Mesh master process problem, it will be repaired by restarting the WiFi driver.

[0063] If the Mesh network connection can be repaired, the child node device switches to the normal working mode within the timeout period C; if the Mesh network connection cannot be repaired, the child node device can only wait for the timeout period C to expire before switching to the fast networking mode.

[0064] In step S27, it is determined whether the first network connection is restored to normal within the second preset time period. If so, step S28 is executed; otherwise, the process ends.

[0065] In step S28, a first network recovery notification message is sent to each sub-node device connected to the master node device.

[0066] If the master node device restores the Mesh network connection to normal within the second preset time period, the master node device sends a first network recovery notification message to each child node device connected to the master node device, so that each child node device can switch from waiting mode to normal working mode, and clear the timing of the second preset time period locally.

[0067] It should be noted that the user can pre-set whether to repair the WAN or Mesh connection when an abnormality occurs on the master node device. If the user chooses not to repair the abnormality, the sub-node device can directly enter the fast networking mode from the normal working mode.

[0068] FIG4 is a schematic diagram of various mode switching of a sub-node device provided in an embodiment of the present disclosure. As shown in FIG4 , the sub-node device has the following four modes: normal operation mode, waiting mode, fast networking mode, and backup master node mode.

[0069] The waiting mode is used to wait for the upstream node to send a network recovery notification message; the fast networking mode is used to perform the Mesh network discovery process and re-establish a WiFi or LAN connection with the new master node device; the backup master node mode can replace the master node device to continue managing the Mesh network when a problem occurs with the master node device.

[0070] The conditions for switching between normal operation mode and standby mode are a Mesh / WAN connection failure or Mesh / WAN connection recovery. The conditions for switching between normal operation mode and backup master node mode are the designation or cancellation of the master node device. Backup master node mode is a special mode for slave nodes. The conditions for switching between this mode and standby mode are the same as those for switching between normal operation mode and standby mode. The conditions for switching between standby mode and fast networking mode are a wait timeout, in which case the switch from standby mode to fast networking mode occurs.

[0071] In some embodiments, after determining the backup master node device based on the sub-node attribute information and sending a backup master node notification message to the backup master node device (i.e., step S12), the method may further include the following steps: in the case of an update of the backup information, sending a backup information notification message carrying the updated backup information to the backup master node device, and receiving a backup information confirmation message sent by the backup master node device; or, receiving a backup information request message sent by the backup master node device, and sending a backup information response message carrying the backup information to the backup master node device. The determined backup master node device can obtain the backup information by actively sending a backup information request message to the master node device, or passively receiving a backup information notification message sent by the master node device, in preparation for the subsequent switch to the master node device.

[0072] Backup information is the information stored in the master node device for use in Mesh networking, and may include at least one of the following: network topology information, identity authentication information (such as Mesh connection key), routing information, connection device information of each sub-node, etc.

[0073] The embodiment of the present disclosure further provides a network fault recovery method, which is applied to a sub-node device in a first network. As shown in FIG5 , the method may include the following steps S31 and S32 .

[0074] In step S31, first sub-node attribute information of the sub-node device is sent to the master node device to which the sub-node device is connected.

[0075] In the case that the upstream node device of the child node device is the master node device, the child node device sends its own first child node attribute information to the master node device.

[0076] In some embodiments, the sub-node device sends the first sub-node attribute information of the sub-node device to the main node device to which the sub-node device is connected according to a preset first period, or sends the first sub-node attribute information of the sub-node device to the main node device to which the sub-node device is connected when the first sub-node attribute information of the sub-node device changes.

[0077] In step S32, upon receiving the second sub-node attribute information sent by the downstream sub-node device of the sub-node device, the second sub-node attribute information is sent to the main node device; wherein the first sub-node attribute information and the second sub-node attribute information include at least one of the following: capability information, load information, and connection information between the sub-node device and the terminal device of the sub-node device, and the first sub-node attribute information and the second sub-node attribute information are used to determine the backup main node device.

[0078] When the upstream node device of the sub-node device is another sub-node device in the first network, the sub-node device sends its own second sub-node attribute information to its upstream sub-node device, so that the second sub-node attribute information is sent to the master node device step by step.

[0079] Each sub-node device can carry sub-node attribute information in the extended information of the 1905 protocol. The sub-node attribute information is the information of the sub-node device itself, which can include one of the following or any combination: capability information of the sub-node device, load information, and connection information between the sub-node device and the terminal device. Capability information can include the theoretical supported bandwidth of the sub-node device and the supported network standard information, for example, whether it supports 5G network, etc.; load information can include the number of data packets sent and received and the transmission rate of the terminal device to which the sub-node device is connected, such as TX and RX data sending and receiving information; the connection information between the sub-node device and the terminal device can include the number of terminal devices connected to the sub-node device, the connection duration, and the traffic used by the terminal device on the sub-node device. The connection information between the sub-node device and the terminal device can be a connection list, which can reflect the user's usage preferences. For example, the terminal device accesses the router in the living room the most times and has the longest connection time, indicating that the user's frequent activities are in the living room.

[0080] The network fault recovery method provided by the embodiment of the present disclosure is that each sub-node device reports its own sub-node attribute information to the main node device in a step-by-step manner, so that the main node device can determine the backup main node device when the network is normal. In this way, when a network failure occurs, the backup main node device can automatically switch to the main node device to re-organize the network, thereby realizing automatic recovery of the network failure; the selection of the backup main node device will take into account at least one of the sub-node device's own capabilities, the load of the entire network, and the terminal device's habitual preference for connecting to the sub-node device. The network performance after recovery is higher, and can provide users with a better data service experience.

[0081] In some embodiments, as shown in FIG6 , the method may further include the following steps S41 to S43 .

[0082] In step S41, a second network abnormality notification message sent by the master node device is received, and the sub-node device is switched from a normal working mode to a waiting mode.

[0083] The second network is a network to which the master node device is connected upstream. In the embodiment of the present disclosure, the second network is the Internet. The child node device receives a second network abnormality notification message from the master node device, indicating that the WAN connection between the master node device and the Internet is abnormal. In this case, the child node device switches from a normal working mode to a standby mode and begins timing a first preset timer, waiting for the master node device's WAN connection to return to normal.

[0084] In step S42, when a second network recovery notification message sent by the master node device is received within the first preset time period, the sub-node device is switched from the waiting mode to the normal working mode.

[0085] If the sub-node device receives the second network recovery notification message sent by the master node device within the first preset time, it means that the master node device has repaired the WAN connection failure in time, then the sub-node device can switch from waiting mode to normal working mode and clear the timing of the first preset time.

[0086] In step S43, when the second network recovery notification message sent by the master node device is not received within the first preset time period, the sub-node device is switched from the waiting mode to the fast networking mode.

[0087] If the sub-node device does not receive the second network recovery notification message sent by the master node device within the first preset time, it means that the master node device cannot repair the WAN connection failure in time, and the sub-node device can switch from the waiting mode to the fast networking mode.

[0088] Because a Mesh network involves multiple sub-node devices, users may be unable to determine which sub-node device has a problem, resulting in a poor user experience. To address this issue, embodiments of the present disclosure add a first network connection detection function to the sub-node devices. Therefore, in some embodiments, as shown in Figure 7, the method may further include the following steps S51 to S55.

[0089] In step S51, the connection status of the first network is detected. If the first network connection is abnormal, step S52 is executed; otherwise, the process ends.

[0090] In the disclosed embodiment, the first network connection abnormality includes: detecting the first network connection abnormality a preset number of times, which may be three times. For example, when the Mesh network connection abnormality is first detected, a second detection is performed. If the connection abnormality is still detected, a third detection is performed. If the connection abnormality is still detected in this detection, the Mesh network connection is considered abnormal. Continuously detecting the Mesh network connection status multiple times can improve the accuracy of network connection status detection, avoid frequent networking, and improve network stability.

[0091] Each sub-node device can detect the connection status of the first network according to a preset period, that is, detect the status of its own access to the first network. In this way, if the connection is abnormal, it can be known which sub-node has the problem.

[0092] In step S52, the sub-node device is switched from the normal working mode to the waiting mode.

[0093] If the sub-node device detects that its own connection in the first network is abnormal, the sub-node device switches from the normal working mode to the waiting mode, and waits for the first network connection to be restored within a second preset time period.

[0094] In step S53, it is determined whether a message broadcast by the master node device or a first network recovery notification message sent by the master node device is received within the second preset time period. If so, step S54 is executed; otherwise, step S55 is executed.

[0095] In step S54, the child node device is switched from the waiting mode to the normal mode.

[0096] If the child node device receives a message broadcast by the master node device or a first network recovery notification message sent by the master node device within the second preset time period, indicating that the master node device has repaired the first network connection in time, the child node device can switch from the waiting mode back to the normal working mode and clear the timing of the second preset time period.

[0097] In step S55, the sub-node device is switched from the waiting mode to the fast networking mode.

[0098] If the sub-node device does not receive the message broadcast by the master node device or the first network recovery notification message sent by the master node device within the second preset time, it means that the master node device cannot repair the first network connection in time, and the sub-node device can switch from waiting mode to fast networking mode.

[0099] In some embodiments, as shown in FIG8 , the method may further include the following steps S61 to S65 .

[0100] In step S61, the connection status of the second network is detected. If the second network connection is abnormal, step S62 is executed; otherwise, the process ends.

[0101] In an embodiment of the present disclosure, the second network connection abnormality includes: detecting the second network connection abnormality a preset number of times, which can be three times. For example, when the WAN connection abnormality is first detected, a second detection is performed. If the connection abnormality is still detected, a third detection is performed. If the WAN connection abnormality is still detected in this detection, the second network connection is considered abnormal. Multiple consecutive detections of the second network connection status can improve the accuracy of network connection status detection, avoid frequent networking, and improve network stability.

[0102] In step S62, the sub-node device is switched from the normal working mode to the waiting mode.

[0103] If the sub-node device detects that the second network connection is abnormal, the sub-node device switches from the normal working mode to the waiting mode and waits for the second network connection to be restored within a first preset time period.

[0104] In step S63, it is determined whether the second network recovery notification message sent by the master node device is received within the first preset time period. If so, step S64 is executed; otherwise, step S65 is executed.

[0105] In step S64, the child node device is switched from the waiting mode to the normal mode.

[0106] If the child node device receives the second network recovery notification message sent by the master node device within the first preset time, it means that the master node device has repaired the second network connection in time, then the child node device can switch from the waiting mode to the normal working mode and clear the timing of the first preset time.

[0107] In step S65, the sub-node device is switched from the waiting mode to the fast networking mode.

[0108] If the sub-node device does not receive the second network recovery notification message sent by the master node device within the first preset time, it means that the master node device cannot repair the second network connection in time, and the sub-node device can switch from the waiting mode to the fast networking mode.

[0109] It should be noted that since the backup master node device is a special sub-node device, the backup master node device will also detect the connection status of the first network and the second network. The backup master node mode of the backup master node device is equivalent to the normal working mode of an ordinary sub-node device. The process of switching between the backup master node mode and the waiting mode of the backup master node device is the same as the switching process between the normal working mode and the waiting mode of the ordinary sub-node device.

[0110] In some embodiments, when the child node device is a backup master node device, as shown in FIG9 , the method may further include the following steps S71 and S72 .

[0111] In step S71, a backup master node notification message sent by the master node device is received, and the sub-node device is switched from the normal working mode to the backup master node mode.

[0112] If a child node device receives a backup master node notification message sent by the master node device, it indicates that the child node device has been selected as the backup master node device, and therefore, can switch from the normal working mode to the backup master node mode.

[0113] In step S72, a child node recovery message sent by the master node device is received, and the child node device is switched from the backup master node mode to the normal working mode.

[0114] If the master node backup device receives the sub-node recovery message sent by the master node device, it means that the master node device has re-determined a new backup master node device and needs to restore the original backup master node device to an ordinary sub-node device. Therefore, after receiving the sub-node recovery message sent by the master node device, the master node backup device switches from the backup master node mode to the normal working mode.

[0115] In some embodiments, after the sub-node device is switched from the normal operating mode to the backup master node mode (step S71), the method may further include the following steps: receiving a backup information notification message sent by the master node device, and obtaining the backup information carried in the backup information notification message; or, sending a backup information request message to the master node device according to a preset second period, receiving a backup information response message sent by the master node device, and obtaining the backup information carried in the backup information response message.

[0116] After determining the backup master node device, in order to subsequently perform Mesh network reorganization, it is necessary to store backup information for network reorganization in the backup master node device. In the embodiment of the present disclosure, the backup master node device uses two methods to obtain the backup information: (1) the master node device actively sends the backup information, in which the backup information update triggers the master node device to send the backup information to the backup master node device, so that the backup master node device can subsequently perform Mesh rapid networking based on the backup information; (2) the master node device passively sends the backup information, in which the backup master node device periodically requests the master node device to obtain the backup information.

[0117] After the waiting mode times out, the backup master node device switches to fast networking mode. Once in fast networking mode, it reconstructs the Mesh master node based on the backup information obtained from the master node device, switching from the backup master node device to the master node device. All sub-node devices in the mesh network successively enter fast networking mode. The new master node device initiates the Mesh master control process and monitoring thread. The master control process reconstructs access information based on the previously acquired Mesh connection key information, waiting for each sub-node device to connect until all connectable sub-node devices on the mesh network have successfully connected. Other sub-node devices then begin the Mesh discovery process, searching for a connectable master node device. The new master node device performs a sub-node backhaul connection verification based on the stored Mesh connection key information. Only if the verification passes will the sub-node device be allowed to access the mesh network.

[0118] Figure 10a is a schematic diagram of the mesh network topology for the initial networking provided by an embodiment of the present disclosure. As shown in Figure 10a, the controller is the master node device, the agent is the sub-node device, and the agent can connect to the terminal device. Figure 10b is a schematic diagram of the mesh network topology after selecting a backup master node device provided by an embodiment of the present disclosure. As shown in Figure 10b, the controller selects an agent as the backup master node device. Figure 10c is a schematic diagram of the topology after rapid reorganization provided by an embodiment of the present disclosure. As shown in Figure 10c, after the Mesh network is reorganized, the backup master node device in Figure 10b becomes the master node device, and the original master node device exits the Mesh network.

[0119] The embodiment of the present disclosure uses two fault recovery technologies (i.e., soft restart and sub-node competition for backup master nodes) when problems occur in the Mesh network master node device, which can effectively solve the problems of Mesh disconnection and unusability caused by WAN connection and Mesh connection encountered by users during the use of the Mesh network. The solution of the embodiment of the present disclosure does not require human participation and can be completed independently, greatly improving the user experience.

[0120] The disclosed embodiments can effectively resolve WiFi Mesh network failures, and are particularly suitable for scenarios where an unrepairable problem with the master node device paralyzes the entire Mesh network. They can repair and resolve Mesh network failures caused by abnormal WAN or WiFi / LAN connections. The disclosed embodiments can quickly restore the Mesh network without requiring user intervention to reconnect. In the disclosed embodiments, the master node device and sub-node devices regularly check whether the WAN and Mesh connections are functioning properly. In addition to regularly reporting connection information between its downstream sub-nodes and terminal devices, each sub-node device also reports its own WAN / LAN service capabilities and service status to its upstream node device, including TX and RX packet transmission and reception status, as well as the actual and theoretical WAN / LAN capacity. The master node device collects the WAN / LAN service capabilities and service status of all sub-node devices, comprehensively assesses the network usage of terminal devices connected to each sub-node device in the Mesh network, identifies user preferences, and selects a backup master node device based on a comprehensive assessment of the network topology and the network load of each sub-node device. The selected backup master node device will broadcast, and each sub-node device will save the information of the backup master node device for use when the Mesh quickly re-networks. The backup master node device can obtain the topology map and other necessary networking information from the master node device and save it locally. If the master node device detects a WAN connection anomaly, it will attempt to repair it by itself and restart the relevant protocol stack. The sub-node device will also switch to waiting mode. If the master node successfully repairs the WAN connection, it will notify each sub-node device that the network connection has been restored, and each sub-node device will switch back to normal working mode; if it cannot be repaired, it means that there is a problem with the master node device itself, and a rapid re-networking operation will be performed at this time. Accordingly, the backup master node device will be converted to the master node device based on the backup information obtained from the master node. The backup information includes fronthaul (forward return) and backhaul (return) information. Other sub-node devices will switch to fast networking mode after the waiting mode times out. The master node device will then receive the network access request from the sub-node device, respond and establish a connection.

[0121] The present disclosure also provides a master node device, as shown in FIG11a . The master node device includes a fault detection module, a message transceiver module, a notification module, a problem repair module, and a backup master node selection module. The fault detection module detects abnormalities in the WAN and Mesh connections. For example, three detections can determine whether a connection is abnormal. The problem repair module performs repairs based on the fault detection module's results. If the WAN connection is abnormal, the WAN-side repair function is activated, performing a soft reboot, including restarting the protocol stack and powering off and on the WAN module. For Mesh connection abnormalities, the repair mechanism includes silently turning WiFi on and off, reloading the WiFi driver and FW (Firewall), and powering on and off the WiFi module. The message transceiver module is used to periodically send broadcast messages, receive topology change information from child nodes, and report messages. The notification module is used to notify each child node of the abnormality when the master node determines that its own abnormality cannot be recovered and the Mesh connection is normal. It also notifies the backup master node device of changes in the topology map or terminal connection. The backup master node selection module is used to select a backup master node device according to the sub-node attribute information obtained by the message transceiver module and the priority order set by the user.

[0122] The disclosed embodiments also provide a child node device, as shown in Figure 11b. The child node device includes a fault detection module, a message transceiver module, a notification module, a mode switching module, an information storage module, a problem repair module, and a backup master node selection module. The fault detection module determines whether the child node's upstream WAN and upstream Mesh connections are functioning properly. Its internal loop checks the WAN and Mesh connections and determines whether a fault has occurred based on messages provided by the message transceiver module. The message transceiver module reports the child node's own and downstream device information to the upstream node and receives broadcast messages within the Mesh to synchronize Mesh topology, master node information, and other information. The notification module broadcasts messages when it detects a WAN connection anomaly and the Mesh is functioning properly. For devices that become backup master nodes, the notification module informs other child node devices of the backup master node's identity to facilitate subsequent Mesh networking. The mode switching module manages the state machines for the four modes of the child node device. The information storage module stores Mesh connection key information, topology maps, and terminal connection lists for each child node device, as reported by the master node device. The problem repair module and the backup master node selection module are enabled when the child node device becomes the master node device, and their specific functions are the same as those of the problem repair module and the backup master node selection module in the master node device.

[0123] The present disclosure also provides a master node device and a sub-node device. As shown in Figure 12, the master node device or sub-node device may include: at least one processor 1201; a memory 1202 storing at least one program, which, when executed by the at least one processor, enables the at least one processor to implement the network fault recovery method provided in the aforementioned embodiments; and at least one I / O interface 1203, connected between the at least one processor and the memory and configured to enable information exchange between the processor and the memory.

[0124] Among them, the processor 1201 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory 1202 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read-write interface) 1203 is connected between the processor 1201 and the memory 1202, and can realize information interaction between the processor 1201 and the memory 1202, including but not limited to a data bus (Bus), etc.

[0125] In some embodiments, the processor 1201 , the memory 1202 , and the I / O interface 1203 are connected to each other via a bus, and further connected to other components of the computing device.

[0126] The embodiments of the present disclosure further provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed, implements the network failure recovery method provided in the aforementioned embodiments.

[0127] It will be appreciated by those skilled in the art that all or some of the steps in the method disclosed above, and the functional modules / units in the device can be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0128] Example embodiments have been disclosed herein, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly indicated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the invention as set forth in the appended claims.

Claims

1. A network failure recovery method, applied to a master node device in a first network, the method comprising: Receiving sub-node attribute information sent by each sub-node device connected to the master node device, the sub-node attribute information including at least one of the following: capability information and load information of the sub-node device, and connection information between the sub-node device and the terminal device; A backup master node device is determined according to the sub-node attribute information, and a backup master node notification message is sent to the backup master node device; the backup master node device is a sub-node device of the first network, and is used to replace the master node device when the first network fails and is re-established.

2. The method of claim 1, wherein: The step of determining a backup master node device according to the subnode attribute information includes: The backup master node device is determined according to the sub-node attribute information and the priority or weight of the preset sub-node attribute information.

3. The method of claim 1, wherein: After determining a backup master node device according to the subnode attribute information and sending a backup master node notification message to the backup master node device, the method further includes: In the case that the backup master node device is different from the backup master node device determined last time, a sub-node recovery message is sent to the backup master node device determined last time, and the sub-node recovery message is used to indicate that the backup master node device determined last time is restored to a sub-node device; wherein, the backup master node device determined last time is determined by the master node device based on the sub-node attribute information received last time.

4. The method of claim 1, wherein: The master node device accesses the second network, and the method further includes: Detecting a connection status of the first network and / or a connection status of the second network.

5. The method of claim 4, further comprising: When the first network connection is normal and the second network connection is abnormal, sending a second network abnormality notification message to each sub-node device connected to the master node device; Restoring the connection of the second network; When the second network connection returns to normal within a first preset time period, a second network recovery notification message is sent to each sub-node device connected to the master node device.

6. The method of claim 4, further comprising: When the first network connection is abnormal, repairing the connection of the first network; When the first network connection returns to normal within a second preset time period, a first network recovery notification message is sent to each sub-node device connected to the master node device.

7. The method of claim 1, wherein: After determining a backup master node device according to the subnode attribute information and sending a backup master node notification message to the backup master node device, the method further includes: In the case where the backup information is updated, sending a backup information notification message carrying the updated backup information to the backup master node device, and receiving a backup information confirmation message sent by the backup master node device; or, A backup information request message sent by the backup master node device is received, and a backup information response message carrying backup information is sent to the backup master node device.

8. A network failure recovery method, applied to a sub-node device in a first network, the method comprising: Sending first sub-node attribute information of the sub-node device to the master node device connected to the sub-node device; When receiving the second sub-node attribute information sent by the downstream sub-node device of the sub-node device, sending the second sub-node attribute information to the master node device; The first subnode attribute information and the second subnode attribute information include at least one of the following: capability information, load information, and connection information between the subnode device and the terminal device. The first subnode attribute information and the second subnode attribute information are used to determine the backup master node device.

9. The method of claim 8, wherein: The sending the first sub-node attribute information of the sub-node device to the master node device connected to the sub-node device includes: The first subnode attribute information of the subnode device is sent to the main node device to which the subnode device is connected according to a preset first period, or the first subnode attribute information of the subnode device is sent to the main node device to which the subnode device is connected when the first subnode attribute information of the subnode device changes.

10. The method of claim 8, further comprising: A second network abnormality notification message sent by the master node device is received, and the sub-node device is switched from a normal working mode to a waiting mode, wherein the second network is a network to which the master node device is connected uplink.

11. The method of claim 10, wherein: After switching the sub-node device from the normal working mode to the waiting mode, the method further includes: When receiving a second network recovery notification message sent by the master node device within a first preset time period, switching the sub-node device from the waiting mode to the normal working mode; or, When the second network recovery notification message sent by the master node device is not received within the first preset time period, the sub-node device is switched from the waiting mode to the fast networking mode.

12. The method of claim 8, further comprising: detecting a connection status of the first network; In case that the first network connection is abnormal, switching the sub-node device from a normal working mode to a waiting mode; When receiving a message broadcast by the master node device or a first network recovery notification message sent by the master node device within a second preset time period, switching the sub-node device from the waiting mode to the normal mode; or, When no message broadcast by the master node device or a first network recovery notification message sent by the master node device is received within a second preset time period, the sub-node device is switched from the waiting mode to the fast networking mode.

13. The method of claim 8, further comprising: Detecting a connection status of a second network, where the second network is a network to which the master node device is connected uplink; In case that the second network connection is abnormal, switching the sub-node device from a normal working mode to a waiting mode; When receiving a second network recovery notification message sent by the master node device within a first preset time period, switching the sub-node device from the waiting mode to the normal mode; or, When the second network recovery notification message sent by the master node device is not received within the first preset time period, the sub-node device is switched from the waiting mode to the fast networking mode.

14. The method of claim 8, further comprising: Receive a backup master node notification message sent by the master node device, and switch the sub-node device from a normal working mode to a backup master node mode.

15. The method of claim 14, wherein: After switching the sub-node device from the normal working mode to the backup master node mode, the method further includes: Receive a sub-node recovery message sent by the master node device, and switch the sub-node device from the backup master node mode to a normal working mode.

16. The method of claim 14, wherein: After switching the sub-node device from the normal working mode to the backup master node mode, the method further includes: receiving a backup information notification message sent by the master node device, and acquiring the backup information carried in the backup information notification message; or, A backup information request message is sent to the master node device according to a preset second period, a backup information response message sent by the master node device is received, and the backup information carried in the backup information response message is acquired.

17. A master node device, comprising: at least one processor; a storage device having at least one program stored thereon; When the at least one program is executed by the at least one processor, the at least one processor implements the network fault recovery method according to any one of claims 1 to 7; At least one I / O interface is connected between the at least one processor and the storage device and is configured to implement information interaction between the at least one processor and the storage device.

18. A subnode device, comprising: at least one processor; a storage device having at least one program stored thereon; When the at least one program is executed by the at least one processor, the at least one processor implements the network fault recovery method according to any one of claims 8 to 16; At least one I / O interface is connected between the at least one processor and the storage device and is configured to implement information interaction between the at least one processor and the storage device.

19. A computer readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the network fault recovery method according to any one of claims 1 to 7 is implemented, or the network fault recovery method according to any one of claims 8 to 16 is implemented.