A message forwarding method and device

By selecting a fault proxy node in the DDC cloud cluster and using the data network link to forward management messages, the stability problem of the DDC cloud cluster when the management network communication fails is solved, ensuring the continuity of communication and the stability of the equipment.

CN119210996BActive Publication Date: 2025-11-07NEW H3C TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411378267.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-11-07
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

In the DDC cloud cluster, equipment failures can cause management network communication interruptions, leading to widespread business disruptions. Therefore, it is necessary to improve equipment stability and fault protection.

Method used

When communication between the main cloud network control unit and the cloud network service unit fails, the second cloud network service unit that is in normal communication with the main cloud network control unit is selected as the fault proxy node. Management messages are forwarded through the data network link to establish an escape channel to ensure communication continuity.

Benefits of technology

It improves the operational stability of the DDC cloud cluster in the event of a management network failure, thus avoiding widespread business interruption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119210996B_ABST
    Figure CN119210996B_ABST
Patent Text Reader

Abstract

The application provides a message forwarding method and device. In one example, the method comprises: in the case that a first cloud network service unit and a master cloud network control unit manage a network communication fault, selecting a fault proxy node from a second cloud network service unit in a distributed decoupling machine frame cloud cluster; in the case that the fault proxy node is determined, for a management message to be forwarded, sending the management message to be forwarded to the fault proxy node through a data network link between the first cloud network service unit and the fault proxy node, and sending the management message to be forwarded by the fault proxy node to a destination device. The application can improve the stability of the DDC cloud cluster.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of network communication technology, and in particular to a message forwarding method and device. BACKGROUND

[0002] DDC (Distributed Disaggregated Chassis) technology solves the upgrade iteration problem of DCI (Data Center Interconnect) network by distributing large chassis devices, using box switches as forwarding line cards and switching network boards, and flexibly distributing in multiple cabinets, thereby solving the upgrade iteration problem of DCI (Data Center Interconnect) network. Compared with traditional chassis switches, DDC not only can build larger scale switching capacity clusters, but also is not limited by single cabinet space, breaking through the capacity bottleneck of chassis switches.

[0003] With the continuous expansion of the device scale, it is particularly important to ensure the stability and effective fault protection of the device. It is necessary to reduce the influence of a single device on the operation of the entire DDC cloud cluster, and to avoid large-scale business abnormal problems caused by partial device failure. SUMMARY

[0004] The present application provides a message forwarding method and device to improve the stability of DDC cloud cluster operation.

[0005] According to a first aspect of the embodiments of the present application, a message forwarding method is provided, applied to a first cloud network service unit in a distributed decoupled chassis cloud cluster, and the method comprises:

[0006] In the case that the first cloud network service unit and the main cloud network control unit manage the network communication fault, a fault proxy node is selected from a second cloud network service unit in the distributed decoupled chassis cloud cluster; wherein the fault proxy node communicates normally with the main cloud network control unit; the second cloud network service unit is other cloud network service unit in the distributed decoupled chassis cloud cluster except the first cloud network service unit;

[0007] In the case that the fault proxy node is determined, for the to-be-forwarded management message, the to-be-forwarded management message is sent to the fault proxy node through the data network link between the first cloud network service unit and the fault proxy node, and the fault proxy node sends the to-be-forwarded management message to the destination device.

[0008] According to a second aspect of the embodiments of the present application, a message forwarding device is provided, deployed in a first cloud network service unit in a distributed decoupled chassis cloud cluster, and the device comprises:

[0009] The selecting unit is used for selecting a fault proxy node from a second cloud network service unit in the distributed decoupling chassis cloud cluster in the case that the first cloud network service unit and the main cloud network control unit management network communication fault, wherein the fault proxy node is in normal communication with the main cloud network control unit, and the second cloud network service unit is other cloud network service unit in the distributed decoupling chassis cloud cluster except the first cloud network service unit;

[0010] The communication unit is used for sending the to-be-transferred management packet to the fault proxy node through the data network link between the first cloud network service unit and the fault proxy node in the case that the fault proxy node is determined, and sending the to-be-transferred management packet to the destination device by the fault proxy node.

[0011] According to the technical scheme disclosed by the application, in the case that the first cloud network service unit and the main cloud network control unit management network communication fault, a second cloud network service unit in normal communication with the main cloud network control unit is selected as a fault proxy node from the distributed decoupling chassis cloud cluster, and the to-be-transferred management packet is sent to the fault proxy node through the data network link between the first cloud network service unit and the fault proxy node, and the to-be-transferred management packet is sent to the destination device by the fault proxy node, so that the data network link between the cloud network service units is used as an escape channel of the management packet of the cloud network service unit, and the stability of the DDC cloud cluster operation is improved. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 is a flow diagram of a packet transfer method provided by an embodiment of the application;

[0013] Figure 2A is a management network normal case management packet transfer path schematic diagram;

[0014] Figure 2B is a management network fault case management packet transfer path schematic diagram provided by an embodiment of the application;

[0015] Figure 3 is a DDC cloud cluster architecture schematic diagram provided by an embodiment of the application;

[0016] Figure 4 is a scene schematic diagram of a small part of NCP management network fault provided by an embodiment of the application;

[0017] Figure 5 is a scene schematic diagram of NCF management network fault provided by an embodiment of the application;

[0018] Figure 6Ais a deployment of a DDC cloud cluster of an escape container provided by an embodiment of the present application.

[0019] Figure 6B is a deployment of an escape container provided by an embodiment of the present application, and a scenario diagram of NCC network failure.

[0020] Figure 7 is a deployment of an escape container provided by an embodiment of the present application, and a scenario diagram of MGT network failure.

[0021] Figure 8 is a structure diagram of a message forwarding device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the DDC device will be briefly described first.

[0023] The DDC device can include a network cloud controller (NCC), a network cloud fabric (NCF), and a network cloud processor (NCP).

[0024] In the DDC cloud cluster architecture, the management network between the DDC devices is connected by the MGT (Management) device to realize the communication between the devices on the board, and the DDC devices establish and maintain the forwarding path of the management network through the Router container.

[0025] In the management network, a forwarding path is usually established independently by the Router container for the inter-board communication between the DDC devices in the cluster, and all the DDC devices are connected to the intermediate MGT device to form a large internal management network. Although the internal management network supports the MGT device to be linked by multiple devices and multiple links, and to reduce the probability of link failure through multiple equivalent paths, there is still a possibility of failure.

[0026] Therefore, in order to improve the stability of the cluster device, an escape channel can be provided for the communication of the NCP in the management network in the case of DDC device management network failure, such as through the cross-board service traffic interaction mode in the data network to complete the communication interaction of the management network data.

[0027] Wherein, in the process of data interaction of the DDC device through the cross-board service traffic interaction mode in the data network, the destination port is not a faceplate port, but an internal port of the destination board, such as a port of a CPU (Center Process Unit) of the destination board.

[0028] For example, referring to Figure 2A Taking the NCP as an example, under the condition that the management network is normal, the management message sent by the service container 2011 of the NCP 201 to the NCP 202 can be sent by the forwarding module (not shown in the figure) of the Router container 2012 to the forwarding module (not shown in the figure) of the Router container 2022 of the NCP 202 through the management channel, and the forwarding module of the Router container 2022 can forward the management message to the service container 2021 of the NCP 202 according to the destination slot number of the management message.

[0029] For example, referring to Figure 2B Under the condition that the management network is faulty, the escape channel can be enabled, the management message sent by the service container 2011 of the NCP 201 to the NCP 202 can be sent by the forwarding module (not shown in the figure) of the Router container 2012 to the forwarding module (not shown in the figure) of the Router container 2022 of the NCP 202 through the management channel, and the forwarding module of the Router container 2022 can forward the management message to the service container 2021 of the NCP 202 according to the destination slot number of the management message.

[0030] In order to make the above-mentioned purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the drawings.

[0031] For example, referring to Figure 1 A flowchart of a packet forwarding method provided by the embodiments of the present application is shown in the figure, wherein the packet forwarding method can be applied to any NCP (which can be referred to as a first NCP) in a DDC cloud cluster, such as Figure 1 The packet forwarding method can include the following steps:

[0032] Step 101, in the case that the first cloud network service unit and the main cloud network control unit manage network communication failure, a failure proxy node is selected from a second cloud network service unit in the distributed decoupling machine frame cloud cluster; wherein the failure proxy node communicates normally with the main cloud network control unit; the second cloud network service unit is other cloud network service unit in the distributed decoupling machine frame cluster except the first cloud network service unit.

[0033] In the embodiment of the application, in the case that the first NCP and the main NCC manage network communication failure, the first NCP can select a second NCP which communicates normally with the main NCC as a failure proxy node from other NCPs (which can be selected from the second NCP) in the DDC cloud cluster except the first NCP, and take the data network link between the first NCP and the failure proxy node as an escape channel of the first NCP packet.

[0034] The first NCP and the main NCC manage network communication failure can include management network link failure of the first NCP, or management network failure of the main NCC.

[0035] It should be noted that, due to the management network failure of the main NCC, and in the case that there is a backup NCC, the backup NCC will re-perform the election of the main NCC, and before the new main NCC is elected, each NCP cannot communicate normally with the main NCC; in the case that the new main NCC is selected, if the first NCP still cannot communicate with the new main NCC through the management network (such as the case of management network failure of the first NCP), and there is a second NCP which can communicate normally with the main NCC in the second NCP, the failure proxy node can be selected; if the first NCP can communicate with the new main NCC through the management network (such as the management network failure of the single device of the original main NCC), the failure proxy node can not be selected.

[0036] Step 102, in the case that the failure proxy node is determined, for the to-be-transferred management packet, the data network link between the first cloud network service unit and the failure proxy node is used to send the to-be-transferred management packet to the failure proxy node, and the failure proxy node sends the to-be-transferred management packet to the destination device.

[0037] In the embodiment of the application, in the case that the failure proxy node is determined, for the to-be-transferred management packet, the data network link between the first NCP and the failure proxy node is used to send the to-be-transferred management packet to the failure proxy node, and the failure proxy node sends the to-be-transferred management packet to the destination device.

[0038] For example, in the case that the failure proxy node and the main NCC manage network communicate normally, the failure proxy node can transfer the management packet of the first NCP through the management network.

[0039] Exemplarily, in the case that the fault agent node also fails to communicate with the main NCC management network, the fault agent node can forward the management message of the first NCP through the data network link.

[0040] It should be noted that in the case that the destination device of the management message of the first NCP is the fault agent node, the fault agent node can not need to forward out again after receiving the management message of the first NCP.

[0041] It can be seen that, in the method flow shown in the figure, Figure 1 In the method flow shown in the figure, in the case that the first cloud network service unit fails to communicate with the main cloud network control unit management network, a second cloud network service unit that normally communicates with the main cloud network control unit is selected from the distributed decoupling machine frame cloud cluster as a fault agent node. For the to-be-forwarded management message, the data network link between the first cloud network service unit and the fault agent node is used to send the to-be-forwarded management message to the fault agent node, and the fault agent node sends the to-be-forwarded management message to the destination device. By using the data network link between the cloud network service units as an escape channel for the management message of the cloud network service unit, the stability of the DDC cloud cluster operation is improved.

[0042] In some embodiments, the above selecting a fault agent node from the second cloud network service unit in the distributed decoupling machine frame cloud cluster comprises:

[0043] sending a fault agent request message to the second cloud network service unit through the data network link between the first cloud network service unit and the second cloud network service unit;

[0044] In the case that the fault agent request response message sent by the second cloud network service unit is received, the fault agent node is selected from the second cloud network service unit that sends the fault agent request response message; wherein the fault agent request response message is sent by the second cloud network service unit to the first cloud network service unit through the data network link in the case that the fault agent request message is received and it is determined that the node normally communicates with the main cloud network control unit.

[0045] Exemplarily, in the case that the first NCP fails to communicate with the main NCC management network, the first NCP can send a fault agent request message to the second NCP through the data network link.

[0046] It should be noted that in the process of sending the fault agent request message by the first NCP, the fault agent request message is sent to all the second NCPs in the DDC cloud cluster before the failure of the first NCP; wherein for the NCPs that are newly added in the DDC cloud cluster after the failure of the first NCP and before the selection of the fault agent node, the first NCP does not obtain that these NCPs are newly added in the DDC cloud cluster, and accordingly, the fault agent request message is not sent to these NCPs.

[0047] In a case where the second NCP receives the fault proxy request message sent by the first NCP, it can be determined whether the node and the master NCC can normally communicate.

[0048] The second NCP and the master NCC can normally communicate includes that the second NCP and the master NCC can communicate through the management network, or the second NCP and the master NCC can communicate through the data network.

[0049] In a case where the second NCP determines that the node and the master NCC can normally communicate, the second NCP can send a fault proxy request response message to the first NCP.

[0050] In a case where the first NCP receives the fault proxy request response message sent by the second NCP, the first NCP selects a fault proxy node from the second NCP that sends the fault proxy request response message.

[0051] In an example, the first NCP can select the second NCP that sends the first fault proxy request response message as the fault proxy node according to the received first fault proxy request response message.

[0052] It should be noted that in the embodiments of the present application, for any NCP, after the NCP selects a proxy fault node, if it is determined that the management network communication between the device and the master NCC is restored, the forwarding of the management message can be switched back to be forwarded through the management network, and there is no need to be forwarded through the fault proxy node again.

[0053] In some embodiments, after the above selecting the fault proxy node from the second cloud network service unit in the distributed decoupling machine frame cloud cluster, the method further includes:

[0054] sending a fault proxy node notification message to the fault proxy node through the data network link between the first cloud network service unit and the fault proxy node, so that the fault proxy node receives the fault proxy node notification message and sends the fault proxy node notification message to other distributed decoupling machine frame devices; wherein the fault proxy node notification message is used to notify the fault proxy node selected by the first cloud network service unit, so that the other distributed decoupling machine frame devices forward the management message sent to the first cloud network service unit through the fault proxy node.

[0055] For example, considering that the first NCP selects a fault proxy node, the management message sent by other DDC devices in the DDC cloud cluster to the first NCP also needs to be forwarded through the fault proxy node, therefore, the fault proxy node selected by the first NCP needs to be notified to other DDC devices in the DDC cloud cluster.

[0056] Correspondingly, in the case that the first NCP selects the fault proxy node, the fault proxy node notification message can be sent to the fault proxy node through the data network link between the first NCP and the fault proxy node.

[0057] The fault proxy node notification message is used to notify the fault proxy node selected by the first NCP, i.e., to notify each DDC device in the DDC cloud cluster which NCP is selected as the fault proxy node by the first NCP.

[0058] The fault proxy node notification message can carry the identification information of the fault proxy node, such as slot number information.

[0059] In the case that the fault proxy node receives the fault proxy node notification message sent by the first NCP, the fault proxy node can forward the fault proxy node notification message to other DDC devices.

[0060] In the case that the other DDC devices receive the fault proxy node notification message, the management network forwarding path of the node can be refreshed, and for the management message that needs to be sent to the first NCP, the management message can be first sent to the fault proxy node, and then forwarded to the first NCP by the fault proxy node.

[0061] In some embodiments, in the case that the escape container is started on the first cloud network service unit, the method further comprises:

[0062] In the case that the current main cloud network control unit management network fails, the escape container is used as a virtual backup cloud network control unit to perform new main cloud network control unit election, and the priority of the escape container to be elected as the main cloud network control unit is lower than the priority of the cloud network control unit to be elected as the main cloud network control unit.

[0063] For example, in order to further improve the stability of the DDC cloud cluster and reduce the influence of the main NCC management network failure on the operation of the DDC cloud cluster, the DDC cloud cluster can deploy multiple NCC devices, and can also select to start an escape container (Escap container) on part of the NCPs.

[0064] For example, 8 NCC devices can be deployed in the DDC cloud cluster, and escape containers can be started on two NCPs, and a 1 master 9 backup specification configuration is adopted, i.e., one main NCC management cluster and nine backup NCCs are used to ensure the stability of the cluster.

[0065] The escape container can be used as a virtual NCC device to participate in the main NCC election.

[0066] The priority of the escape container to be elected as the main NCC is lower than the priority of the NCC to be elected as the main NCC.

[0067] In some embodiments, selecting a fault proxy node from the second cloud network service unit in the distributed decoupling machine frame cloud cluster can include:

[0068] selecting a target cloud network service unit in the second cloud network service unit as the fault proxy node; wherein the target cloud network service unit starts an escape container, and the escape container is elected as the main cloud network control unit as a virtual backup cloud network control unit.

[0069] For example, in the case that the escape container on the NCP is elected as the main NCC, and all the management networks of the NCC devices are faulty, in order to improve the interaction efficiency of the management message, the first NCP can select the second NCP (which can be referred to as a target NCP) in which the started escape container is elected as the main NCC as the fault proxy node in the process of selecting the fault proxy node.

[0070] In order to make the skilled in the art better understand the technical solutions provided by the embodiments of the present application, the technical solutions provided by the embodiments of the present application will be described below in conjunction with specific application scenarios.

[0071] Please refer to Figure 3 , an architecture schematic diagram of a DDC cloud cluster provided by the embodiments of the present application is shown, as Figure 3 indicated, the DDC cloud cluster can include a plurality of NCCs (only two NCCs, i.e., NCC311 and NCC312, are shown in the figure), an MGT device 340, a plurality of NCFs (only two NCFs, i.e., NCF321 and NCF322, are shown in the figure), and a plurality of NCPs (only three NCPs, i.e., NCP331, NCP332, and NCP333, are shown in the figure).

[0072] As shown in Figure 3 , the NCPs realize data network interaction through the NCFs, and the NCPs, the NCFs, and the NCCs realize management network interaction through the MGT device.

[0073] In Figure 3 , the solid line is a data network link, and the dashed line is a management network link.

[0074] Among them, for any DDC device, a plurality of links can be adopted in a backup manner to support link ECMP (Equal-Cost Multi-Path, equal-cost multi-path).

[0075] It should be noted that Figure 3Taking the example of data network links between each NCP and each NCF (which can be called full connectivity), in actual applications, considering the large number of DDC devices, full connectivity may not be possible due to the number of device ports. Instead, a connection between NCP and NCF can be established according to actual needs, ensuring that the link has backups and improving reliability while reducing the limitations caused by the number of device ports.

[0076] Based on the above scenario, the following section will explain the message forwarding implementation with specific examples.

[0077] Example 1: A small number of NCP management network failures.

[0078] In the event of a partial agent escape function when a small number of NCPs experience a management network failure and are unable to communicate with the current primary NCC, the system will activate the partial agent escape function. The NCP with the management network failure will notify other NCPs and select one of them as the fault agent node. The messages of the NCP with the management network failure will be forwarded to the primary NCC through the fault agent node to ensure uninterrupted communication.

[0079] by Figure 4 Taking the scenario shown as an example, assuming line 5 fails and the NCP332 management network fails, the NCP332 management network can no longer communicate with other DDC devices, while the management networks of other DDC devices are normal, the system will activate some escape functions. Here, it is assumed that NCC311 is the current master NCC.

[0080] exist Figure 4 In the scenario shown, the NCP332 needs to select a fault proxy device first. The process is as follows:

[0081] 1) NCP322 provides emergency access to all other NCPs (in the area where the malfunction occurred) via the escape route. Figure 4 The device (including NCP331 and NCP333) sends a fault proxy request message and sets the device to fault election state, with the election time set to T1.

[0082] 1.1 The escape route from NCP332 to NCP331 includes:

[0083] Direction 1: via line 9 to NCF321, then via line 8 to NCP331;

[0084] Direction 2: via line 12 to NCF322, then via line 11 to NCP331.

[0085] 1.2 The escape route from NCP332 to NCP333 includes:

[0086] Direction 1: via line 9 to NCF321, then via line 10 to NCP333;

[0087] Direction 2: through line 12 to NCF 322, and then through line 11 to NCP 331.

[0088] 2) In the case that NCP 331 or NCP 333 receives the proxy request message of NCP 332, if the device (NCP 331 or NCP 333) is in normal communication with the current master NCC (i.e. NCC 311), the fault proxy request is answered, carrying the slot information of the board, and the fault proxy request answer message is sent to NCP 332 through the escape channel; otherwise, no fault proxy request answer message is sent.

[0089] 3) If NCP 332 receives the fault proxy request answer message within the election time T1, the first fault proxy request answer message is selected, NCP 332 cancels the fault election state, and the NCP (assuming NCP 331) that sent the fault proxy request answer message is selected as the proxy fault node.

[0090] NCP 332 will send a fault proxy node notification message to NCP 331 through the escape channel, notifying NCP 331 to act as the proxy fault node of NCP 332.

[0091] 4) NCP 331 is selected as the fault proxy node of NCP 332, and NCP 331 forwards the fault proxy node notification message to other DDC devices through the management network. The DDC devices receiving the fault proxy node notification message refresh the management network forwarding path of the device, and the management message that needs to be sent to NCP 332 will be sent to NCP 331 through the management network first, and then NCP 331 will send it to NCP 332 through the escape channel, realizing the interaction of management information between NCP 332 and other DDC devices.

[0092] Among them, the comparison of the switching of the management message interaction path of NCC 311 and NCP 332 is as follows:

[0093] Before the management network of NCP 332 fails: NCC 311-line 1-MGT 340-line 5-NCP 332;

[0094] After the management network of NCP 332 fails, assuming that NCP 331 is the fault proxy node of NCP 332, then the escape channel is line 8 and line 9:

[0095] NCC 311-line 1-MGT 340-line 4-NCP 331 (fault proxy node)-escape channel (line 8 and 9)-NCP 332.

[0096] It should be noted that in the case of NCP 331 being a fault agent node of NCP 332, lines 8 and 9, or lines 11 and 12 can be selected as escape channels in a load balancing manner.

[0097] Example II, NCF manages network failure.

[0098] As shown in Figure 5 , assuming that NCF 321 and NCF 322 both manage network failure, the NCF will be split from the cluster of the DDC, and the NCF can be designed with a default enabled traffic forwarding function and a default multicast table item, and no longer respond to the dynamic changes of the NCP. Since the port connection between the NCP and the NCF is a self-negotiation mechanism, it also does not depend on the state change of the NCP, therefore, when the NCF is split out and there is no master, it will continue to work and will not affect the entire DDC cluster network.

[0099] In this embodiment, for the ports on the NCF used to connect the NCP (which can be referred to as target ports), message forwarding channels between the target ports can be established in advance.

[0100] For any target port, when it is detected that the target port is connected to the NCP, the state of the target port is controlled to be in an open state.

[0101] Based on the above design, the NCF can realize the forwarding of the message of the NCP according to the actual connection of the NCP without the need for management data interaction with the master NCC, so that the NCF can continue to work in the case of NCF managing network failure.

[0102] Example III, NCC manages network failure.

[0103] The escape container function is introduced on the DDC cloud cluster architecture for the stability of the device, and 8 NCCs and 2 NCPs are allocated as escape containers in the specification design, which can support up to 8+2 master controls, equivalent to 10 NCCs (including two virtual NCCs), 1 master and 9 standby specifications, the master NCC manages the cluster, and the 9 standby NCCs are backed up to ensure the stability of the cluster.

[0104] After the DDC device is started, the cloud cluster will start the escape container in addition to the regular business container and the Rrouter container according to the resource priority of the NCP, which has the same role as the master container (Master) and the slave container (Standby), and is used for backup and management of the DDC device, and is a separate container externally.

[0105] In the case where there are more than two NCPs, the escape container can be started on at most two NCPs with high priority according to the classification of the NCP resource priority.

[0106] As Figure 6A shown, in the case of no failure of the NCC311 and NCC312 management network, if the escape container is distributed to the NCP332 and NCP333, there are 2 NCC devices, 2 NCF devices, 3 NCP devices and 2 NCP escape containers in the network (assuming that the escape container 3321 is started on the NCP332 and the escape container 3331 is started on the NCP333), which is equivalent to the frame device, 4 master controls (2 NCCs, 2 NCP escape containers), 2 network boards (NCFs) and 3 service boards (3 NCPs).

[0107] In the case of failure of the NCC311 and NCC312 management network, the escape container of the NCP will take over the entire cluster, as Figure 6B shown, assuming that line 1 and line 2 are both failed, the escape containers of the NCP332 and NCP333 will perform master NCC election to select the master NCC and take over the cluster to ensure the stability of the cluster.

[0108] Example four, MGT network failure.

[0109] MGT device network failure, that is, the entire management network is completely failed, and the management data interaction between devices cannot be completed through the management network. As Figure 7 shown, in the case of MGT340 network failure, line 1, line 2, line 3, line 4, line 5, line 6 and line 7 are all failed, the escape scheme in this scenario is as follows:

[0110] 1) The escape containers of the NCP332 and NCP333 perform master NCC election to elect the master NCC to take over the entire DDC cluster, for details, see example three. Among them, the management message is switched to the escape channel.

[0111] 2) The interaction of the NCP management message is realized through the escape channel to ensure the normal operation of the NCP device, for reference, see the scenario of example one.

[0112] Taking NCP331 as an example, NCP331 can trigger the failure agent process to communicate with other NCPs through the escape channel and complete the communication with the master NCC.

[0113] Assuming that the escape container started on the NCP332 is elected as the master NCC, the NCP331 and NCP333 can preferentially select the NCP332 as the failure agent node.

[0114] Among them, assuming that the failure agent node of the NCP331 is NCP332 and the failure agent node of the NCP333 is NCP332, the management message interaction path is as follows:

[0115] NCP331 and NCP332 manage the message interaction path: NCP331-escape channel (such as line 8 and line 9)-NCP332.

[0116] NCP333 and NCP332 manage the message interaction path: NCP333-escape channel (such as line 13 and line 12)-NCP332.

[0117] NCP331 and NCP333 manage the message interaction path: NCP331-escape channel (such as line 8 and line 9)-NCP332-escape channel (such as line 12 and line 13)-NCP333.

[0118] 3) NCF device is separated from the cluster and runs alone, see example two for implementation.

[0119] See Figure 8 , a structure diagram of a packet forwarding device provided by an embodiment of the present application is shown in Figure 8 , the packet forwarding device 80 can include:

[0120] The selection unit 810 is configured to select a fault proxy node from a second cloud network service unit in the distributed decoupling chassis cloud cluster in the case that the first cloud network service unit and the master cloud network control unit manage a network communication fault; wherein the fault proxy node communicates normally with the master cloud network control unit; and the second cloud network service unit is a cloud network service unit other than the first cloud network service unit in the distributed decoupling chassis cloud cluster.

[0121] The communication unit 820 is configured to, in the case that the fault proxy node is determined, send, for a to-be-forwarded management packet, the to-be-forwarded management packet to the fault proxy node through a data network link between the first cloud network service unit and the fault proxy node, and send, by the fault proxy node, the to-be-forwarded management packet to a destination device.

[0122] In some embodiments, the selection unit 810 selects the fault proxy node from the second cloud network service unit in the distributed decoupling chassis cloud cluster, including:

[0123] sending a fault proxy request packet to the second cloud network service unit through a data network link between the first cloud network service unit and the second cloud network service unit;

[0124] selecting the fault proxy node from the second cloud network service unit that sends the fault proxy request response packet in the case that the fault proxy request response packet sent by the second cloud network service unit is received; wherein the fault proxy request response packet is sent by the second cloud network service unit to the first cloud network service unit through the data network link in the case that the fault proxy request packet is received and it is determined that the node communicates normally with the master cloud network control unit.

[0125] In some embodiments, the communication unit 820 is further configured to send a fault proxy node notification message to the fault proxy node through a data network link between the first cloud network service unit and the fault proxy node, so that the fault proxy node sends the fault proxy node notification message to other distributed decoupling chassis devices when receiving the fault proxy node notification message; wherein the fault proxy node notification message is used to notify the fault proxy node selected by the first cloud network service unit, so that the other distributed decoupling chassis devices forward the management message sent to the first cloud network service unit through the fault proxy node.

[0126] In some embodiments, the message forwarding device can further include:

[0127] The election unit is configured to, when the escape container is started on the first cloud network service unit and the current master cloud network control unit manages the network fault, take the escape container as a virtual backup cloud network control unit to perform new master cloud network control unit election; wherein the priority of the escape container being elected as the master cloud network control unit is lower than the priority of the cloud network control unit being elected as the master cloud network control unit.

[0128] In some embodiments, the selection unit 810 selects the fault proxy node from a second cloud network service unit in the distributed decoupling chassis cloud cluster, including:

[0129] The selection unit 810 selects a target cloud network service unit in the second cloud network service unit as the fault proxy node; wherein the target cloud network service unit starts an escape container, and the escape container is elected as the master cloud network control unit as a virtual backup cloud network control unit.

[0130] The implementation process of the functions and roles of each unit in the device 80 is specifically described in the implementation process of the corresponding steps in the above method, which will not be repeated here. Each unit in the device 80 can be a hardware module, which can include a permanently designed special circuit or logic device (such as a special processor, such as an FPGA or an ASIC) for completing a specific operation. The hardware module can also include a programmable logic device or circuit temporarily configured by software (such as including a general-purpose processor or other programmable processor) for performing a specific operation.

[0131] The embodiment of the present application also provides a distributed decoupling chassis cloud cluster, including a cloud network control unit, a cloud network switching unit, a cloud network service unit and a management device; the cloud network control unit, the cloud network switching unit and the cloud network service unit realize management network interaction through the management device; the cloud network service unit realizes data network interaction through the cloud network switching unit; wherein:

[0132] The cloud network service unit is configured to select a fault proxy node from a second cloud network service unit in the distributed decoupling chassis cloud cluster when the cloud network service unit is a first cloud network service unit and the first cloud network service unit has a network communication fault with the master cloud network control unit; the fault proxy node has a normal communication with the master cloud network control unit; the second cloud network service unit is a cloud network service unit other than the first cloud network service unit in the distributed decoupling chassis cloud cluster; and the first cloud network service unit is configured to send the to-be-transmitted management message to the fault proxy node through a data network link between the first cloud network service unit and the fault proxy node, and send the to-be-transmitted management message to a target device by the fault proxy node.

[0133] In some embodiments, for target ports on the cloud network switching unit for connecting the cloud network service units, message forwarding channels are established between the target ports in advance.

[0134] For any target port, when it is detected that the target port is connected to a cloud network service unit, the state of the target port is controlled to be in an open state.

[0135] For example, based on the above settings of the cloud network switching unit, when the cloud network switching unit has a network fault, the cloud network switching unit can be designed to have a default enabled traffic forwarding function and establish a default multicast table item, and no longer respond to dynamic changes of the cloud network service units. The cloud network switching unit does not need to rely on state changes of the cloud network service units, and can continue to work without a master (i.e., having a network communication fault with the master cloud network control unit), and will not affect the cluster network of the entire DDC.

[0136] The implementation processes of the functions and roles of the units in the cluster are specifically described in the implementation processes of the corresponding steps in the above method, and will not be described here. Each unit in the cluster can be a hardware module, which can include a permanently designed special circuit or logic device (such as a special processor, such as an FPGA or an ASIC) for completing specific operations. The hardware module can also include a programmable logic device or circuit temporarily configured by software (such as including a general-purpose processor or other programmable processor) for performing specific operations.

Claims

1. A packet forwarding method, characterized by, A method applied to a first cloud network service unit in a distributed decoupling chassis cloud cluster, the method comprising: In the case that the first cloud network service unit and a master cloud network control unit management network communication fault, selecting a fault proxy node from a second cloud network service unit in the distributed decoupling chassis cloud cluster; wherein the fault proxy node and the master cloud network control unit communication is normal; the second cloud network service unit is other cloud network service unit in the distributed decoupling chassis cloud cluster except the first cloud network service unit; In the case that the fault proxy node is determined, for the to-be-forwarded management packet, sending the to-be-forwarded management packet to the fault proxy node through the data network link between the first cloud network service unit and the fault proxy node, and sending the to-be-forwarded management packet to a destination device by the fault proxy node.

2. The method of claim 1, wherein, The method further comprises: sending a fault proxy request packet to the second cloud network service unit through the data network link between the first cloud network service unit and the second cloud network service unit; In the case that a fault proxy request response packet sent by the second cloud network service unit is received, selecting a fault proxy node from the second cloud network service unit sending the fault proxy request response packet; wherein the fault proxy request response packet is sent to the first cloud network service unit by the second cloud network service unit through the data network link in the case that the fault proxy request packet is received and it is determined that the node and the master cloud network control unit communication is normal.

3. The method of claim 1, wherein, The method further comprises: sending a fault proxy node notification packet to the fault proxy node through the data network link between the first cloud network service unit and the fault proxy node, so that the fault proxy node sends a fault proxy node notification packet to other distributed decoupling chassis devices in the case that the fault proxy node notification packet is received; wherein the fault proxy node notification packet is used to notify the fault proxy node selected by the first cloud network service unit, so that other distributed decoupling chassis devices forward the management packet sent to the first cloud network service unit through the fault proxy node.

4. The method of claim 1, wherein, In the case that an escape container is started on the first cloud network service unit, the method further comprises: In the case that the current master cloud network control unit management network fault, taking the escape container as a virtual backup cloud network control unit to perform new master cloud network control unit election; wherein the priority of the escape container being elected as the master cloud network control unit is lower than the priority of the cloud network control unit being elected as the master cloud network control unit.

5. The method of claim 1, wherein, The method further comprises: selecting a target cloud network service unit in the second cloud network service unit as a fault proxy node; wherein the target cloud network service unit starts an escape container, and the escape container is elected as a master cloud network control unit as a virtual backup cloud network control unit.

6. A packet forwarding device, comprising: A first cloud network service unit deployed in a distributed decoupling chassis cloud cluster, the device comprising: a selection unit configured to select a fault proxy node from a second cloud network service unit in the distributed decoupling chassis cloud cluster in a case where the first cloud network service unit and a master cloud network control unit manage a network communication fault; wherein the fault proxy node is in normal communication with the master cloud network control unit; and the second cloud network service unit is a cloud network service unit other than the first cloud network service unit in the distributed decoupling chassis cloud cluster; a communication unit configured to, in a case where the fault proxy node is determined, send, to the fault proxy node, a to-be-forwarded management packet through a data network link between the first cloud network service unit and the fault proxy node, and send, by the fault proxy node, the to-be-forwarded management packet to a target device.

7. The apparatus of claim 6, wherein, The selection unit selects the fault proxy node from the second cloud network service unit in the distributed decoupling chassis cloud cluster, comprising: sending a fault proxy request packet to the second cloud network service unit through a data network link between the first cloud network service unit and the second cloud network service unit; in a case where a fault proxy request response packet sent by the second cloud network service unit is received, selecting the fault proxy node from the second cloud network service unit that sends the fault proxy request response packet; wherein the fault proxy request response packet is sent by the second cloud network service unit to the first cloud network service unit through the data network link in a case where the fault proxy request packet is received and it is determined that the node is in normal communication with the master cloud network control unit.

8. The device of claim 6, wherein the communication unit is further configured to send a fault proxy node notification packet to the fault proxy node through the data network link between the first cloud network service unit and the fault proxy node, so that, in a case where the fault proxy node receives the fault proxy node notification packet, the fault proxy node sends the fault proxy node notification packet to other distributed decoupling chassis devices; wherein the fault proxy node notification packet is used to notify the fault proxy node selected by the first cloud network service unit, so that other distributed decoupling chassis devices forward a management packet sent to the first cloud network service unit through the fault proxy node.

9. The apparatus of claim 6, wherein, The device further comprises: an election unit configured to, in a case where an escape container is started on the first cloud network service unit and a current master cloud network control unit manages a network fault, perform new master cloud network control unit election with the escape container as a virtual backup cloud network control unit; wherein a priority of the escape container to be elected as the master cloud network control unit is lower than a priority of a cloud network control unit to be elected as the master cloud network control unit.

10. The apparatus of claim 6, wherein, The selection unit selects the fault proxy node from the second cloud network service unit in the distributed decoupling chassis cloud cluster, comprising: Select a target cloud network service unit in the second cloud network service unit as a fault proxy node; wherein the target cloud network service unit starts an escape container, and the escape container is elected as a main cloud network control unit as a virtual backup cloud network control unit.

11. A distributed decoupled machine frame cloud cluster, characterized in that, Comprise: A cloud network control unit, a cloud network switching unit, a cloud network service unit, and a management device; The cloud network control unit, the cloud network switching unit, and the cloud network service unit realize management network interaction through the management device; the cloud network service unit realizes data network interaction through the cloud network switching unit; wherein: The cloud network service unit is configured to, as a first cloud network service unit, and in the case of a management network communication fault between the first cloud network service unit and a main cloud network control unit, select a fault proxy node from a second cloud network service unit in the distributed decoupling chassis cloud cluster; wherein the fault proxy node is in normal communication with the main cloud network control unit; the second cloud network service unit is other cloud network service units in the distributed decoupling chassis cloud cluster except the first cloud network service unit; for a to-be-forwarded management packet, the to-be-forwarded management packet is sent to the fault proxy node through a data network link between the first cloud network service unit and the fault proxy node, and the fault proxy node sends the to-be-forwarded management packet to a destination device.

12. The distributed decoupler chassis cloud cluster of claim 11, wherein, For target ports on the cloud network switching unit for connecting cloud network service units, message forwarding channels between the target ports are pre-established; For any target port, in the case of detecting that the target port is connected to a cloud network service unit, the state of the target port is controlled to be an open state.

Citation Information

Patent Citations

  • A message forwarding method, a switch, and a computer-readable storage medium

    CN109088819A

  • VM availability during management and VM network failures in host computing systems

    US20150278041A1