Network address assignment failover

By introducing a high availability mechanism into the controller cluster, the DHCP server failure caused by communication interruption is solved, the uniqueness of IP address allocation and normal operation of client devices are ensured, and the reliability of network address allocation is improved.

CN120528893APending Publication Date: 2025-08-22HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510100534.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-22
Filing Date
2025-01-22
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

In the controller cluster, communication interruption causes the DHCP server to be unable to determine the cause of the failure, causing the client device to be unable to obtain the IP address or multiple IP addresses, causing an error operation.

Method used

By introducing a high availability (HA) mechanism into the controller cluster, after detecting communication interruptions, it is determined whether the partner controller is still providing services. If so, it will not take over; if services are no longer provided, it will take over to ensure the uniqueness and correctness of IP address allocation.

Benefits of technology

It reduces the possibility of incorrect operations on client devices, improves the reliability and stability of network address allocation services, and avoids the problem of redundant IP address allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120528893A_ABST
    Figure CN120528893A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to network address allocation failover. In some examples, a first controller of a controller cluster provides a first controller network address assignment service to a first set of client devices. The first controller detects a communication interruption with a second controller of the controller cluster, and determines whether a second controller network address assignment service of the second controller is unavailable based on a determination of whether network address assignment is performed at the second controller within a specified recent time interval. Based on determining that a second controller network address assignment service of the second controller is unavailable, the first controller transitions to a partner off state as part of a network address assignment failover, wherein the first controller provides a first controller network address assignment service to a first set of client devices and provides the first controller network address assignment service to a second set of client devices associated with the second controller.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] A computing environment, such as a data center, a cloud environment, or other type of computing environment, can provide services to client devices. The client devices can access the computing environment through a network. BRIEF DESCRIPTION OF THE DRAWINGS

[0002] Some implementations of the present disclosure are described with respect to the following figures.

[0003] Figure 1 is a block diagram of a network arrangement according to some examples, the network arrangement including a controller cluster, a computing environment, a management system, and client devices.

[0004] Figure 2A and Figure 2B Depicted is a flow diagram of a process for handling communication disruptions between controllers of a controller cluster, according to some examples.

[0005] Figure 3 is a flow chart of a process for handling communication disruptions between controllers of a controller cluster according to other examples.

[0006] Figure 4 is a block diagram of a controller according to some examples.

[0007] Figure 5 is a block diagram of a system according to some examples.

[0008] Figure 6 is a flowchart of a process according to some examples.

[0009] Throughout the drawings, the same reference numerals denote similar, but not necessarily identical, elements. The drawings are not necessarily to scale, and the dimensions of some parts may be exaggerated to more clearly illustrate the examples shown. In addition, the drawings provide examples and / or implementations consistent with this specification; however, this specification is not limited to the examples and / or implementations provided in the drawings. DETAILED DESCRIPTION

[0010] The client device can communicate data with the computing environment through network devices of the network arrangement. Examples of network devices include switches, wireless access points, gateways, concentrators, or other network devices. In some examples, the network arrangement may include a controller having a Dynamic Host Configuration Protocol (DHCP) server that is configured to assign an Internet Protocol (IP) address to the client device. For example, the DHCP server can receive a DHCP message from the client device, wherein the DHCP message includes a Media Access Control (MAC) address of the client device. Based on the MAC address, the DHCP server sends a response DHCP message to the client device that includes an IP address associated with the MAC address.

[0011] Multiple controllers with corresponding DHCP servers can be included in a controller cluster for redundancy to prevent failures. For example, a two-node controller cluster can include two controllers with their corresponding DHCP servers. The DHCP servers of the controller cluster can include a primary DHCP server and a secondary DHCP server. Both the primary DHCP server and the secondary DHCP server can actively provide DHCP services, with the primary DHCP server serving a first set of client devices and the secondary DHCP server serving a second set of client devices.

[0012] Multiple DHCP servers in the controller cluster form a DHCP high availability (DHCP-HA) arrangement. If a first DHCP server in the DHCP-HA arrangement (serving a first set of client devices) fails, a second DHCP server in the DHCP-HA arrangement can take over and provide DHCP services to both the first and second sets of client devices.

[0013] In some cases, communication between the primary and secondary DHCP servers may be interrupted. This interruption in communication may be due to a failure in the communication path between the primary and secondary DHCP servers. However, even if communication between the primary and secondary DHCP servers is interrupted, the primary and secondary DHCP servers may continue to be active (i.e., they can continue to provide DHCP services). In other cases, the interruption in communication may be caused by a failure in one of the primary and secondary DHCP servers; in this latter case, the failed DHCP server may no longer be able to provide DHCP services.

[0014] In response to detecting a communication interruption between the primary and secondary DHCP servers, the first DHCP server in the DHCP-HA arrangement (referred to as "DHCP server A") may be unable to determine whether the communication interruption was caused by a communication path failure or a failure of a partner DHCP server in the DHCP-HA arrangement (referred to as "DHCP server B"). In some cases, even after a communication interruption has occurred, DHCP server A may simply continue to serve its set of client devices ("Set A"), while not serving the other set of client devices associated with DHCP server B ("Set B"). In this case, if the communication interruption was caused by a failed DHCP server B that is no longer able to provide DHCP service, the client devices of Set B associated with DHCP server B will not be able to obtain an IP address assignment (referred to as a "lease") from DHCP server B. If the client devices of Set B are unable to obtain an IP address from DHCP server B, the client devices will not be able to operate over the network.

[0015] On the other hand, if DHCP server A takes over providing DHCP services to client devices of set B in response to detecting a disruption in communication between DHCP servers A and B (and if DHCP server B takes over providing DHCP services to client devices of set A), then if both DHCP servers remain active, the client device may receive multiple IP address assignments (potentially including different IP addresses) from DHCP servers A and B. A client device receiving multiple IP address assignments from different DHCP servers in response to the same DHCP message may result in erroneous operation at the client device because the client device may interpret the receipt of multiple IP address assignments as an error condition.

[0016] According to some implementations of the present disclosure, a technique or mechanism is provided in a cluster of controllers ("controller cluster") that provides a network address allocation service to determine an action to take in response to detecting a communication disruption between controllers of the controller cluster. Typically, a network address allocation service allocates a network address to a client device in response to a request from a client device. An example of a network address allocation service is a DHCP service that allocates an IP address to a client device in response to a DHCP message from the client device. The communication disruption can occur in various scenarios, including, for example: (1) a first scenario in which a failure in a communication path prevents communication between controllers of the controller cluster, but the controller is still actively providing the network address allocation service, and (2) a second scenario in which a controller of the controller cluster has experienced a failure that prevents the controller from providing its network address allocation service. According to some implementations of the present disclosure, a given controller of the controller cluster may perform the following actions: (A) not take over the network address allocation service of a partner controller of the controller cluster in response to the communication disruption if the partner controller is still actively providing the partner controller's network address allocation service, and (B) take over the network address allocation service of the partner controller in response to the communication disruption if the partner controller is no longer actively providing the partner controller's network address allocation service. Action (A) avoids the situation where multiple controllers of a controller cluster provide redundant network address allocation in response to requests from client devices. Action (B) allows a functional controller of a controller cluster to provide network address allocation services to client devices that are part of the set of client devices served by the failed partner controller.

[0017] Techniques or mechanisms according to some implementations may improve related techniques and computer functions associated with providing network address allocation services using a controller cluster by handling communication interruptions and controller failures in a manner that reduces the likelihood of redundant network address allocations that can cause errors at client devices or in network communications, and reduces the likelihood of not providing network address allocation services to client devices when a controller in the controller cluster experiences a failure.

[0018] Examples of client devices may include any or some combination of the following: computers (e.g., desktop computers, laptop computers, server computers, tablet computers, etc.), smartphones, gaming appliances, vehicles, home appliances, Internet of Things (IoT) devices, or other types of electronic devices.

[0019] The controller can be implemented using one or more physical computers. In some cases, multiple controllers of a controller cluster can be implemented on a common computer or a collection of common computers. In this latter example, the controller can be executed as a virtual computing entity (e.g., a virtual machine or VM, a container, etc.) in one or more computers.

[0020] Figure 1 is a block diagram of an example network arrangement including a controller cluster 102, which includes a controller 102-1 and a controller 102-2. Figure 1 The example shows a controller cluster 102 having a pair of controllers, but in other examples, a controller cluster may include more than two controllers.

[0021] Each controller 102-1 and 102-2 includes a corresponding network address allocation server to allocate network addresses to corresponding client devices. In some examples, the network address allocation server includes a DHCP server 104-1 in controller 102-1 and a DHCP server 104-2 in controller 102-2.

[0022] Each controller 102-1 and 102-2 also includes a high availability engine to manage failures associated with controller cluster 102. A failure in controller cluster 102 may include a failure in one or more communication paths between controllers 102-1 and 102-2. Another failure in controller cluster 102 may include a failure in a controller of controller cluster 102. A failure may refer to a hardware failure or error, a failure or error in a machine-readable instruction, or any other condition that causes a component of controller cluster 102 to not function according to a target function. In some examples, the high availability engine in controller 102-1 is DHCP-HA engine 106-1, while the high availability engine in controller 102-2 is DHCP-HA engine 106-2.

[0023] Controller 102-1 also includes a branch gateway 108-1, and controller 102-2 also includes a branch gateway 108-2. Branch gateways 108-1 and 108-2 are discussed further below. DHCP server 104-1, DHCP-HA engine 106-1, and branch gateway 108-1 may be implemented using hardware processing circuitry of controller 102-1, or alternatively, using machine-readable instructions executable by processing resources of controller 102-1. Similarly, DHCP server 104-2, DHCP-HA engine 106-2, and branch gateway 108-2 may be implemented using hardware processing circuitry of controller 102-2, or alternatively, using machine-readable instructions executable by processing resources of controller 102-2.

[0024] Each DHCP server 104-1 or 104-2 is capable of assigning IP addresses to corresponding client devices. When both controllers 102-1 and 102-2 are operational (i.e., the controllers operate according to their intended functions), DHCP server 104-1 is responsible for assigning IP addresses to client devices (CDs) 110-1 in a first set of client devices, and DHCP server 104-2 is responsible for assigning IP addresses to client devices 110-2 in a second set of client devices.

[0025] Client devices 110-1 and 110-2 are capable of communicating in a wireless network 112, which includes access points (APs) 114-A, 114-B, and 114-C, with which corresponding wireless client devices can establish wireless connections. In other examples, wireless network 112 may include a different number of APs, which may be fewer than three APs or more than three APs. A client device may be associated with an AP, which refers to the client device establishing a wireless connection with the AP so that the client device can use the communication resources of wireless network 112 to communicate. Wireless network 112 may include a wireless local area network (WLAN), a cellular network, or any other type of wireless network. APs 114-A, 114-B, and 114-C may operate in accordance with the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard (also known as the Wi-Fi standard). Alternatively, APs 114-A, 114-B, and 114-C may operate in accordance with a cellular protocol or other wireless technology.

[0026] although Figure 1 An example is shown with a wireless network including an AP, but it should be noted that in other examples, client devices 110-1 and 110-2 may include wired devices connected to a wired network. Techniques or mechanisms according to some examples of the present disclosure may be applied to arrangements including wired client devices.

[0027] APs 114-A, 114-B, and 114-C are connected to a switch infrastructure 116 that includes one or more switches. A switch is a network device that forwards data packets along a communication path based on the MAC address in the data packet. A "data packet" (also known as a "data frame," "data unit," etc.) can refer to any unit of data that can be individually transmitted in a communication path.

[0028] Switch infrastructure 116 connects APs 114-A, 114-B, and 114-C to controller cluster 102. Figure 1In the example of FIG, controller 102-1 may be connected to port 118-1 of switch infrastructure 116, and controller 102-2 may be connected to port 118-2 of switch infrastructure 116. Ports 118-1 and 118-2 may be part of the same switch, or may be part of different switches of switch infrastructure 116. A "port" may refer to any communication interface through which an electronic device can communicate.

[0029] A client device (110-1 or 110-2) associated with an AP can access a controller in controller cluster 102 to obtain a network address, such as an IP address assigned by DHCP server 104-1 or 104-2. In some examples, the client device can broadcast a client request (e.g., a DHCP message) to controllers 102-1 and 102-2 of controller cluster 102. The broadcast message is intended to be received by multiple target entities, which in this case include DHCP servers 104-1 and 104-2.

[0030] In response to a client request, a designated DHCP server in controller cluster 102 allocates an IP address to the client device. The allocation of an IP address by a DHCP server is also called a lease, which refers to the temporary allocation of an IP address. The allocated IP address is sent to the client device by the designated DHCP server.

[0031] In some examples, the DHCP server designated to allocate an IP address in response to a client request is a DHCP server selected based on the MAC address in the client request. The MAC address in the client request is the MAC address of the client device that issued the client request. The DHCP server selected from the DHCP servers in controller cluster 102 can be based on a value derived from the MAC address in the client request. For example, a hash function can be applied to the MAC address in the client request to generate a hash value. If the hash value falls within a first range of hash values, controller 102-1 is selected to respond to the client request. On the other hand, if the hash value falls within a second range of hash values, controller 102-2 is selected to respond to the client request. In other examples, in addition to applying a hash function to the MAC address in the client request, a different function can be applied to the MAC address in the client request (or any other parameter value) to generate a selection value used to select one of the multiple controllers in controller cluster 102. The DHCP server in the selected controller is used to respond to the client request.

[0032] One or more inter-controller communication paths may be established between the controllers of the controller cluster 102. Figure 1In the example shown, two inter-controller communication paths are established between controllers 102-1 and 102-2. In some examples, the inter-controller communication paths include different tunnels between controllers 102-1 and 102-2, where the tunnels are established through switch infrastructure 116. A tunnel refers to a communication path in which data packets are encapsulated within packets of another protocol.

[0033] Examples of tunnels include Internet Protocol Security (IPsec) tunnels and Generic Routing Encapsulation (GRE) tunnels. An IPsec tunnel refers to a tunnel in which data packets are encrypted and encapsulated according to the IPsec protocol, as described in various Request for Comments (RFCs), including RFC 2401, entitled "Security Architecture for the Internet Protocol," dated November 1998. An IPsec tunnel may perform encapsulation according to the Encapsulating Security Payload (ESP) protocol, as described in RFC 4303, entitled "IP Encapsulating Security Payload (ESP)," dated December 2005.

[0034] GRE tunnels allow the source and destination endpoints of the GRE tunnel to operate as if the source and destination endpoints (in this case, controllers 102-1 and 102-2) have a virtual point-to-point connection with each other. GRE is described in RFC 2784, entitled "Generic Routing Encapsulation (GRE)" dated March 2000.

[0035] Although reference is made to IPsec tunnels and GRE tunnels, it should be noted that other types of communication paths may be employed in other examples. In the following discussion, the first inter-controller communication path between controllers 102-1 and 102-2 is referred to as cluster tunnel 120, which may be implemented as an IPsec tunnel. The second inter-controller communication path between controllers 102-1 and 102-2 is referred to as peer tunnel 122, which may be implemented as a GRE tunnel.

[0036] In some examples, cluster tunnel 120 and peer tunnel 122 are examples of downstream communication paths between controllers 102-1 and 102-2 because cluster tunnel 120 and peer tunnel 122 are established through switch infrastructure 116, which is downstream of controller cluster 102. A network entity, such as switch infrastructure 116, is downstream of controller cluster 102 if the network entity is closer to client devices using services of controller cluster 102 than another entity.

[0037] In contrast, entities upstream of controller cluster 102 include a computing environment 130 that provides various services accessible by client devices 110-1 and 110-2. An example of computing environment 130 is a data center. In other examples, computing environment 130 may include a cloud computing environment or any other type of computing environment.

[0038] The computing environment 130 can be accessed through a network 132, such as the Internet or any other type of network. Controller 102-1 can access the network 132 through a router infrastructure 134-1 including one or more routers, and controller 102-2 can access the network 132 through a router infrastructure 134-2 including one or more routers. In other examples, controllers 102-1 and 102-2 can access the network 132 through the same router infrastructure.

[0039] In addition to the DHCP servers, the controllers 102-1 and 102-2 include corresponding branch gateways. For example, the controller 102-1 includes the branch gateway 108-1, and the controller 102-2 includes the branch gateway 108-2. The branch gateway is a type of edge router that can be coupled to a computing environment 130, such as via multiple network paths, through the router infrastructure 134-1, 134-2. The network path from the branch gateway to the computing environment 130 is called an "uplink path". The branch gateway can select one of the uplink paths to use for any client device that is accessing the services 131 of the computing environment 130. The branch gateways 108-1 and 108-2 can be part of a software-defined wide area network (SD-WAN) for managing which uplink paths to use for data communications between client devices and the computing environment 130.

[0040] In some examples, computing environment 130 includes a virtual private network (VPN) concentrator (VPNC) 136. An uplink path is connected to VPNC 136. A VPNC refers to a network device that manages multiple VPN connections from client devices to computing environment 130. A VPN comprises a secure connection between a client device and computing environment 130. Although reference is made to the use of VPNs and VPNCs in accordance with some examples, it should be noted that in other examples, other types of connections can be established between a client device and computing environment 130.

[0041] Figure 1 Also depicted is a management system 138 connected to the network 132. The management system 138 may reside in a cloud computing environment, or more generally, may refer to any management system that is remote from the controller cluster 102 and that may be used to manage the controller cluster 102. The management system 138 may be implemented using one or more computers.

[0042] DHCP servers use a lease database to record IP addresses assigned in response to client requests from client devices. DHCP server 104-1 uses lease database 105-1, and DHCP server 104-2 uses lease database 105-2. The lease database comprises a repository of leases (IP addresses) that have been assigned by the DHCP servers. The lease database can be in the form of a file or any other type of data structure stored in the memory of the corresponding controller 102-1 or 102-2. A DHCP lease is a temporary assignment of an IP address to a client device.

[0043] Each DHCP server has a pool of IP addresses from which the DHCP server can assign IP addresses to client devices. DHCP server 104-1 can assign IP addresses from a first IP address pool, and DHCP server 104-1 can assign IP addresses from a second IP address pool, where the first pool and the second pool are different from each other.

[0044] Lease databases 105-1 and 105-2 are synchronized with each other. In other words, lease database 105-1 includes the IP address assigned to client device 110-1 by DHCP server 104-1 and the IP address assigned to client device 110-2 by DHCP server 104-2, and similarly, lease database 105-2 includes the IP address assigned to client device 110-2 by DHCP server 104-2 and the IP address assigned to client device 110-1 by DHCP server 104-1.

[0045] Multiple communication paths (including cluster tunnel 120 and peer tunnel 122) can be used for different purposes. In some examples, controllers 102-1 and 102-2 can use cluster tunnel 120 (e.g., including an IPsec tunnel) to securely exchange information with each other, such as for synchronizing lease databases 105-1 and 105-2. Peer tunnels (e.g., including GRE tunnels) can be used to support high availability features of controller cluster 102, causing a first branch gateway of a first controller of controller cluster 102 to use an uplink path of a partner controller's branch gateway in the event that the first controller's uplink path fails or is overloaded.

[0046] In some examples, controllers 102-1 and 102-2 can use a heartbeat mechanism to detect a communication interruption between controllers 102-1 and 102-2. The heartbeat mechanism relies on the use of heartbeat messages exchanged between controllers 102-1 and 102-2. A "heartbeat message" can refer to a packet, information element, signal, or any other type of indicator sent from a first entity to a second entity to indicate that the first entity is active. DHCP-HA engine 106-1 in controller 102-1 can periodically send heartbeat messages to controller 102-2, and similarly, DHCP-HA engine 106-2 in controller 102-2 can periodically send heartbeat messages to controller 102-1.

[0047] In some examples, heartbeat messages may be exchanged between controllers 102-1 and 102-2 in each of multiple communication paths between the controllers, including cluster tunnel 120 and peer tunnel 122. In other examples, heartbeat messages are exchanged over only one communication path between controllers 102-1 and 102-2.

[0048] If the DHCP-HA engine in a given controller fails to receive a heartbeat message from a partner controller of controller cluster 102 within a specified time interval (referred to as a "heartbeat interval," which may be set by an administrator or another entity such as a program or machine), the given controller of controller cluster 102 enters a communication-disconnected state. For example, the DHCP-HA engine periodically sends heartbeat messages at each heartbeat interval. If the given controller does not receive a heartbeat message from a partner controller within the heartbeat interval, the given controller may indicate that a communication disruption has occurred after the heartbeat interval expires. The failure to receive heartbeat messages results in a potential partition of controller cluster 102, wherein controllers 102-1 and 102-2 become isolated from each other due to the communication disruption between controllers 102-1 and 102-2.

[0049] According to some implementations of the present disclosure, a given controller of controller cluster 102 may perform the following actions: (A) not take over the DHCP service of the partner controller of the controller cluster in response to a communication disruption if the partner controller is still actively providing the partner controller's DHCP service, and (B) take over the DHCP service of the partner controller in response to a communication disruption if the partner controller is no longer actively providing the partner controller's DHCP service. To support actions (A) and (B), each of DHCP-HA engines 106-1 and 106-2 may perform a partner activity confirmation procedure in response to detecting a communication disruption between controllers 102-1 and 102-2 to check whether the partner controller is still able to provide the partner's DHCP service.

[0050] Figure 2A and Figure 2B A flow chart depicts a process for handling a communication interruption between controllers 102-1 and 102-2, wherein DHCP failover may be performed if one of DHCP servers 104-1 and 104-2 fails. Figure 2A and Figure 2B A particular order of tasks is depicted, but in other examples, a different order of tasks can be used, some tasks can be omitted, and other tasks can be added.

[0051] In the following discussion, reference is made to the various DHCP-HA states described in "DHCP Failover Protocol" by Ralph Droms et al., draft-ietf-dhc-failover-12.txt, March 2003. Controller 102-1 or 102-2 may enter any DHCP-HA state. Although some example DHCP states are referenced, in other examples, controllers 102-1 and 102-2 may transition between other states.

[0052] DHCP-HA states include normal, communication interrupted, partner down, recovering, and down.

[0053] The normal state refers to a state in which controllers 102-1 and 102-2 are synchronized with each other, or more specifically, lease databases 105-1 and 105-2 are synchronized with each other. Each synchronized lease database includes leases assigned by the DHCP server associated with the lease database, as well as leases assigned by partner DHCP servers. Other DHCP-HA states are discussed further below.

[0054] Assuming that controller 102-1 is in a normal state, DHCP-HA engine 106-1 in controller 102-1 detects (at 202) a communication interruption in cluster tunnel 120 between controllers 102-1 and 102-2. This detection is based on DHCP-HA engine 106-1 failing to receive a heartbeat message through cluster tunnel 120 within the heartbeat interval.

[0055] In response to detecting a communication disruption through cluster tunnel 120, DHCP-HA engine 106-1 checks (at 204) whether a communication disruption between controllers 102-1 and 102-2 also exists in peer tunnel 122. If DHCP-HA engine 106-1 fails to receive a heartbeat message within the heartbeat interval of peer tunnel 122, a communication disruption in peer tunnel 122 is indicated.

[0056] If DHCP-HA engine 106-1 determines (at 204) that there is no communication disruption in peer tunnel 122 (e.g., DHCP-HA engine 106-1 receives a heartbeat message in peer tunnel 122 within the heartbeat interval of peer tunnel 122), this is an indication that partner controller 102-2 is active. The ability to receive heartbeat messages through peer tunnel 122 indicates that partner DHCP server 104-2 is active, but communication through cluster tunnel 120 is experiencing some issues (e.g., the TCP connection may be having issues in cluster tunnel 120). In this case, DHCP-HA engine 106-1 may issue an alarm indicating a problem with cluster tunnel 120, and controller 102-1 may perform tasks (at 206) to resolve the failed cluster tunnel 120.

[0057] However, if DHCP-HA engine 106-1 determines (at 204) that there is a communication interruption in peer tunnel 122, DHCP-HA engine 106-1 transitions controller 102-1 (at 208) to a communication interruption state. In the communication interruption state, DHCP server 104-1 can assign IP addresses from the first IP address pool. However, DHCP server 104-1 in the communication interruption state cannot assign IP addresses from the second IP address pool because it is possible that DHCP server 104-2 in partner controller 102-2 may still be active even if there is a communication interruption in cluster tunnel 120 and peer tunnel 122. When the first IP address pool is exhausted, i.e., DHCP server 104-1 has assigned all IP addresses in the first IP address pool to client devices, DHCP server 104-1 will no longer be able to assign any additional IP addresses.

[0058] After entering the communication-disabled state, DHCP-HA engine 106-1 continues to check whether DHCP server 104-2 in partner controller 102-2 is active. It is possible that connectivity between controllers 102-1 and 102-2 through switch infrastructure 116 is down, but controllers 102-1 and 102-2 may still have connectivity to APs 114-A through 114-C through switch infrastructure 116.

[0059] Even if there is a communication disruption in the cluster tunnel 120 and the peer tunnel 122, the controller 102-1 still has another communication path in the upstream direction through the computing environment 130. The branch gateway 108-1 in the controller 102-1 has an uplink path to the computing environment 130, and the branch gateway 108-2 in the controller 102-2 has an uplink path to the computing environment 130.

[0060] DHCP-HA engine 106-1 may send (at 210) a health query via an uplink path through branch gateway 108-1. A tunnel (referred to as an "uplink tunnel") may be established via the uplink path between controller 106-1 and VPNC 136 of computing environment 130. The uplink tunnel may be, for example, an IPsec tunnel or another type of tunnel. The health query is sent via the uplink tunnel to VPNC 136, which forwards (at 212) the health query via another tunnel (such as an uplink tunnel, such as an IPsec tunnel) between VPNC 136 and partner controller 106-2.

[0061] The health query includes a source network address that identifies controller 102-1 as the source and a destination network address that identifies controller 102-2 as the destination. The health query also includes a timestamp recorded by DHCP-HA engine 106-1, which indicates the time when the health query was created or sent. The health query is a request to partner controller 102-2 to indicate whether the DHCP server 104-2 of partner controller 102-2 is active.

[0062] In response to receiving the health query, partner controller 102-2 checks (at 214) whether there has been recent lease activity at controller 102-2. For example, controller 102-2 may check lease database 105-2 to determine whether lease activity has occurred during a specified lease activity duration relative to the timestamp in the health query. The lease activity may include an update of a lease or an allocation of a new lease. The specified lease activity duration may include a first time interval before the timestamp, a second time interval after the timestamp, or both the first time interval and the second time interval. The length of the specified lease activity duration may be set by an administrator or by another entity.

[0063] The lease database includes entries that include the IP addresses assigned to corresponding MAC addresses and a timestamp indicating when the IP addresses were assigned. A lease is a temporary IP address assignment, where the assigned IP address expires after a set expiration time. Lease renewal occurs when the DHCP server renews the assignment of an IP address and resets the expiration time for the IP address. When a lease is renewed, the timestamp in the entry corresponding to the updated lease in the lease database is updated to reflect the time of the update. When a new lease is assigned, a new entry is added to the lease database, which includes the assigned IP address and a timestamp indicating the new lease time.

[0064] If controller 102-2 determines that there is lease activity during the specified duration relative to the timestamp in the health query, controller 102-2 creates (at 216) a health query response (in response to the health query from controller 102-1) indicating that DHCP server 104-2 is active. For example, partner controller 102-2 may set an indicator (I) (e.g., a field, an information element, etc.) in the health query response to a first value (e.g., "1") to indicate that DHCP server 104-2 is active.

[0065] If partner controller 102-2 determines that there has been no lease activity during the specified duration relative to the timestamp in the health query, partner controller 102-2 creates (at 218) a health query response indicating that DHCP server 104-2 is unavailable. For example, partner controller 102-2 may set an indicator (I) in the health query response to a second value (e.g., “0”) (different from the first value) to indicate that DHCP server 104-2 is unavailable.

[0066] Partner controller 102-2 sends a health query response with an indicator (set to the first value or the second value) to VPNC 136 (at 220), and VPNC 136 forwards the health query response to controller 102-1 (at 222). In response to receiving the health query response, controller 102-1 checks (at 224) the value of the indicator (I) in the health query response. If the indicator (I) is set to the first value (e.g., "1") indicating that partner controller 102-2's DHCP server 104-2 is active, DHCP-HA engine 106-1 maintains controller 102-1 in a communication interrupted state (at 226).

[0067] Note that although controller 102-1 is described above as detecting a communication disruption in the inter-controller communication path and sending a health query to partner controller 102-2, partner controller 102-2 will also perform similar tasks in response to detecting a communication disruption. If both DHCP servers 104-1 and 104-2 are active, and cluster tunnel 120 and peer tunnel 122 are disrupted, both controllers 102-1 and 102-2 are operating in a communication disruption state, meaning that each DHCP server 104-1 or 104-2 allocates IP addresses from its own IP address pool, rather than from the IP address pool associated with the DHCP server in the partner controller.

[0068] If the indicator (I) is set to a second value (e.g., "0") indicating that the DHCP server 104-2 of the partner controller 102-2 is unavailable, the DHCP-HA engine 106-1 sends (at 228) a service request to the partner controller 102-2 to cause the partner controller 102-2 to transition to the down state. The service request is sent to the VPNC 136, which forwards (at 230) the service request to the partner controller 102-2.

[0069] Upon transitioning to the Off state in response to the service request, controller 102-2 sends (at 232) a service response to confirm that controller 102-2 has transitioned to the Off state. The service response is sent from controller 102-2 to VPNC 136, which forwards (at 234) the service response to controller 102-1.

[0070] In response to the service response from controller 102-2, DHCP-HA engine 106-1 transitions controller 102-1 to the Partner Down state (at 236). In the Partner Down state, DHCP server 104-1 in controller 102-1 is able to allocate IP addresses from its first IP address pool and the second IP address pool associated with partner controller 102-2. Transitioning controller 102-1 to the Partner Down state allows DHCP server 104-1 to perform a DHCP service failover, wherein DHCP server 104-1 assumes the responsibilities of both DHCP servers 104-1 and 104-2 during the period when DHCP server 104-2 is unavailable.

[0071] In the buddy-off state, DHCP server 104-1 can serve client devices having MAC addresses that map to both hash ranges of DHCP servers 104-1 and 104-2. In other words, DHCP server 104-1 can assign IP addresses to client devices that would normally be served by DHCP server 104-2.

[0072] As part of transitioning to the Buddy Down state, DHCP-HA engine 106-1 in controller 102-1 may send (at 238) a Buddy Down Indication message. The Buddy Down Indication message is sent to VPNC 136, which forwards (at 240) the Buddy Down Indication message to controller 102-2. The Buddy Down Indication message indicates to controller 102-2 that controller 102-1 has transitioned to the Buddy Down state. In response to the Buddy Down Indication message, controller 102-2 may store (at 242) in a memory of controller 102-2 an indication that controller 102-2 will begin in the Resume state when controller 102-2 resumes from the Shutdown state.

[0073] At some later time, communication is resumed over cluster tunnel 120 and peer tunnel 122. For example, heartbeat messages may again be received by respective DHCP-HA engines 106-1 and 106-2 via cluster tunnel 120 and peer tunnel 122. In response to detecting (at 244) the resumption of communication over cluster tunnel 120 and peer tunnel 122, and after some specified amount of time (e.g., a Maximum Client Delivery Time, or MCLT) from detecting the resumption of communication, DHCP-HA engine 106-1 transitions (at 246) controller 102-1 from the Communication Disconnected state (set at 226) or the Partner Down state (set at 236) to the Normal state. The MCLT is a configured maximum amount of time that a DHCP server can extend a lease for a client device beyond the time known by the partner DHCP server. Effectively, the MCLT sets the maximum amount of time that an IP address assigned by DHCP server 104-1 to a client device (that would otherwise have been served by partner DHCP server 104-1) is valid.

[0074] In response to detecting (at 248) the resumption of communications of the cluster tunnel 120 and the peer tunnel 122, if the controller 102-2 was in the shutdown state, the DHCP-HA engine 106-2 in the partner controller 102-2 transitions (at 250) the controller 102-2 from the shutdown state to the resumed state. As described above, the controller 102-2 may store (at 240) in a memory of the controller 102-2 an indication that the controller 102-2 is to begin in the resumed state when the controller 102-2 is resumed from the shutdown state.

[0075] In the recovery state, DHCP server 104-2 recovers (at 252) lease information from lease database 105-1 of controller 102-1, such as through cluster tunnel 120. After recovering the lease information, controller 102-2 may transition (at 254) to the normal state.

[0076] Note that if DHCP server 104-2 was active during the communication interruption between controllers 102-1 and 102-2, controller 102-2 would have transitioned to the Communication Interrupted state (and would not have transitioned to the Off state based on tasks 224, 228, and 230). In this alternative scenario, controller 102-2 would transition from the Communication Interrupted state to the Normal state, such as after a specified amount of time (e.g., MCLT).

[0077] In task 222, assume that controller 102-1 receives a health query response from controller 102-2. In other examples, it is possible that controller 102-1 does not receive any health query responses, which may be due to controller 102-2 being inoperable. If controller 102-1 does not receive a health query response in response to one or more health queries (e.g., after a threshold number of retries), controller 102-1 may assume that DHCP server 104-2 of controller 102-2 is out of service, and DHCP server 104-1 in controller 102-1 enters a partner-down state and allocates IP addresses from both the first IP address pool and the second IP address pool.

[0078] Figure 3 1 is a flowchart of a management system controlled process for handling a communication interruption between controllers 102-1 and 102-2, wherein a DHCP service failover may be performed if one of DHCP servers 104-1 and 104-2 fails. The management system controlled process involves Figure 1 management system 138. Although Figure 3 A particular order of tasks is depicted, but in other examples, a different order of tasks can be used, some tasks can be omitted, and other tasks can be added.

[0079] DHCP-HA engines 106-1 and 106-2 of respective controllers 102-1 and 102-2 send status updates to management system 138. DHCP-HA engine 106-1 sends (at 302) a status update to management system 138, and DHCP-HA engine 106-2 sends (at 304) a status update to management system 138. The status update may include information about the current DHCP-HA state of the controller. The status update may be sent in response to a change in the DHCP-HA state at the controller. The status update may also be sent by the DHCP-HA engine when the DHCP server initially starts up.

[0080] Controllers 102-1 and 102-2 also send client device information to management system 138. Controller 102-1 sends (at 306) client device information for client devices served by controller 102-1 to management system 138, and controller 102-2 sends (at 308) client device information for client devices served by controller 102-2 to management system 138. The client device information may include information about new leases or lease renewals, as well as other information. The client device information may be sent periodically by the controllers or may be sent in response to another event.

[0081] Based on the status updates from controllers 102-1 and 102-2, management system 138 determines (at 310) whether controllers 102-1 and 102-2 are in a communication-disconnected state (e.g., because the controllers are not receiving heartbeat messages from each other). If controllers 102-1 and 102-2 are not in a communication-disconnected state, management system 138 returns to receive more status updates from controllers 102-1 and 102-2.

[0082] In response to detecting that controllers 102-1 and 102-2 are in a communication interruption state, management system 138 determines (at 312) whether DHCP server 104-1 or 104-2 is unavailable based on client device information from controllers 102-1 and 102-2. This determination can be based on management system 138 detecting that a given DHCP server has not participated in any lease activity for a specified lease activity duration. The given DHCP server is identified as an unavailable DHCP server.

[0083] If both DHCP servers 104-1 and 104-2 are unavailable, management system 138 takes no further action and returns to receive further status updates from controllers 102-1 and 102-2. However, if management system 138 identifies one of the DHCP servers as unavailable, management system 138 sends (at 314) a shutdown message to the controller (e.g., controller 102-2) that includes the unavailable DHCP server to cause the controller to transition to the shutdown state.

[0084] In response to transitioning to the shutdown state, controller 102-2 sends (at 316) a shutdown transition message to controller 102-1 (such as via an uplink path of VPNC 136) to notify controller 102-1 that controller 102-2 has entered the shutdown state. Alternatively, instead of or in addition to controller 102-2 sending the shutdown transition message to controller 102-1, management system 138 may send the shutdown transition message to controller 102-1.

[0085] In response to the shutdown transition message, controller 102-1 transitions (at 318) to the Partner Down state. In the Partner Down state, DHCP server 104-1 in controller 102-1 is able to allocate IP addresses from its first IP address pool and the second IP address pool associated with partner controller 102-2. Controller 102-1 transitioning to the Partner Down state allows DHCP server 104-1 to perform a DHCP service failover, wherein DHCP server 104-1 assumes the responsibilities of both DHCP servers 104-1 and 104-2 during the period when DHCP server 104-2 is unavailable.

[0086] As part of transitioning to the Buddy Off state, controller 102-1 may send (at 320) a Buddy Off Indication message to controller 102-2 (such as via an uplink path of VPNC 136). The Buddy Off Indication message indicates to controller 102-2 that controller 102-1 has transitioned to the Buddy Off state. In response to the Buddy Off Indication message, controller 102-2 may store (at 322) in a memory of controller 102-2 an indication that controller 102-2 will begin in a Resume state when controller 102-2 resumes from the Shutdown state.

[0087] When communication between controllers 102 - 1 and 102 - 2 is restored, tasks similar to 244 , 246 , 248 , 250 , 252 , and 254 may be performed to transition back to a normal state.

[0088] Note that controllers 102 - 1 and 102 - 2 send status updates to management system 138 when controllers 102 - 1 and 102 - 2 transition between different states.

[0089] exist Figures 2A to 2B and Figure 3 In the depicted process, an uplink path between controller clusters 102 through computing environment 130 is used by a first controller to verify whether a partner DHCP server of a partner controller is active during a communication outage period in one or more inter-controller communication paths. The ability to verify whether the partner DHCP server is active allows a first DHCP server of the first controller to take appropriate action during the communication outage period, which may include (A) performing a DHCP service failover if the partner DHCP server is unavailable, or (B) continuing to provide DHCP service if the partner DHCP server is active, wherein the first DHCP server allocates leases from its own IP address pool (rather than the IP address pool of the partner DHCP server). In this way, if the partner DHCP server is unavailable, client devices that should be served by the partner DHCP server will not be deprived of DHCP service. On the other hand, if the partner DHCP server is active during the communication outage period, the first DHCP server does not take over the DHCP service of the partner DHCP server, which may result in redundant network address allocations, thereby causing client device errors.

[0090] In another example of the present disclosure, a DHCP-HA engine in a first controller may detect that a partner DHCP server is unavailable based on detecting that an unusually large number of client requests (e.g., DHCP messages) for leases that should have been serviced by a partner controller are received at the first controller. A client device broadcasts a client request for a lease that is received by both controllers 102-1 and 102-2. Each controller that receives the client request generates a hash value based on the MAC address in the client request. If the hash value falls within a first range of hash values, controller 102-1 is selected to respond to the client request. On the other hand, if the hash value falls within a second range of hash values, controller 102-2 is selected to respond to the client request. If controller 102-1 repeatedly receives client requests from the same client device (or the same set of client devices) that should have been serviced by controller 102-2, DHCP-HA engine 106-1 in controller 102-1 may determine that partner DHCP server 104-2 is unavailable, and in response, DHCP-HA engine 106-1 may transition controller 102-1 to a partner-down state to perform a DHCP service failover.

[0091] DHCP-HA engine 106-1 may track the number of client requests from the same client device (or the same set of client devices) that should have been served by controller 102-2 within a specified time interval. The client requests may include client requests for new leases and client requests for lease renewals. If the tracked number exceeds a specified threshold, this is an indication that DHCP service 104-2 is unavailable, and DHCP server 104-1 in controller 102-1 may take over DHCP service from partner DHCP server 104-2.

[0092] A client request for a new lease is broadcast by the client device for receipt by multiple controllers of controller cluster 102. A client request for a lease update is unicast by the client device to the controller that previously issued the new lease to the client device. However, if the client device does not receive a response to the client request for a lease update for a threshold number of retries, the client device may broadcast a client request for a lease update for receipt by multiple controllers of controller cluster 102.

[0093] Figure 4 is a block diagram of a first controller 400 of a controller cluster. The first controller 400 may be, for example Figure 1Controller 102-1 or 102-2. The controller cluster may include multiple (two or more) controllers. The first controller 400 includes a hardware processor 402 (or multiple hardware processors). The hardware processor may include a microprocessor, a core of a multi-core microprocessor, a microcontroller, a programmable integrated circuit, a programmable gate array, or another hardware processing circuit.

[0094] The first controller 400 includes a non-transitory machine-readable or computer-readable storage medium 404 storing machine-readable instructions executable on a hardware processor 402 to perform various tasks. The machine-readable instructions executable on a hardware processor may refer to instructions executable on a single hardware processor or instructions executable on multiple hardware processors.

[0095] The machine-readable instructions include network address allocation instructions 406 for providing a first controller network address allocation service to client devices in the first set of client devices (such as based on characteristics of the client devices in the first set of client devices). For example, the characteristics may include a MAC address of the client device. In other examples, other characteristics of the client device may be used to determine whether the client device is to be served by the first controller 400 or by another controller in the controller cluster.

[0096] The machine-readable instructions include communication interruption detection instructions 408 to detect a communication interruption with a second controller of the controller cluster, wherein communication via a first communication path between the first controller and the second controller is interrupted. For example, the first communication path may include Figure 1 In some examples, the detection of the communication interruption is further based on detecting that there is also a communication interruption through a second communication path between the first controller and the second controller. For example, the second communication path may include Figure 1 peer tunnel 122.

[0097] The machine-readable instructions include second controller network address allocation service availability determination instructions 410 to determine whether a second controller network address allocation service of a second controller is unavailable. The availability determination is based on determining whether network address allocation at the second controller was performed within a specified recent time interval. The availability determination may be performed relative to a health query (e.g., a query sent from a first controller to a second controller) sent to the second controller. Figure 2A The specified most recent time interval may be measured using a timestamp of 210, 212 in FIG. 1 , or alternatively, the specified most recent time interval may be a time interval during which no leasing activity occurs at the second controller, as indicated by Figure 1 The same as monitored by the management system 138 (e.g., Figure 3 312 in ).

[0098] The machine-readable instructions include buddy down transition instructions 412 to transition the first controller to a buddy down state as part of a network address allocation failover based on determining that a second controller network address allocation service of the second controller is unavailable, wherein the first controller provides the first controller network address allocation service to client devices in a first set of client devices and provides the first controller network address allocation service to client devices in a second set of client devices associated with the second controller.

[0099] In some examples, the partner down state is a DHCP-HA partner down state. More generally, the partner down state is a state of a controller in which a controller is about to take over (or has taken over) a second controller network address allocation service of a second controller as part of a network address allocation failover.

[0100] In some examples, the first controller network address allocation service and the second controller network address allocation service are such as provided by Figure 1 DHCP servers 104-1 and 104-2 provide DHCP services.

[0101] In some examples, the determination of whether a network address allocation was performed within a specified recent time interval is based on a query of a database such as a DHCP lease database (e.g., Figure 1 Check of the network address database of 105-1 or 105-2).

[0102] In some examples, the network address database includes a first network address database used by the second controller, and the first controller uses the second network address database, wherein the second network address database is synchronized with the first network address database.

[0103] In some examples, in response to the first controller providing the first controller network address assignment service to a client device in the second set of client devices, the first network address database and the second network address database become out of synchronization. The machine-readable instructions may detect that communication is reestablished between the first controller and the second controller, and based on detecting that communication is reestablished between the first controller and the second controller, initiate synchronization of the first network address database and the second network address database (e.g., Figure 2B 252 in ).

[0104] In some examples, the first controller performs a first controller network address allocation service for a first set of client devices associated with a first subset of MAC addresses, and the second controller performs a second controller network address allocation service for a second set of client devices associated with a second subset of MAC addresses.

[0105] In some examples, the first controller may detect a communication disruption with the second controller based on detecting a failure to receive a first heartbeat indication over a first communication path, such as a heartbeat message over cluster tunnel 120 (eg, an IPsec tunnel).

[0106] In some examples, the first controller may further detect a communication interruption with the second controller based on detecting a failure to receive a second heartbeat indication over a second communication path different from the first communication path. The second communication path may be Figure 1 peer tunnel 122 (eg, a GRE tunnel).

[0107] In some examples, in response to detecting a communication interruption with the second controller, the first controller may send a health query (e.g., Figure 2A 210, 212 in the example), wherein the uplink communication path is different from the first and second communication paths. For example, the uplink communication path includes the first controller and the VPNC in the computing environment (e.g., Figure 1 In response to the health query, the first controller may receive an indication that the second controller network address allocation service of the second controller is unavailable via the uplink communication path (e.g., Figure 2A Determining that the second controller network address allocation service of the second controller is unavailable is based on receiving an indication. The indication is based on the second controller determining that no network address allocation has been performed within a specified recent time interval.

[0108] In some examples, determining whether the network address allocation was performed within a specified recent time interval is relative to a timestamp of the health query.

[0109] In some examples, based on determining that the second controller network address allocation service of the second controller is unavailable, the machine-readable instructions may cause a shutdown indication (e.g., Figure 2A 228, 230) to request the second controller to enter the shutdown state.

[0110] In some examples, the first controller may detect a further interruption in communication with the second controller. In response to detecting the further interruption in communication with the second controller, the first controller may send a health query to the second controller via the second communication path. Based on detecting that the second controller does not respond to the health query after a threshold number of retries, the first controller may determine that the second controller network address allocation service of the second controller is unavailable.

[0111] In some examples, determining whether network address allocation is performed within a specified recent time interval is performed by a management system (e.g., Figure 1 Based on the management system determining that the second controller has not performed network address allocation within a specified recent time interval, the first controller may receive an indication that the management system has caused the second controller to shut down (e.g., Figure 3 316 in ).

[0112] In some examples, the first controller may detect greater than a threshold number of client requests for network address allocation from the same client device or set of client devices within a specified time interval, where the client requests from the same client device or set of client devices should have been serviced by the second controller. Based on the detection of greater than the threshold number of client requests from the same client device or set of client devices that should have been serviced by the second controller, the first controller indicates a disruption in communication between the first controller and the second controller.

[0113] Figure 5 5 is a block diagram of a system 500 that can be implemented using one or more computers. The system 500 includes a controller cluster 502 having a first controller 504 providing a first controller network address allocation service 506 and a second controller 508 providing a second controller network address allocation service 510. The first controller 504 has an inter-controller communication path 512 to the second controller 508.

[0114] The first controller 504 can perform various tasks according to some examples of the present disclosure. The tasks include a communication interruption detection task 514 to detect a communication interruption with the second controller 508 on the inter-controller communication path 512 .

[0115] Tasks include a second controller network address allocation service availability determination task 516 to determine whether the second controller network address allocation service 510 of the second controller 508 is unavailable. The determination is based on determining whether network address allocation was performed at the second controller 508 within a specified recent time interval.

[0116] The tasks include a partner down transition task 518 that transitions the first controller 504 to a partner down state as part of a network address allocation failover based on determining that a second controller network address allocation service 510 of the second controller 508 is unavailable, wherein the first controller 504 provides the first controller network address allocation service 506 to client devices in a first set of client devices associated with the first controller and to client devices in a second set of client devices associated with the second controller.

[0117] In some examples, in response to transitioning to the buddy off state, the first controller 504 sends a buddy off indication (e.g., Figure 2B In response to the partner shutdown indication from the first controller 504, the second controller 508 stores (e.g., Figure 2B 242) indicates that the second controller will enter a recovery state when transitioning from the unavailable state. In the recovery state, the second controller 508 synchronizes a database of assigned network addresses associated with the second controller 508 (e.g., a DHCP lease database) with a database of assigned network addresses associated with the first controller 504 (e.g., a DHCP lease database).

[0118] Figure 6 6 is a flow diagram of a process 600 according to some examples. The process 600 includes providing (at 602) a first DHCP service by a first controller of a controller cluster to a first set of client devices having characteristics mapped to the first controller, wherein the first DHCP service uses a first lease database. The characteristics may include MAC addresses or other characteristics.

[0119] Process 600 includes providing (at 604 ), by a second controller of the controller cluster, a second DHCP service to a second set of client devices having characteristics mapped to the second controller, wherein the first DHCP service uses a second lease database.

[0120] Process 600 includes synchronizing (at 606) a first lease database and a second lease database via an inter-controller communication path between the first controller and the second controller. An example of an inter-controller communication path is Figure 1 Cluster tunnel 120.

[0121] Process 600 includes detecting (at 608), by a first controller, a disruption in communication with a second controller via an inter-controller communication path. Process 600 includes determining (at 610), by the first controller, whether a second DHCP service of the second controller is unavailable based on a determination of whether DHCP lease activity at the second controller has occurred within a specified recent time interval.

[0122] Process 600 includes, based on determining that a second DHCP service of a second controller is unavailable, transitioning (at 612) a first controller to a partner down state as part of a DHCP service failover, wherein the first controller provides a first DHCP service to a first set of client devices and a second set of client devices.

[0123] Storage media (e.g. Figure 4404) may include any or some combination of the following: semiconductor memory devices, such as dynamic or static random access memory (DRAM or SRAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory; magnetic disks, such as fixed disks, floppy disks, and removable disks; another magnetic medium including tape; optical media, such as compact disks (CDs) or digital video disks (DVDs); or another type of storage device. Note that the instructions discussed above may be provided on one computer-readable or machine-readable storage medium, or alternatively, may be provided on multiple computer-readable or machine-readable storage media or multiple media distributed in a large system, possibly having multiple nodes. Such computer-readable or machine-readable storage media are considered part of an article of manufacture (or an article of industry). An article of manufacture or an article of industry may refer to any manufactured single component or multiple components. The storage medium or multiple storage media may be located in the machine that runs the machine-readable instructions, or at a remote site from which the machine-readable instructions for execution may be downloaded over a network.

[0124] In the present disclosure, the use of the terms "a", "an" or "the" is also intended to include plural forms, unless the context clearly indicates otherwise. In addition, the terms "include", "contain", "comprise", "cover", "have" or "have" when used in the present disclosure specify the presence of the elements described, but do not exclude the presence or addition of other elements.

[0125] In the foregoing description, numerous details have been set forth to provide an understanding of the subject matter disclosed herein. However, implementations may be practiced without some of these details. Other implementations may include modifications and variations to the details discussed above. The appended claims are intended to cover such modifications and variations.

Claims

1. A first controller of a controller cluster, the first controller comprising: Hardware processor; as well as a non-transitory storage medium storing instructions executable on the hardware processor to: providing a first controller network address allocation service to client devices in the first set of client devices; detecting a communication disruption with a second controller in the controller cluster, wherein communication through a first communication path between the first controller and the second controller is disrupted; determining whether a second controller network address allocation service of the second controller is unavailable based on a determination of whether network address allocation was performed at the second controller within a specified recent time interval; and Based on determining that the second controller network address allocation service of the second controller is unavailable, transitioning the first controller to a partner down state as part of a network address allocation failover, wherein the first controller provides the first controller network address allocation service to the client devices in the first set of client devices and provides the first controller network address allocation service to client devices in the second set of client devices associated with the second controller. 2 . The first controller according to claim 1 , wherein the first controller network address allocation service and the second controller network address allocation service are Dynamic Host Configuration Protocol (DHCP) services. 3 . The first controller of claim 1 , wherein the determination of whether the network address allocation was performed within the specified most recent time interval is based on checking a network address database.

4. The first controller of claim 3, wherein the network address database comprises a first network address database used by the second controller, and the instructions are executable on the hardware processor to: A second network address database is used at the first controller, the second network address database being synchronized with the first network address database.

5. The first controller of claim 4 , wherein in response to the first controller providing the first controller network address assignment service to the client device in the second set of client devices, the first network address database and the second network address database become out of sync, and wherein the instructions are executable on the hardware processor to: detecting re-establishment of communication between the first controller and the second controller; and Based on detecting that the communication is re-established between the first controller and the second controller, synchronization of the first network address database and the second network address database is initiated. 6 . The first controller according to claim 4 , wherein the first network address database and the second network address database are Internet Protocol (IP) address lease databases.

7. The first controller of claim 1 , wherein the first controller performs the first controller network address allocation service for the first set of client devices associated with a first subset of media access control (MAC) addresses, and the second controller performs the second controller network address allocation service for the second set of client devices associated with a second subset of MAC addresses.

8. The first controller of claim 1 , wherein the instructions are executable on the hardware processor to: The communication interruption with the second controller is detected based on detecting a failure to receive a first heartbeat indication over the first communication path. 9 . The first controller of claim 8 , wherein the failure to receive the first heartbeat indication over the first communication path comprises a failure to receive the first heartbeat indication in an Internet Protocol Security (IPsec) tunnel of the first communication path.

10. The first controller of claim 9, wherein the instructions are executable on the hardware processor to: The communication interruption with the second controller is also detected based on detecting a failure to receive a second heartbeat indication over a second communication path different from the first communication path.

11. The first controller of claim 1 , wherein the instructions are executable on the hardware processor to: responsive to said detecting a disruption in communications with said second controller, causing a health query to be sent from said first controller to said second controller via a second communication path between said first controller and said second controller, wherein said second communication path is different from said first communication path; as well as In response to the health query, receiving at the first controller via the second communication path an indication that the second controller network address allocation service of the second controller is unavailable; Wherein determining that the second controller network address allocation service of the second controller is unavailable is based on receiving the indication, and wherein the indication is based on the second controller determining that the network address allocation was not performed within the specified most recent time interval. 12 . The first controller of claim 11 , wherein the determination of whether the network address allocation was performed within the specified most recent time interval is relative to a timestamp of the health query.

13. The first controller of claim 1 , wherein the instructions are executable on the hardware processor to: Based on determining that the second controller network address allocation service of the second controller is unavailable, causing a shutdown indication to be sent from the first controller to the second controller to request the second controller to enter a shutdown state.

14. The first controller of claim 1 , wherein the instructions are executable on the hardware processor to: detecting a disruption in further communication with the second controller; responsive to said detecting a disruption of said further communications with said second controller, causing a health query to be sent from said first controller to said second controller via a second communications path; and The second controller network address assignment service of the second controller is determined to be unavailable based on detecting that the second controller has not responded to the health query after a threshold number of retries.

15. The first controller of claim 1 , wherein the determination of whether the network address allocation was performed within the specified most recent time interval is performed by a management system separate from the first controller and the second controller, and wherein instructions are executable on the hardware processor to: An indication is received at the first controller that the management system has shut down the second controller based on the management system determining that the second controller has not performed the network address allocation within the specified most recent time interval.

16. The first controller of claim 1 , wherein the instructions are executable on the hardware processor to: detecting, at the first controller, greater than a threshold number of client requests for network address allocation from the same client device or set of client devices within a specified time interval, wherein the client requests from the same client device or set of client devices should have been serviced by the second controller; and Based on the detection of greater than a threshold number of client requests from the same client device or set of client devices that should have been serviced by the second controller, a disruption in communication between the first controller and the second controller is indicated.

17. A system comprising: A controller cluster, the controller cluster comprising a first controller and a second controller, the first controller being configured to provide a first controller network address allocation service, the second controller being configured to provide a second controller network address allocation service, wherein the first controller has an inter-controller communication path with the second controller, The first controller is used for: detecting a disruption in communication with the second controller via the inter-controller communication path; determining whether the second controller network address allocation service of the second controller is unavailable based on a determination of whether network address allocation was performed at the second controller within a specified recent time interval; and Based on determining that the second controller network address allocation service of the second controller is unavailable, transitioning the first controller to a partner down state as part of a network address allocation failover, wherein the first controller provides the first controller network address allocation service to client devices in a first set of client devices associated with the first controller and provides the first controller network address allocation service to client devices in a second set of client devices associated with the second controller.

18. The system of claim 17, wherein the first controller is configured to: In response to transitioning to the buddy off state, sending a buddy off indication from the first controller to the second controller, The second controller is used for: In response to the buddy shutdown indication from the first controller, storing an indication that the second controller is to enter a recovery state when transitioning from an unavailable state; and In the recovery state, a database of assigned network addresses associated with the second controller and a database of assigned network addresses associated with the first controller are synchronized.

19. A method comprising: providing, by a first controller of a controller cluster, a first Dynamic Host Configuration Protocol (DHCP) service to a first set of client devices having characteristics mapped to the first controller, wherein the first DHCP service uses a first lease database; providing, by a second controller of the controller cluster, a second DHCP service to a second set of client devices having characteristics mapped to the second controller, wherein the first DHCP service uses a second lease database; synchronizing the first lease database and the second lease database via an inter-controller communication path between the first controller and the second controller; detecting, by the first controller, a disruption in communication with the second controller via the inter-controller communication path; determining, by the first controller, whether the second DHCP service of the second controller is unavailable based on a determination of whether DHCP lease activity has occurred at the second controller within a specified recent time interval; as well as Based on determining that the second DHCP service of the second controller is unavailable, transitioning the first controller to a partner down state as part of a DHCP service failover, wherein the first controller provides the first DHCP service to the first set of client devices and provides the first DHCP service to the second set of client devices.

20. The method of claim 19, wherein the determination at the second controller whether the DHCP lease activity occurred within the specified most recent time interval is based on information communicated over a second communication path through a computing environment, the computing environment including services to be accessed by client devices in the first set of client devices and the second set of client devices.