Main-standby switching method, device and system and storage medium
By utilizing fault messages and energy storage devices after a primary equipment failure, the backup equipment can be quickly switched to the primary equipment, solving the problem of insufficient switching speed in existing technologies and improving business continuity and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2026-03-10
AI Technical Summary
In a dual-homed network, when the primary device fails and loses power, the backup device can only switch to the primary device at a sub-second response time, which affects the user experience of low-latency services such as voice services.
After detecting a fault, the master device sends a fault message to the backup device. Upon receiving the message, the backup device quickly switches to the master device and sends the fault message using existing ports such as stacked links or peer-to-peer links. Combined with the power support provided by the energy storage device, the fault message is delivered in a timely manner.
The speed of switching from backup equipment to primary equipment has been improved to the millisecond level, reducing service interruptions or delays and enhancing the user experience.
Smart Images

Figure CN121644328A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a method, apparatus, system and storage medium for primary / standby switching. Background Technology
[0002] When the primary device in a dual-homed network fails and loses power, the backup device can determine that the primary device is in a faulty state by detecting a port failure (down) or a fast heartbeat timeout. The backup device can then switch over to become the primary device to ensure service continuity and reliability. Specifically, the backup device monitors the primary device's port status to determine port down times in seconds or sub-seconds, and uses heartbeat messages to determine fast heartbeat timeout times in seconds. Therefore, after a primary device fails and loses power, the fastest response time for the backup device to switch over to become the primary device can only reach sub-seconds. This may be noticeable for low-latency services (such as voice services), thus impacting user experience. Summary of the Invention
[0003] This application provides a method, apparatus, system, and storage medium for switching between primary and backup devices, which can improve the response speed of switching backup devices to primary devices to the millisecond level in the event of a primary device failure, thereby improving the user experience.
[0004] To achieve the above objectives, the embodiments of this application provide the following technical solutions:
[0005] Firstly, a master / slave switching method is provided. This method can be executed by a first device acting as the master device; or it can be executed by a module applied to the first device, such as a chip, chip system, or circuit; or it can be implemented by a logic module or software capable of implementing all or part of the functions of the first device, without limitation. For ease of description, the following explanation uses execution by the first device as an example.
[0006] The method includes: after detecting a failure in the first device, sending a failure message to the second device and switching from the primary device to the backup device.
[0007] The fault message is used to indicate that the first device is in a fault state and the second device is a backup device.
[0008] Through the above technical solution, after detecting a fault, the first device can promptly send a fault message to the second device, enabling the second device to quickly switch from backup to primary, effectively ensuring service continuity. Typically, sending a fault message takes only a few milliseconds or tens of milliseconds, meaning the speed is usually on the millisecond level. Therefore, the above technical solution improves the speed at which the backup device switches to primary after a primary device failure to the millisecond level, thus increasing the fault response speed to the millisecond level. This ensures that users do not experience noticeable interruptions or delays when using related services, thereby improving the user experience.
[0009] In one alternative implementation, the first device supports the dying-gasp function, and the fault message is a dying-gasp message.
[0010] The above provides a specific implementation method for fault messages, which can effectively improve the feasibility of this application.
[0011] In one optional implementation, the first device includes an energy storage device. When the first device is detected to have malfunctioned, sending a fault message to the second device may include: after detecting a malfunction in the first device, using the electrical energy released by the energy storage device to send a fault message to the second device.
[0012] With the above technical solution, even if the main power supply fails due to a failure of the first device, the energy storage device can provide power to send fault messages, ensuring the timely transmission of fault messages and effectively avoiding the situation where fault messages cannot be sent due to a sudden power outage.
[0013] In one alternative implementation, the energy storage device is a capacitor.
[0014] Compared to some other energy storage devices, capacitors typically have a smaller size and lighter weight, thus reducing the space occupied by the energy storage device. Furthermore, capacitors have a longer lifespan and can withstand multiple charge-discharge cycles without significant performance degradation. This means that when the first device uses capacitors as energy storage devices, it is not necessary to frequently replace energy storage components, which can effectively reduce the maintenance costs and downtime of the first device.
[0015] In one optional implementation, the first device is the master device in the stacked network, and the port on the first device used to connect to the second device is a stacking link port. Based on this, the aforementioned sending of a fault message to the second device may include: sending a fault message to the second device through the stacking link port.
[0016] In stacked networking, primary and backup devices are typically connected via a stacking link. This means the port on the primary device used to connect to the backup device is the stacking link port. Therefore, sending fault messages to the second device via the stacking link port utilizes the existing port, eliminating the need for a dedicated fault message transmission channel and reducing the complexity and cost of stacked networking. Furthermore, compared to other connection methods, stacking links often offer stronger stability and interference resistance. Sending fault messages to the second device via the stacking link port effectively reduces the risk of fault messages being lost or corrupted during transmission.
[0017] In one optional implementation, the first device is the master device in the stacked network, and the port on the first device used to connect to the second device is a stacking link port. Furthermore, the method may further include: sending a fault message to a network management device so that the network management device can update the device status of the first device.
[0018] In the above technical solution, after receiving a fault message, the network management device can update the device status of the first device stored in its own storage to ensure that the network topology and device status information stored in the network management device are consistent with the actual situation, so that the staff can accurately understand the device status of each switch in the stacked network through the network management device.
[0019] In one optional implementation, the first device is the master device in a multichassis link aggregation group (M-LAG) network, and the port used by the first device to connect to the second device is a peer-to-peer link port. Based on this, sending a fault message to the second device may include: sending a fault message to the second device through the peer-to-peer link port.
[0020] In M-LAG networking, the primary and backup devices are typically connected via peer-to-peer links. That is, the port used by the primary device to connect to the backup device is the peer-to-peer link port. Therefore, the method of sending fault messages to the second device through the peer-to-peer link port is to use the existing port to send fault messages, without the need to configure an additional dedicated fault message transmission channel, which can reduce the complexity and cost of M-LAG networking.
[0021] In one alternative implementation, the fault message is pre-configured in the first device.
[0022] With the above technical solution, after the first device fails, the pre-configured fault message in the first device can be sent out immediately without temporary generation. This can effectively shorten the time interval between the occurrence of the fault and the sending of the fault message.
[0023] Secondly, a primary / standby switching method is provided. This method can be executed by a second device acting as a standby device; alternatively, it can be executed by a module applied to the second device, such as a chip, chip system, or circuit; or it can be implemented by a logic module or software capable of realizing all or part of the functions of the second device, without limitation. For ease of description, the following explanation uses the execution by the second device as an example.
[0024] The method includes: receiving a fault message from a first device and switching from a standby device to a primary device.
[0025] Among them, the fault message is used to indicate that the first device is in a fault state, and the first device is the master device.
[0026] Through the above technical solution, the second device can switch from backup to master after receiving a fault message from the first device, effectively ensuring service continuity. Generally, sending a fault message takes only a few milliseconds or tens of milliseconds, meaning the speed is typically on the millisecond level. Therefore, the above technical solution improves the speed at which the backup device switches to master after the master device sends a fault to the millisecond level, thus increasing the fault response speed to the millisecond level. This ensures that users do not experience noticeable interruptions or delays when using related services, thereby improving the user experience.
[0027] In one alternative implementation, the fault message is a dying-gasp message.
[0028] The above provides a specific implementation method for fault messages, which can effectively improve the feasibility of this application.
[0029] In one optional implementation, the fault message carries a device identifier. Based on this, the above-mentioned switch from the standby device to the primary device may include: switching from the standby device to the primary device when the device identifier carried in the fault message is the same as the device identifier of the first device.
[0030] In the above technical solution, the second device only switches from the backup device to the master device if the device identifier carried in the fault message is consistent with the device identifier of the first device it is connected to, thus ensuring the accuracy of the switch and avoiding accidental switching.
[0031] In one alternative implementation, the second device may also update its own forwarding chip entries.
[0032] By updating the forwarding chip entries, service traffic can be quickly and effectively switched from the original primary device (i.e., the first device) to the new primary device (i.e., the second device), thereby ensuring service continuity and stability and reducing service interruptions or performance degradation caused by changes in the primary device.
[0033] Thirdly, a primary / backup switching device is provided. The device is located in a first device where the role is the primary device, and includes: a functional unit for executing any one of the methods provided in the first aspect, wherein the actions performed by each functional unit are implemented by hardware or by executing corresponding software through hardware.
[0034] The device includes a transceiver module and a processing module. The transceiver module is used to send a fault message to the second device after detecting a fault in the first device. The processing module is used to switch from the primary device to the backup device.
[0035] The fault message is used to indicate that the first device is in a fault state and the second device is a backup device.
[0036] Fourthly, a primary / backup switching device is provided, which is located in a second device that acts as a backup device, and includes: a functional unit for performing any of the methods provided in the second aspect, wherein the actions performed by each functional unit are implemented by hardware or by hardware executing corresponding software.
[0037] The device includes a transceiver module and a processing module. The transceiver module receives fault messages from the first device. The processing module is used for switching from the backup device to the primary device.
[0038] Among them, the fault message is used to indicate that the first device is in a fault state, and the first device is the master device.
[0039] Fifthly, a primary / standby switching system is provided, comprising: a first device and a second device, wherein the first device is the primary device and the second device is the standby device. The first device is used to execute any primary / standby switching method provided in the first aspect, and the second device is used to execute any primary / standby switching method provided in the second aspect.
[0040] In a sixth aspect, a computer program product is provided, the computer program product including instructions, which, when executed on a computer, enable the computer to perform any one of the primary / standby switching methods provided in the first or second aspect.
[0041] In a seventh aspect, a computer-readable storage medium is provided, including computer-executable instructions that, when executed on a computer, cause the computer to perform any of the primary / standby switching methods provided in the first to second aspects.
[0042] It should be noted that the technical effects of any of the implementation methods in aspects three through seven can be found in the technical effects of the corresponding implementation methods in aspects one and two, and will not be repeated here. Attached Figure Description
[0043] Figure 1This is a schematic diagram of the M-LAG networking structure provided in related technologies;
[0044] Figure 2 A schematic diagram illustrating the sending of heartbeat messages between the primary and backup devices provided in the related technologies;
[0045] Figure 3 This is a schematic diagram of a stacked network structure provided in related technologies;
[0046] Figure 4 A system architecture diagram of a primary / standby switching system provided in this application embodiment;
[0047] Figure 5 This is a schematic diagram of the composition of a primary / standby switching device provided in an embodiment of this application;
[0048] Figure 6 A schematic diagram of the interaction flow of a primary / standby switching method provided in an embodiment of this application;
[0049] Figure 7 This is a schematic diagram of a stacked network structure provided in an embodiment of this application;
[0050] Figure 8 This is a schematic diagram of the M-LAG network structure provided in an embodiment of this application;
[0051] Figure 9 A schematic diagram of the interaction flow of the master-slave switching method in the stacked networking provided in the embodiments of this application;
[0052] Figure 10 A schematic diagram of the interaction flow of the primary / standby switching method in M-LAG networking provided in the embodiments of this application;
[0053] Figure 11 This is a schematic diagram of the structure of a first device provided in an embodiment of this application;
[0054] Figure 12 This is a schematic diagram of the structure of a second device provided in an embodiment of this application. Detailed Implementation
[0055] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.
[0056] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0057] "Used for indication" can include direct and indirect indications, as well as explicit and implicit indications. When describing "indication information used to indicate A" or "indication information of A," it can include whether the indication information directly or indirectly indicates A, but does not necessarily mean that the indication information carries A. The information indicated by a certain piece of information is called the information to be indicated. In the specific implementation process, there are many ways to indicate the information to be indicated, such as, but not limited to, directly indicating the information to be indicated, such as the information to be indicated itself or its index. It can also indirectly indicate the information to be indicated by indicating other information, where there is a relationship between the other information and the information to be indicated. It can also indicate only a part of the information to be indicated, while the other parts are known or pre-agreed. At the same time, it is possible to identify the common parts of various pieces of information and unify the indication to reduce the indication overhead caused by individually indicating the same information. Furthermore, the specific indication method can also be any existing indication method, such as, but not limited to, the above-mentioned indication methods and their various combinations. Specific details of various indication methods can be found in existing technologies, and will not be elaborated upon here.
[0058] It is understood that the term "embodiment" used throughout the specification means that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, throughout the specification, various embodiments do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It is understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0059] In this application, unless otherwise specified, the same or similar parts between the various embodiments can be referred to each other. In the various embodiments of this application, unless otherwise specified or logically conflicting, the terminology and / or descriptions between different embodiments are consistent and can be mutually referenced. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships. The following embodiments of this application do not constitute a limitation on the scope of protection of this application.
[0060] Commonly used reliability networking in current data centers includes M-LAG networking implemented using M-LAG technology in dual-homed networking and stacked networking. In a dual-homed network, if the primary device experiences a power outage or hardware failure and needs to be reset to restore normal operation, the backup device can determine that the primary device is in a faulty state after detecting a port down or fast heartbeat timeout. Subsequently, the backup device can switch over to the primary device to ensure service continuity and reliability. The following descriptions of the primary and backup device switching processes will use M-LAG networking and stacked networking as examples.
[0061] In an M-LAG network, two devices can be included, such as... Figure 1 As shown, the two devices serve as the master and backup devices, respectively. They can be connected via a dual active detection link (DAD link) and a peer-link. The peer-link is used to send synchronization information (such as heartbeat messages) between the master and backup devices. Specifically, as... Figure 2 As shown, when both the primary and backup devices are in normal working order, they can send heartbeat messages to each other via peer-link to inform each other of their status, such as the absence of a fault. For example, the primary device can send a heartbeat message to the backup device via peer-link to inform the backup device that the primary device is functioning normally and has not experienced a fault. Furthermore, when both the primary and backup devices are in normal working order, third-party devices connected to the M-LAG network (such as…) Figure 1 The switch shown can send messages to the network through the primary device or through the backup device.
[0062] When the primary device fails, it cannot send heartbeat messages to the backup device. If the backup device does not receive heartbeat messages from the primary device within a specified time interval, it can determine that the primary device is faulty. Once the backup device determines that the primary device is faulty, it can switch its role to become the primary device. Accordingly, third-party devices connected to the M-LAG network can then send messages to the network through the new primary device.
[0063] In a stacked network, up to three devices can be included, such as... Figure 3As shown, these three devices play the roles of master, standby, and slave, respectively. They are interconnected via a stack trunk. The standby device can monitor the master device's port status to determine its overall status. For example, if the standby device detects that the master device's port is online, it confirms that the master device is functioning correctly. If the standby device detects that the master device's port is down, it confirms that the master device is faulty. In this case, the standby device can switch to become the new master, the slave device can switch to become the new standby, and the original master device becomes the new slave.
[0064] However, the backup device detects the port status of the primary device to determine if the primary device's port status is down within seconds or sub-seconds. It also determines a fast heartbeat timeout within seconds by not receiving heartbeat packets from the primary device within a specified time interval. Therefore, after a power outage caused by a primary device failure, the backup device's response time to switch back to primary is at most sub-seconds. This may be noticeable for low-latency services (such as voice services), thus impacting the user experience.
[0065] In view of this, this application provides a primary / backup switching method. After detecting a fault in itself, the first device (i.e., the primary device) in the network can send a fault message to the second device (i.e., the backup device). After receiving the fault message, the second device can switch from the backup device to the primary device. Correspondingly, the first device switches from the primary device to the backup device.
[0066] Using the above method, after detecting a fault, the first device can promptly send a fault message to the second device, enabling the second device to quickly switch from standby to primary, effectively ensuring service continuity. Typically, sending a fault message takes only a few milliseconds or tens of milliseconds, meaning the speed is usually on the millisecond level. Therefore, the above technical solution improves the speed at which the standby device switches to primary after the primary device sends a fault to the millisecond level, thus increasing the fault response speed to the millisecond level. This ensures that users do not experience noticeable interruptions or delays when using related services, thereby improving the user experience.
[0067] The technical solution provided in this application will now be described with reference to the accompanying drawings.
[0068] Figure 4 This is a system architecture diagram of a primary / standby switchover system provided in an embodiment of this application. Figure 4 As shown, this primary / standby switching system can include multiple devices, such as... Figure 4The first device 401 and the second device 402 shown are communicatively connected. The first device 401 is currently the master device, and the second device 402 is currently the backup device.
[0069] In the embodiments of this application, the first device 401 and the second device 402 may be two switches located in a network (such as an M-LAG network or a stacked network), two physical servers or two cloud servers located in a server cluster, or two databases in a database cluster, and there is no limitation on this.
[0070] In this embodiment of the application, after detecting a fault in itself, the first device 401 can send a fault message to the second device 402. After receiving the fault message, the second device 402 can switch from a standby device to a primary device, and correspondingly, the first device 401 switches from a primary device to a standby device.
[0071] The hardware included in the first device 401 and the second device 402 will be described below.
[0072] The first device 401 can be adopted Figure 5 The shown composition structure, or including Figure 5 The components shown. Figure 5 This is a schematic diagram illustrating the composition of a primary / standby switching device 500 provided in an embodiment of this application. The primary / standby switching device 500 can be a first device 401 or a chip or system-on-a-chip within the first device 401. For example... Figure 5 As shown, the master / slave switching device 500 may include a processor 501, a power detection device 502, a power storage device 503, a communication interface 504, and a communication line 505.
[0073] The processor 501, power detection device 502, energy storage device 503, and communication interface 504 can be connected via communication line 505.
[0074] The processor 501 can be a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 501 can also be other devices with processing capabilities, such as circuits, devices, or software modules, without limitation.
[0075] The power detection device 502 may include multiple electronic components and circuits for real-time monitoring of the status of the external input power supply, such as detecting whether the current input by the external input power supply is less than a preset current threshold. If the current input by the external input power supply is detected to be less than the preset current threshold, it indicates that the main / standby switching device 500 is in a fault state.
[0076] In this embodiment, the power detection device 502 can be a power detection circuit, a power management chip, etc., and is not limited thereto.
[0077] In one alternative implementation, after the power detection device 502 determines that the primary / backup switching device 500 is in a fault state, the power detection device 502 can send an interrupt signal to the processor 501. After receiving the interrupt signal, the processor 501 can send a fault message to the backup device.
[0078] In another optional implementation, the primary / standby switching device 500 may further include a chip 507. The chip 507 can receive an interrupt signal sent by the power detection device 502 and, in response to the interrupt signal, send a fault message to the standby device. Specifically, after confirming that the primary / standby switching device 500 is in a fault state, the power detection device 502 can send an interrupt signal to the chip 507. After receiving the interrupt signal, the chip 507 can send a fault message to the standby device.
[0079] It is understandable that chip 507 can be a standalone device, such as Figure 5 As shown, the chip 507 can be connected to the processor 501, the power detection device 502, the energy storage device 503, and the communication interface 504 via a communication line 505. The chip 507 can also be a chip integrated into the processor 501.
[0080] In the embodiments of this application, chip 507 can be a forwarding chip or an application-specific integrated circuit (ASIC), without limitation.
[0081] The energy storage device 503 can store electrical energy when the main / standby switching device 500 is in a normal state. When the main / standby switching device 500 is in a fault state, such as when the voltage difference across its terminals determines that the main / standby switching device 500 is in a fault state, it releases the electrical energy it stores to provide energy support for the main / standby switching device 500 to send a fault message.
[0082] In this embodiment, the energy storage device 503 can be a capacitor, supercapacitor, battery, or other components, and is not limited thereto.
[0083] Communication interface 504 is used for communication with other devices or other communication networks. These other communication networks can be Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc. Communication interface 504 can be a module, circuit, transceiver, or any device capable of enabling communication.
[0084] Communication line 505 is used to transmit information between the components included in the primary / standby switching device 500.
[0085] In one alternative implementation, processor 501 may include one or more CPUs, for example... Figure 5 CPU0 and CPU1 in the CPU.
[0086] In one alternative implementation, the primary / standby switching device 500 includes multiple processors, for example, besides Figure 5 In addition to processor 501, it may also include processor 506.
[0087] It should be noted that the primary / standby switching device 500 can be a switch, router, server, or other similar device. Figure 5 Equipment with a similar structure. Furthermore... Figure 5 The structural composition shown does not constitute a limitation on the master / slave switching device, except... Figure 5 In addition to the components shown, the main / standby switching device may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0088] The second device 402 may include a processor, a communication interface, and a communication line. The processor and the communication interface can be connected via the communication line.
[0089] In this embodiment, the roles of the first device 401 and the second device 402 can be switched. Specifically, after receiving a fault message, the second device 402 can switch from a backup device to a new master device. After becoming the new master device, the second device 402 can also send a fault message to the backup device connected to it after detecting a fault, so that the backup device can switch to the next master device. Based on this, the second device 402 may also include the aforementioned... Figure 5 The power detection device and energy storage device shown, as well as the connection relationships between the processor, communication interface, power detection device, energy storage device, and communication line, and the functions of the processor, communication interface, power detection device, energy storage device, and communication line, can all be referenced. Figure 5 As shown, it will not be elaborated further here.
[0090] Furthermore, the actions, terms, etc., involved in the various embodiments of this application can be referenced interchangeably without limitation. The message names or parameter names in the messages exchanged between the various devices in the embodiments of this application are merely examples, and other names may be used in specific implementations without limitation.
[0091] The following is combined with Figure 4 ,by Figure 4 Taking the first and second devices as examples, the primary / backup switching method provided in this application embodiment is described. Figure 6 This is a schematic diagram of the interaction flow of a primary / standby switchover method provided in an embodiment of this application, as shown below. Figure 6 As shown, the method includes:
[0092] S601, after the first device detects a fault in itself, it sends a fault message to the second device.
[0093] The first device is the primary device, and the second device is the backup device.
[0094] A fault message is used to indicate that the first device is in a fault state. The fault message may carry the device identifier of the first device.
[0095] When the first device is in normal working order, the external power supply connected to it can continuously supply current to the first device, providing power to ensure its normal operation. If the first device malfunctions, the external power supply connected to it can no longer supply current. If the first device detects that the current supplied by the external power supply is less than a preset current threshold, it can send a fault message to the second device through a power-off protection mechanism.
[0096] Among them, the power failure protection mechanism refers to a series of measures taken after the current supplied by the external power supply disappears in order to send a fault message to the second device.
[0097] In the embodiments of this application, the preset current threshold is not limited. For example, the preset current threshold can be 1 milliampere (mA) or 5 mA, and is not limited.
[0098] In one optional implementation, the first device may include a power detection device and a chip. The power detection device may include multiple electronic components and circuits for real-time monitoring of the status of the external input power supply. The chip may send a fault message to the second device. Based on this, the power detection device can monitor the status of the external input power supply in real time, and if it detects that the current input by the external power supply is less than a preset current threshold, it confirms that the first device is in a fault state. Subsequently, the power detection device may send an interrupt signal to the chip, and upon receiving the interrupt signal, the chip may send a fault message to the second device.
[0099] It is understandable that the above-mentioned behavior of sending an interrupt signal to the chip when the power detection device detects that the current input of the external power supply is less than the preset current threshold, and the behavior of the chip sending a fault message to the second device after receiving the interrupt signal, is an optional power-loss protection mechanism.
[0100] In the embodiments of this application, the power detection device can be a power detection circuit, a power management chip, etc., and is not limited thereto.
[0101] In the embodiments of this application, the chip can be a forwarding chip or an ASIC, without limitation. The following description uses a forwarding chip as an example.
[0102] In an optional implementation, the first device may include an energy storage device that can store electrical energy when the first device is in a normal state and release the stored electrical energy when the first device is in a fault state, providing energy support for the first device to send a fault message. Accordingly, the above-described S601 can be replaced by: after detecting a fault in the first device, using the electrical energy released by the energy storage device to send a fault message to the second device.
[0103] The embodiments of this application do not limit the energy storage device. For example, the energy storage device can be a capacitor, supercapacitor, battery or other components. The following description takes a capacitor as an example of an energy storage device.
[0104] For example, when the first device is in a faulty state, the external input power supply connected to the first device cannot continue to supply current to the first device, causing a voltage difference to appear across the capacitor of the first device. Therefore, the capacitor of the first device can begin to discharge and release electrical energy when there is a voltage difference across it. In addition, the power detection device can use the electrical energy released by the capacitor to send an interrupt signal to the forwarding chip. After receiving the interrupt signal, the forwarding chip can use the electrical energy released by the capacitor to send a fault message to the second device.
[0105] In an alternative implementation, the first device may further include a buffer circuit, which can provide short-term power support to the first device when the current supplied by the external input power supply connected to the first device is less than a preset current threshold, so that the first device can perform some necessary emergency operations, such as data saving.
[0106] In the case where the first device includes a power storage device, if the current supplied by the external input power supply connected to the first device is less than a preset current threshold, the buffer circuit can also use the electrical energy released by the power storage device to extend its own operating time, so as to deliver the electrical energy released by the power storage device to various devices (such as forwarding chips, processors, etc.) in the first device and maintain the operation of various devices in the first device.
[0107] In this embodiment of the application, the duration of operation extended by the energy released by the capacitor in the cache circuit of the first device is not limited, such as 100 milliseconds (ms), 200 ms, etc. The specific duration can be determined according to factors such as the capacitance of the capacitor and the power consumption of the cache circuit.
[0108] In an alternative implementation, if the first device supports the dying-gasp function, the aforementioned fault message can be a dying-gasp message.
[0109] In one alternative implementation, the aforementioned fault message may be pre-configured in the first device. For example, the first device may be configured by developers before it leaves the factory. Alternatively, the aforementioned fault message may be pre-generated and stored in the first device before being sent to the second device. For example, the first device may generate and store the fault message while in normal operation.
[0110] In one alternative implementation, the first device can send a fault message to the second device via a transmission port. The transmission port may be pre-configured in the first device.
[0111] The sending port refers to the port used by the first device to connect to the second device.
[0112] Specifically, before sending a fault message to the second device, the first device can first determine the port (i.e., the sending port) it uses to connect to the second device, and then send the fault message through that sending port to achieve the sending of the fault message to the second device.
[0113] In one alternative implementation, the first device can send a fault message to the second device via a transmission port and a transmission link. The transmission port and transmission link may be pre-configured in the first device.
[0114] The transmission link refers to any physical line or transmission medium used for data transmission between the first device and the second device.
[0115] Specifically, before the first device sends a fault message to the second device, it can first determine the sending port and sending link used to connect to the second device. Then, the first device can send the fault message through the sending port and sending link to achieve the sending of the fault message to the second device.
[0116] S602, the first device switches from the primary device to the backup device.
[0117] After determining that it is in a faulty state, the first device can switch from being the primary device to the backup device. For example, the first device can passively switch to the backup device after the electrical energy stored in the energy storage device is depleted. Alternatively, after sending a fault message to the second device, the first device can actively switch to the backup device using the electrical energy released by the energy storage device before the electrical energy stored in the energy storage device is depleted.
[0118] S603, the second device receives a fault message from the first device.
[0119] S604, the second device switches from the backup device to the master device.
[0120] Specifically, after the first device determines that it is in a faulty state and sends a fault message to the second device, the second device can receive the fault message from the first device. The second device can parse the fault message to obtain the device identifier carried in the fault message. If the device identifier is consistent with the device identifier of the master device to which it is connected, the second device can determine that the master device in this network is in a faulty state. At this time, the second device can switch from the standby device to the master device.
[0121] In one optional implementation, the second device may include a role arbitration module. The role arbitration module can parse fault messages and, if the device identifier carried in the fault message is consistent with the device identifier of the master device to which the second device is connected, switch the role of the second device, that is, switch the second device from a standby device to a master device.
[0122] In this embodiment, the roles of the first device and the second device can be switched. That is, after sending a fault message to the second device, the first device can switch from being the primary device to a new backup device. After becoming the new backup device, the first device can also parse the fault message upon receiving it subsequently. Therefore, the first device can also include a role arbitration module, and the function of the role arbitration module can refer to the function of the role arbitration module in the second device described above, which will not be repeated here.
[0123] In one optional implementation, the second device may include forwarding chip entries, which may contain data forwarding related information, such as the destination address of the data packet, the sending port, and the forwarding path. After switching from a backup device to a master device, the second device may update its own forwarding chip entries, such as updating the data forwarding paths contained in the forwarding chip entries to conform to its own network topology.
[0124] For details on updating the entries in the forwarding chip table, please refer to relevant technologies; they will not be elaborated here.
[0125] By updating the forwarding chip entries, service traffic can be quickly and effectively switched from the original primary device (i.e., the first device) to the new primary device (i.e., the second device), thereby ensuring service continuity and stability and reducing service interruptions or performance degradation caused by changes in the primary device.
[0126] Through the above technical solution, after detecting a fault, the first device can promptly send a fault message to the second device, enabling the second device to quickly switch from backup to primary, effectively ensuring service continuity. Typically, sending a fault message takes only a few milliseconds or tens of milliseconds, meaning the speed is usually on the millisecond level. Therefore, the above technical solution can improve the speed of switching from backup to primary to the millisecond level, thus improving the fault response speed to the millisecond level. This ensures that users do not experience noticeable interruptions or delays when using related services, thereby improving the user experience.
[0127] In this embodiment, the first device and the second device can be two switches located in a network or two servers located in a server cluster; there is no limitation in this regard. The following description uses two switches located in a network as an example.
[0128] The above networking can be either stacked networking or M-LAG networking.
[0129] like Figure 7 As shown, a stacked network typically includes three devices (such as switches), designated as the master, backup, and slave. These devices are interconnected via stack trunks. The sending port on the master device used to connect to the backup device is called the stack trunk port. The master device can also connect to a network management device via an in-band link; this in-band link port is called the master device's sending port. The network management device stores device information and status data for the master, backup, and slave devices.
[0130] like Figure 8As shown, an M-LAG network typically includes two devices (such as switches), designated as the primary and backup devices. The primary and backup devices are connected via a peer-link. The sending port on the primary device used to connect to the backup device is called the peer-link port. Furthermore, the primary and backup devices can also be connected to a peer device (also known as an eth-trunk peer) using Ethernet link aggregation (eth-trunk) technology.
[0131] The following description uses the dying-gasp fault message as an example, the chip as a forwarding chip, and the energy storage device as a capacitor to illustrate the master / slave switching methods in stacked networking and M-LAG networking.
[0132] Combination Figure 7 , Figure 9 This diagram illustrates the interaction flow of the master / slave switchover method in a stacked network provided in this embodiment. The method is implemented through interaction between switch A, switch B, a network management device, and switch C. Switch A acts as the master device in the stacked network, switch B as the standby device, and switch C as the slave device. Figure 9 As shown, the method includes:
[0133] S901, when switch A is in the up state of its own stack trunk port, it generates and stores the dying-gasp message template.
[0134] In a stacked network, the dying-gasp message template can include the dying-gasp message and the stack trunk port.
[0135] S902, the capacitor of switch A determines that a fault has occurred in switch A by the voltage difference across its two ends, and then releases electrical energy.
[0136] Device 1 is the first device described above, that is: Device 1 is the master device in the stacked network.
[0137] S903, when the power detection device of switch A detects that the current input of the external power supply is less than the preset current threshold, it sends an interrupt signal to the forwarding chip using the electrical energy released by the capacitor.
[0138] After receiving the interrupt signal, the forwarding chip of switch A sends a dying-gasp message to the second device through the stack trunk port in the dying-gasp message template.
[0139] S905, switch A sends dying-gasp messages to the network management device through its in-band link port.
[0140] S906, the network management device updates the device status of switch A.
[0141] Specifically, such as Figure 9 As shown, after receiving the interrupt signal, the forwarding chip of switch A can also send a dying-gasp message to the network management device connected to switch A through the in-band link port. After receiving the dying-gasp message, the network management device can update the device status of switch A stored in itself to ensure that the network topology and device status information stored in the network management device are consistent with the actual situation, so that the staff can accurately understand the device status of each switch in the stacked network through the network management device.
[0142] In this embodiment, after receiving the interrupt signal, the forwarding chip of switch A can first send a dying-gasp message to switch B through the stack trunk port, and then send a dying-gasp message to the network management device through the in-band link port. Alternatively, it can send a dying-gasp message to switch B through the stack trunk port while simultaneously sending a dying-gasp message to the network management device through the in-band link port. No limitation is imposed.
[0143] S907, after receiving the dying-gasp message, the role arbitration module of switch B parses the dying-gasp message.
[0144] S908, if the device identifier carried in the dying-gasp message is the same as the device identifier of the master device connected to switch B, the role arbitration module of switch B will switch from the backup device to the master device.
[0145] S909, switch A switches from the primary device to the backup device.
[0146] S910, switch B sends an announcement message to switch C through the stack trunk port.
[0147] After switch B switches from standby to master, it can send a notification message to switch C through the stack trunk port to inform switch C that it (referring to switch B) has switched to master. This allows switch C to adjust its operating mode and communication method with switch B in a timely manner. For example, when switch C needs to send data, request resources, or report status to the master, it can communicate with switch B.
[0148] Combination Figure 8 , Figure 10This is a schematic diagram of the interaction flow of the master-slave switching method in the M-LAG network provided in this application embodiment. This method is implemented through interaction between switch 1, switch 2, and the eth-trunk peer. Switch 1 acts as the master device in the M-LAG network, and switch 2 acts as the standby device in the M-LAG network. Figure 10 As shown, the method includes:
[0149] S1001, when its peer-link port is in the up state, switch 1 generates and stores the dying-gasp message template.
[0150] In M-LAG networking, the dying-gasp message template can include the dying-gasp message and the peer-link port.
[0151] S1002, the capacitor of switch 1 determines that a fault has occurred in switch 1 by the voltage difference between its two ends, and then releases electrical energy.
[0152] S1003, when the power detection device of switch 1 detects that the current input of the external power supply is less than the preset current threshold, it sends an interrupt signal to the forwarding chip using the electrical energy released by the capacitor.
[0153] S1004 After receiving the interrupt signal, the forwarding chip of switch 1 sends a dying-gasp message to switch 2 through the peer-link port in the dying-gasp message template.
[0154] S1005, after receiving the dying-gasp message, the role arbitration module of switch 2 parses the dying-gasp message.
[0155] S1006, if the device identifier carried in the dying-gasp message of switch 2 is the same as the device identifier of the master device connected to switch 2, the role arbitration module of switch 2 will switch from the backup device to the master device.
[0156] S1007, Switch 2 sends a Link Aggregation Control Protocol (LAC) message to the peer on the eth-trunk.
[0157] S1008, Switch 1 switches from the primary device to the backup device.
[0158] In an M-LAG network, after switch 2 switches from a backup device to a primary device, in order to re-optimize the load balancing of network links or adapt to new network conditions, such as... Figure 10As shown, switch 2 can also send information about the switching of primary and backup devices to the peer of the eth-trunk via link aggregation control protocol (LACP) messages, so that the peer of the eth-trunk can switch the member ports of the eth-trunk according to the instructions of the LACP messages, so as to ensure normal data transmission and stable network operation.
[0159] The above mainly describes the solutions provided by the embodiments of this application from the perspective of interaction between various devices. Correspondingly, the embodiments of this application also provide a primary / backup switching device, which is used to implement the various methods described above. This primary / backup switching device can be the first device in the above method embodiments, or a component that can be used in the first device; or, the primary / backup switching device can be the second device in the above method embodiments, or a component that can be used in the second device. It is understood that, in order to achieve the above functions, the primary / backup switching device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0160] This application embodiment can divide the primary / standby switching device into functional modules according to the above method embodiment. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be understood that the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0161] For example, taking the primary / standby switching device as the first device in the above method embodiment as an example, Figure 11 A schematic diagram of a first device is shown, which includes a transceiver module 1101 and a processing module 1102. The transceiver module 1101, also known as a transceiver unit, is used to implement transceiver functions, and may be, for example, a transceiver circuit, a transceiver, a transceiver device, or a communication interface.
[0162] The transceiver module 1101 is used to send a fault message to the second device after detecting a fault in the first device. The fault message indicates that the first device is in a fault state and the second device is a backup device.
[0163] Processing module 1102 is used to switch from the primary device to the backup device.
[0164] The transceiver module 1101 can be used to implement the transceiver function of the first device in the above method embodiment, and the processing module 1102 can be used to implement the processing function of the first device in the above method embodiment. Therefore, all relevant content of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module, and will not be repeated here.
[0165] In this embodiment, the first device is presented as an integrated unit divided into functional modules. Here, "module" can refer to a specific ASIC, circuitry, a processor and memory executing one or more software or firmware programs, integrated logic circuitry, and / or other devices that can provide the aforementioned functions. In a simplified embodiment, those skilled in the art will recognize that the first device can employ... Figure 5 The main / standby switching device 500 shown is in the form of this device.
[0166] Since the first device provided in this application embodiment can execute the above-described master-slave switching method, the technical effects it can achieve can be referred to the above-described method embodiment, and will not be repeated here.
[0167] Alternatively, for example, taking the primary / standby switching device as the second device in the above method embodiment as an example, Figure 12 A schematic diagram of a second device is shown, which includes a transceiver module 1201 and a processing module 1202. The transceiver module 1201, also known as a transceiver unit, is used to implement transceiver functions, and may be, for example, a transceiver circuit, a transceiver, a transceiver device, or a communication interface.
[0168] The transceiver module 1201 is used to receive fault messages from the first device. The fault message indicates that the first device is in a fault state, and the first device is the master device.
[0169] Processing module 1202 is used for switching from standby equipment to master equipment.
[0170] The transceiver module 1201 can be used to implement the transceiver function of the second device in the above method embodiment, and the processing module 1202 can be used to implement the processing function of the second device in the above method embodiment. Therefore, all relevant content of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module, and will not be repeated here.
[0171] In this embodiment, the second device is presented as an integrated unit divided into functional modules. Here, "module" can refer to a specific ASIC, circuitry, a processor and memory executing one or more software or firmware programs, integrated logic circuitry, and / or other devices that can provide the aforementioned functions. In a simplified embodiment, those skilled in the art will recognize that the second device can employ... Figure 5 The main / standby switching device 500 shown is in the form of this device.
[0172] Since the second device provided in this application embodiment can execute the above-described master-slave switching method, the technical effects it can achieve can be referred to the above-described method embodiment, and will not be repeated here.
[0173] It should be understood that one or more of the above modules or units can be implemented by software, hardware, or a combination of both. When any of the above modules or units are implemented by software, the software exists as computer program instructions and is stored in memory. The processor can be used to execute the program instructions and implement the above method flow. The processor can be built into a SoC (System-on-a-Chip) or ASIC, or it can be a separate semiconductor chip. In addition to the core that executes software instructions for computation or processing, the processor may further include necessary hardware accelerators, such as field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), or logic circuits that implement dedicated logic operations.
[0174] When the above modules or units are implemented in hardware, the hardware can be any one or any combination of a CPU, microprocessor, digital signal processing (DSP) chip, micro controller unit (MCU), artificial intelligence processor, ASIC, SoC, FPGA, PLD, application-specific digital circuit, hardware accelerator, or non-integrated discrete device, which can run the necessary software or perform the above method flow independently of software.
[0175] Optionally, embodiments of this application also provide a primary / backup switching device (e.g., the primary / backup switching device may be a chip or a chip system), which includes a processor for implementing the methods in any of the above method embodiments. In one possible design, the primary / backup switching device further includes a memory. The memory is used to store necessary program instructions and data, and the processor can call the program code stored in the memory to instruct the primary / backup switching device to execute the methods in any of the above method embodiments. Of course, the memory may not be included in the primary / backup switching device. When the primary / backup switching device is a chip system, it may be composed of chips or may include chips and other discrete devices; embodiments of this application do not specifically limit this.
[0176] In one possible implementation, this application also provides a computer-readable storage medium storing a computer program or instructions that, when run on a primary / standby switching device, enable the primary / standby switching device to execute the methods described in any of the above method embodiments or any implementation thereof.
[0177] In one possible implementation, this application embodiment also provides a communication system, which includes the first device and the second device described in the above method embodiments.
[0178] In one possible implementation, this application also provides a primary / backup switching method, which includes the method described in any of the above-described method embodiments or any of their implementations.
[0179] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
[0180] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, the disclosure, and the appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.
[0181] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.
Claims
1. A method for master-backup switchover, characterized in that, The method is applied to a first device, the first device is a master device, and the method comprises: sending a failure message to a second device after detecting that the first device fails; the failure message is used to indicate that the first device is in a failure state; the second device is a backup device; switching from the master device to the backup device.
2. The method of claim 1, wherein, The first device supports a dying-gasp function, and the failure message is a dying-gasp message.
3. The method according to claim 1 or 2, characterized in that, The first device comprises a power storage device. The step of sending the failure message to the second device after detecting that the first device fails comprises: sending the failure message to the second device by using power released by the power storage device after detecting that the first device fails.
4. The method of claim 3, wherein, The power storage device is a capacitor.
5. The method according to any one of claims 1-4, characterized in that, The first device is a master device in a stacked networking group; and a port of the first device used to connect the second device is a stacked link port. The step of sending the failure message to the second device comprises: sending the failure message to the second device through the stacked link port.
6. The method according to any one of claims 1-4, characterized in that, The first device is a master device in a multi-chassis link aggregation group (M-LAG) networking group; and a port of the first device used to connect the second device is a peer-to-peer link port. The step of sending the failure message to the second device comprises: sending the failure message to the second device through the peer-to-peer link port.
7. A method for master-backup switchover, characterized in that, The method is applied to a second device, the second device is a backup device, and the method comprises: receiving a failure message from a first device; the failure message is used to indicate that the first device is in a failure state; the first device is a master device; switching from the backup device to the master device.
8. The method of claim 7, wherein, The failure message is a dying-gasp message.
9. The method according to claim 7 or 8, characterized in that, The failure message carries a device identifier. The step of switching from the backup device to the master device comprises: switching from the backup device to the master device in a case where the device identifier carried in the failure message is the same as a device identifier of the first device.
10. A master-backup switchover apparatus characterized by comprising: The device is located in a first device, the first device is a master device, and the device comprises: a transceiver module, configured to send a failure message to a second device after detecting that the first device fails; the failure message is used to indicate that the first device is in a failure state; the second device is a backup device; a processing module, configured to switch from the master device to the backup device.
11. A master-backup switchover apparatus characterized by comprising: The device is located in a second device, the second device is a backup device, and the device comprises: a transceiver module, configured to receive a failure message from a first device; the failure message is used to indicate that the first device is in a failure state; the first device is a master device; a processing module, configured to switch from the backup device to the master device.
12. A master-backup switching system, characterized by comprising: The system comprises a first device and a second device; the first device is a master device; and the second device is a backup device. The first device is configured to perform the master-backup switching method in any one of claims 1 to 6. The second device is configured to perform the master-backup switching method in any one of claims 7 to 9.
13. A computer-readable storage medium, characterized in that, comprises program code which, when running on a computer or processor, causes the computer or the processor to carry out the master-backup switchover method according to any one of claims 1-6 or the master-backup switchover method according to any one of claims 7-9.