Fault isolation method and apparatus based on white box switch, electronic device, and readable storage medium
By establishing a monitoring group on a white box switch and configuring preset ports, automated fault detection and isolation are solved, and the problems of slow failure convergence speed and dependence on external controllers in the DCN network are solved, and fast and thorough fault isolation and recovery are achieved.
Patent Information
- Application Number
- PCT/CN2024/135499
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-11-29
- Publication Date
- 2025-06-19
AI Technical Summary
In current DCN networks, as the number of white box switches increases, automated failure isolation and recovery becomes increasingly important, and prior art is difficult to achieve rapid and thorough failure convergence and relies on external controllers and complex protocol interactions.
By establishing a monitoring group on the white box switch, configuring preset ports and data, and storing configuration information using the hash type and key-value structure of Redis, it can automatically detect faults and trigger batch port shutdown and recovery, avoiding the security risks caused by manual operations.
It realizes fault convergence based on port admin status, does not rely on protocol interaction, and the convergence speed is faster and more thorough. It supports delayed online to prevent network jitter, iterative development efficiency is high, deployment is fast, and no need to rely on external controllers.
Smart Images

Figure CN2024135499_19062025_PF_FP_ABST
Abstract
Description
A white box switch-based fault isolation method, device, electronic device, and readable storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 12, 2023, with application number 202311703073.2 and invention name “A fault isolation method based on white box switch”, the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the technical field of monitoring and operation and maintenance of data communication white-box switches, and in particular to a fault isolation method, device, electronic device, and readable storage medium based on a white-box switch. Background Art
[0004] A DCN, or Data Communication Network, supports the seven-layer protocol stack, with Layer 1 (physical layer), Layer 2 (data link layer), and Layer 3 (network layer). It primarily carries management information and distributed signaling messages. A DCN features a distributed network computing environment and a multi-level distributed data warehouse. The number of switches in current DCNs is rapidly increasing, with white-box switches making up a growing portion of the network. White-box switches are easy to customize, making it crucial to provide a universal fault isolation method. Summary of the Invention
[0005] This application aims to at least partially address one of the technical problems in the related art. To this end, one objective of this application is to provide a white-box switch-based fault isolation method, apparatus, electronic device, and readable storage medium. This application can automatically discover and isolate faults across a large number of white-box switches in a data center, while eliminating the potential safety hazards associated with manual operations and improving the self-healing capabilities of data center network services.
[0006] A first aspect disclosed in the present application provides a fault isolation method based on a white box switch, the method comprising:
[0007] Create a monitoring group on the white box switch;
[0008] The monitoring group is configured with preset ports and data, and status data of the monitoring group is configured;
[0009] When the white box switch detects a fault on itself or a protocol anomaly, it triggers the monitoring group to perform a batch shutdown on the preset ports to isolate the fault;
[0010] After the fault or protocol abnormality is recovered, batch no shutdown is executed on the preset ports, and the white box switch can continue to provide services.
[0011] The step of establishing a monitoring group on the white box switch includes:
[0012] Establish a monitoring group in the white box switch SONIC system;
[0013] Log in to the control interface of the white box switch SONIC system;
[0014] Click the "Monitoring Group" option on the control interface;
[0015] Click the "New" button and enter the name and description of the monitoring group;
[0016] Click the "OK" button to create the monitoring group.
[0017] The monitoring group configures preset ports and data, including:
[0018] Use the hash type and key filed value structure of redis to configure the preset port and data, and save the preset port and data of the monitoring group configuration in CONFIG_DB.
[0019] The use of Redis's hash type and key filed value structure to configure the preset port and data includes:
[0020] Key="MONITOR_GROUP|MGxx"
[0021] Field Value
[0022] "ports"<list_value>
[0023] "up_delay_time" <value>
[0024] "MGxx": xx is an int number that identifies the monitoring group number;
[0025] "ports": preset ports, the port group for executing shutdown and no shutdown, which can be one or more ports;
[0026] "up_delay_time": Delay operation when switching from admin down to admin up.
[0027] The status data of the configuration monitoring group includes:
[0028] The state data of the configuration monitoring group is stored in the STATE_DB database, and there are three states: init, up, and down;
[0029] Key="MONITOR_GROUP_STATE|MGxx";
[0030] Field Value
[0031] "state" <init up down>
[0032] MG port status, stored in the STATE_DB database;
[0033] Key = "MG_PORT|{port}";
[0034] Field Value
[0035] "MonitorGroupxx" <up down>
[0036] Subscribe to MONITOR_GROUP_STATE|MGxx status and set the status of MG|port}.
[0037] Subscribing to the MONITOR_GROUP_STATE|MGxx status and setting the status of MG|port} include:
[0038] Subscribe to MONITOR_GROUP_STATE|MGxx status and set the status of MG|port}. Portmgrd determines the port admin value sent to the hardware based on the user-configured shutdown status and MG_PORT|port} data.
[0039] The steps of subscribing to the MONITOR_GROUP_STATE|MGxx state and setting the MG|port} state; and determining the port admin value sent to the hardware based on the user-configured shutdown state and MG_PORT|port} data by portmgrd include:
[0040] The white box switch is initialized, the program enters the working state, and starts subscribing to the Redis STATE_DB database MonitorGroup*;
[0041] The default state of MG is init, which is the initial state. It initializes the soft data and does not perform other operations.
[0042] Upon receiving subscription information and the MONITOR_GROUP_STATE state is down, a shutdown flag is written to each port in the preset ports list, recorded in STATE_DB, and the MG_PORT|port} of the corresponding port is set to MonitorGroupxx down;
[0043] Upon receiving subscription information and the MONITOR_GROUP_STATE state is up, the corresponding port's MG_PORT|{port} is set to MonitorGroupxx down. If the state changes from down to up, a delay process is executed. The delay time is the configured up_delay_time. This can effectively prevent network jitter caused by frequent port updowns. If the state changes from init to up, no delay is required.
[0044] The portmgr process determines the port admin status that is ultimately sent to the hardware. Based on user configuration, admin status, and the status of each monitoring group in MG_PORT|port}, if one is marked as down, the status sent to the hardware admin is down; if all are up, the status sent to the hardware admin is up.
[0045] A second aspect disclosed in the present application provides a fault isolation device based on a white box switch, the device comprising:
[0046] Establish monitoring group module, used to establish monitoring groups on white box switches;
[0047] Configuration data module, used for configuring preset ports and data for the monitoring group, and configuring status data of the monitoring group;
[0048] The monitoring group executes a batch shutdown module on the preset ports, which is used to trigger the monitoring group to execute a batch shutdown module on the preset ports to isolate the fault when the white box switch detects its own fault or detects a protocol abnormality;
[0049] The preset port executes a batch no shutdown module, which is used to execute a batch no shutdown on the preset port after the failure or protocol abnormality is recovered, and the white box switch can continue to provide services.
[0050] The third aspect disclosed in the present application is an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in a fault isolation method based on a white box switch are implemented.
[0051] The fourth aspect disclosed in the present application is a readable storage medium, wherein the readable storage medium stores a computer program, wherein the computer program is suitable for being loaded by a processor to execute the steps in the method for isolating faults based on a white box switch.
[0052] Compared with the existing technology, the present application proposes a fault isolation method based on a white box switch. The advantages of the present application are:
[0053] This application completes fault convergence based on the port admin up and down status, does not rely on protocol interaction, and has faster and more thorough convergence. It supports delayed up to effectively prevent frequent fault jitter scenarios;
[0054] This application does not rely on controllers or other external devices. All functional operations can be completed only on white box switches, so iterative development efficiency is high and deployment is fast;
[0055] This application supports configuring multiple monitoring groups, which are fully decoupled and easy to dynamically expand and shrink. At the same time, commands can be displayed, and the status information of each monitoring group can be viewed at a glance;
[0056] This application does not require network request interaction with the network management platform and is completed on the white box switch, which has high timeliness, greatly shortens fault isolation time, and speeds up network recovery.
[0057] This application can be flexibly added according to the network model of different scenarios. The status source of the monitoring group can also be probe results, BUFFER status and master-slave discovery, which has high flexibility. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] FIG1 is a schematic diagram of a fault isolation method based on a white box switch provided in one embodiment of the present application;
[0059] FIG2 is a schematic diagram of state transitions of three MGs provided in one embodiment of the present application;
[0060] FIG3 is a schematic diagram of message transmission logic provided by an embodiment of the present application;
[0061] FIG4 is a schematic diagram of a TOR switch MR monitoring a complete disconnection of a northbound BGP peer and shutting down a southbound port, provided by an embodiment of the present application;
[0062] FIG5 is a schematic diagram of a fault isolation device based on a white box switch provided by an embodiment of the present application;
[0063] FIG6 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application;
[0064] FIG7 is a schematic diagram of the structure of a computer-readable storage medium provided by an embodiment of the present application. DETAILED DESCRIPTION
[0065] To better understand this application, various aspects of this application will be described in more detail with reference to the accompanying drawings. It should be understood that these detailed descriptions are merely descriptions of exemplary embodiments of this application and are not intended to limit the scope of this application in any way. Throughout this specification, like reference numerals refer to like elements. The expression "and / or" includes any and all combinations of one or more of the associated listed items.
[0066] In the accompanying drawings, the size, dimensions, and shapes of elements have been slightly adjusted for ease of illustration. The drawings are for illustration only and are not drawn strictly to scale. As used herein, the terms "substantially," "approximately," and similar terms are intended to indicate approximations, not degrees, and are intended to illustrate the inherent deviations in measurements or calculations that would be recognized by those of ordinary skill in the art. Furthermore, in this application, the order in which the steps are described does not necessarily represent the order in which these steps will occur in actual operation, unless otherwise specified or inferred from the context.
[0067] It should also be understood that expressions such as "including", "comprising", "having", "containing" and / or "comprising" in this specification are open rather than closed expressions, which indicate the presence of the stated features, elements and / or components, but do not exclude the presence of one or more other features, elements, components and / or combinations thereof. In addition, when expressions such as "at least one of..." appear after a list of listed features, they modify the entire list of features rather than just the individual elements in the list. In addition, when describing embodiments of the present application, "may" is used to mean "one or more embodiments of the present application". And, the term "exemplary" is intended to refer to an example or illustration.
[0068] Unless otherwise defined, all words used herein (including engineering terms and scientific and technological terms) have the same meaning as commonly understood by those skilled in the art to which this application belongs. It should also be understood that, unless otherwise specified in this application, words defined in commonly used dictionaries should be interpreted as having the same meaning as they do in the context of the relevant technology, and should not be interpreted in an idealized or overly formal sense.
[0069] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0070] Example 1
[0071] FIG1 is a schematic diagram of a fault isolation method based on a white-box switch according to an embodiment of the present application. As shown in FIG1 , a fault isolation method based on a white-box switch includes:
[0072] Create a monitoring group on the white box switch;
[0073] The step of establishing a monitoring group on the white box switch includes:
[0074] Establish a monitoring group in the white box switch SONIC system;
[0075] Log in to the control interface of the white box switch SONIC system;
[0076] Click the "Monitoring Group" option on the control interface;
[0077] Click the "New" button and enter the name and description of the monitoring group;
[0078] Click "OK" to create the monitoring group.
[0079] Among them, the SONIC system, whose full name is Software for Open Networking in the Cloud, is an open source network operating system.
[0080] The monitoring group is configured with preset ports and data, and status data of the monitoring group is configured;
[0081] The monitoring group configures preset ports and data, including:
[0082] Use the hash type and key filed value structure of redis to configure the preset port and data, and save the preset port and data of the monitoring group configuration in CONFIG_DB.
[0083] Among them, redis is a database in the form of key-value pairs.
[0084] The use of Redis's hash type and key filed value structure to configure the preset port and data includes:
[0085] Key="MONITOR_GROUP|MGxx"
[0086] Field Value
[0087] "ports"<list_value>
[0088] "up_delay_time" <value>
[0089] "MGxx": xx is an int number that identifies the monitoring group number;
[0090] "ports": preset ports, the port group for executing shutdown and no shutdown, which can be one or more ports;
[0091] "up_delay_time": Delay operation when switching from admin down to admin up;
[0092] MG stands for Monitor Group, which means monitoring group.
[0093] The status data of the configuration monitoring group includes:
[0094] The state data of the configuration monitoring group is stored in the STATE_DB database. There are three states: init, up, and down. As shown in Figure 2, the state transitions of the three MGs are shown.
[0095] Key="MONITOR_GROUP_STATE|MGxx";
[0096] Field Value
[0097] "state" <init up down>
[0098] MG port status, stored in the STATE_DB database;
[0099] Key = "MG_PORT|{port}";
[0100] Field Value
[0101] "MonitorGroupxx" <up down>
[0102] Subscribe to MONITOR_GROUP_STATE|MGxx status and set the status of MG|port}.
[0103] Subscribing to the MONITOR_GROUP_STATE|MGxx status and setting the status of MG|port} include:
[0104] Subscribe to MONITOR_GROUP_STATE|MGxx status and set the status of MG|port}. Portmgrd determines the port admin value sent to the hardware based on the user-configured shutdown status and MG_PORT|port} data.
[0105] The steps of subscribing to the MONITOR_GROUP_STATE|MGxx state and setting the MG|port} state; and determining the port admin value sent to the hardware based on the user-configured shutdown state and MG_PORT|port} data by portmgrd include:
[0106] The white box switch is initialized, the program enters the working state, and starts subscribing to the Redis STATE_DB database MonitorGroup*;
[0107] The default state of MG is init, which is the initial state. It initializes the soft data and does not perform other operations.
[0108] Upon receiving subscription information and the MONITOR_GROUP_STATE state is down, a shutdown flag is written to each port in the preset ports list, recorded in STATE_DB, and the MG_PORT|port} of the corresponding port is set to MonitorGroupxx down;
[0109] Upon receiving subscription information and the MONITOR_GROUP_STATE state is up, the corresponding port's MG_PORT|{port} is set to MonitorGroupxx down. If the state changes from down to up, a delay process is executed. The delay time is the configured up_delay_time. This can effectively prevent network jitter caused by frequent port updowns. If the state changes from init to up, no delay is required.
[0110] The portmgr process determines the port admin status that is ultimately sent to the hardware. Based on user configuration, admin status, and the status of each monitoring group in MG_PORT|port}, if one is marked as down, the status sent to the hardware admin is down; if all are up, the status sent to the hardware admin is up.
[0111] When the white box switch detects a fault on itself or a protocol anomaly, it triggers the monitoring group to perform a batch shutdown on the preset ports to isolate the fault;
[0112] Among them, shutdown means to close.
[0113] After the fault or protocol abnormality is recovered, batch no shutdown is executed on the preset ports, and the white box switch can continue to provide services.
[0114] Among them, no shutdown means not to shut down.
[0115] Example 2
[0116] FIG5 is a schematic diagram of a fault isolation device based on a white box switch provided by an embodiment of the present application. As shown in FIG5 , a fault isolation device based on a white box switch includes:
[0117] Establish monitoring group module, used to establish monitoring groups on white box switches;
[0118] Configuration data module, used for configuring preset ports and data for the monitoring group, and configuring status data of the monitoring group;
[0119] The monitoring group executes a batch shutdown module on the preset ports, which is used to trigger the monitoring group to execute a batch shutdown module on the preset ports to isolate the fault when the white box switch detects its own fault or detects a protocol abnormality;
[0120] The preset port executes a batch no shutdown module, which is used to execute a batch no shutdown on the preset port after the failure or protocol abnormality is recovered, and the white box switch can continue to provide services.
[0121] Example 3
[0122] FIG3 is a schematic diagram of message transmission logic provided by an embodiment of the present application. As shown in FIG3 , the message transmission logic, when the TOR switch and the northbound BGP peer are all down, the specific steps of isolating all southbound ports of the device include:
[0123] Read the configuration database, read the configuration, and the status is init;
[0124] APP writes MG_STATE, writes down when a fault occurs, and writes up when recovery occurs;
[0125] The MonitorGroup process subscribes to MG status changes, triggering state machine changes;
[0126] MonitorGroup will write the status of the preset port to MG_PORT immediately or after a delay based on the status decision;
[0127] The portmgrd process receives the subscription message and determines the final hardware admin value based on the port shutdown status configured by the user and the status of each MG group in MG_port sent to this port.
[0128] TOR stands for Top of Rack (TOP of Rack), which is typically an access switch. BGP stands for Border Gateway Protocol.
[0129] Example 4
[0130] FIG4 is a schematic diagram of a TOR switch MR monitoring a complete northbound BGP peer disconnection and shutting down a southbound port, according to an embodiment of the present application. As shown in FIG4 , the TOR switch MR monitoring a complete northbound BGP peer disconnection and shutting down a southbound port specifically includes:
[0131] When a TOR switch acts as a server gateway and a northbound BGP peer is established, all northbound routes are deleted, resulting in continuous packet loss for southbound traffic. To address this issue, add a monitoring group (MonitorGroup) to monitor all northbound BGP peers, with all southbound ports as the default. If all northbound BGP peers are not in the ESTABLISHED state, set the MonitorGroup state to down, triggering an admin down on all southbound interfaces. This isolates the TOR devices, redirects server traffic to another TOR, eliminates packet loss, and restores the network.
[0132] Example 5
[0133] If the MCLAG master / standby device loses connectivity in a dual-master scenario, isolate all ports on the device except peerlink.
[0134] When MCLAG networking is in progress, the protocol between the primary and standby devices is interrupted, resulting in a dual master state and continuous packet loss on both masters. To address this failure scenario, a new connection line is added to detect the role of the peer device. This sends the local device's role information (master / standby) to the peer device. A monitoring group, MonitorGroup, is added to monitor both the peer and local roles. If both roles are master, and the local iccpd device uses a smaller IP address (previously in standby), the MonitorGroup status is set to down, triggering all non-peerlink ports to go down, isolating the original standby device and restoring the network. MCLAG stands for Mclag Multichassis Link Aggregation Group, a cross-device link aggregation group.
[0135] Example 6
[0136] Key docker and process exceptions, isolate the device.
[0137] When key Docker containers or processes exit abnormally, the device may experience various unexpected errors. For example, when SWS exits, the hardware and software forwarding tables may be inconsistent, leading to network packet loss. Add a monitoring group, MonitorGroup, to trigger an admin down on all ports when key Docker containers or processes exit abnormally, isolating the device.
[0138] Example 7
[0139] Active and standby port groups.
[0140] For scenarios where network elements operate in active / standby mode, add a monitoring group called MonitorGroup. This group monitors the primary port, with the default port designated as the backup port. When the primary port is up, the backup port is down; when the primary port is down, the backup port is up.
[0141] Example 8
[0142] Figure 6 is a schematic diagram of the structure of an electronic device provided in one embodiment of the present application. As shown in Figure 6, according to another aspect of the present application, an electronic device 500 is provided. The electronic device 500 may include one or more processors and one or more memories. The memories may store computer-readable code that, when executed by the one or more processors, may implement a white-box switch-based fault isolation method.
[0143] The method or system according to the embodiment of the present application can also be implemented with the help of the architecture of the electronic device shown in Figure 6. As shown in Figure 6, the electronic device 500 may include a bus 501, one or more CPUs 502, a read-only memory (ROM) 503, a random access memory (RAM) 504, a communication port 505 connected to the network, an input / output component 506, a hard disk 507, etc. The storage device in the electronic device 500, such as ROM 503 or hard disk 507, can store a method for isolating faults based on a white box switch provided in the present application. A method for isolating faults based on a white box switch may, for example, include: establishing a monitoring group on the white box switch; the monitoring group configures preset ports and data, and configures status data of the monitoring group; when the white box switch detects its own fault or detects a protocol anomaly, it triggers the monitoring group to perform a batch shutdown on the preset ports to isolate the fault; after the fault or protocol anomaly is recovered, a batch no shutdown is performed on the preset ports, and the white box switch can continue to provide services. Further, the electronic device 500 may also include a user interface 508. Of course, the architecture shown in FIG6 is merely exemplary. When implementing different devices, one or more components in the electronic device shown in FIG6 may be omitted according to actual needs.
[0144] Example 9
[0145] FIG7 is a schematic diagram of the computer-readable storage medium structure provided by one embodiment of the present application. As shown in FIG7 , a computer-readable storage medium 600 according to one embodiment of the present application is shown. Computer-readable instructions are stored on the computer-readable storage medium 600. When the computer-readable instructions are executed by a processor, a fault isolation method based on a white box switch according to an embodiment of the present application described with reference to the above figures can be executed. The storage medium 600 includes, but is not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, and the like.
[0146] It should be understood that the methods, apparatuses, and devices of the present application can be implemented in many ways. For example, the methods, apparatuses, and devices of the present application can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps used for the method is for illustration only, and the steps of the method of the present application are not limited to the order specifically described above unless otherwise specified. In addition, in some embodiments, the present application can also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the method according to the present application. Therefore, the present application also covers recording media that store programs for executing the method according to the present application.
[0147] In addition, the parts of the above technical solutions provided in the embodiments of the present application that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail to avoid excessive redundancy.
[0148] The above-described specific embodiments further illustrate the purpose, technical solutions, and beneficial effects of this application. It should be understood that the above description is merely a specific embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this application shall be included within the scope of protection of this application.< / up> < / init> < / value> < / up> < / init> < / value>
Claims
1. A fault isolation method based on a white box switch, characterized in that: The following steps are involved: Create a monitoring group on the white box switch; The monitoring group is configured with preset ports and data, and status data of the monitoring group is configured; When the white box switch detects a fault of itself or a protocol anomaly, it triggers the monitoring group to perform batch shutdown on the preset ports to isolate the fault; After the failure or protocol abnormality is recovered, batch no shutdown is performed on the preset ports, and the white box switch can continue to provide services.
2. The fault isolation method based on a white box switch according to claim 1, characterized in that: The step of establishing a monitoring group on the white box switch includes: Establish a monitoring group in the white box switch SONIC system; Log in to the control interface of the white box switch SONIC system; Click the "Monitoring Group" option on the control interface; Click the "New" button and enter the name and description of the monitoring group; Click the "OK" button to create the monitoring group.
3. The fault isolation method based on a white box switch according to claim 1, characterized in that: The monitoring group configures preset ports and data, including: Use the hash type and key filed value structure of redis to configure the preset port and data, and save the preset port and data of the monitoring group configuration in CONFIG_DB.
4. The fault isolation method based on a white box switch according to claim 3, characterized in that: The use of redis hash type and key filed value structure to configure the preset port and data includes: Key="MONITOR_GROUP|MGxx" Field Value "ports"<list_value> "up_delay_time" <value>< / value> "MGxx": xx is an int number, identifying the sequence number of the monitoring group; "ports": default port, the port group for executing shutdown and no shutdown, which can be one or more ports; "up_delay_time": Delay operation when switching from admin down to admin up.
5. The method for isolating faults based on a white box switch according to claim 1, characterized in that: The configuration monitoring group status data includes: The state data of the configuration monitoring group is stored in the STATE_DB database, and there are three states: init, up, and down; Key="MONITOR_GROUP_STATE|MGxx"; Field Value "state" <init up down>< / init> MG port status, stored in the STATE_DB database; Key = "MG_PORT|{port}"; Field Value "MonitorGroupxx" <up down>< / up> Subscribe to MONITOR_GROUP_STATE|MGxx status and set the status of MG|port}.
6. The method for isolating faults based on a white box switch according to claim 5, characterized in that: The subscribing to the MONITOR_GROUP_STATE|MGxx state and setting the state of MG|port} include: Subscribe to MONITOR_GROUP_STATE|MGxx status, set the status of MG|port}, and portmgrd determines the port admin value sent to the hardware based on the shutdown status configured by the user and MG_PORT|port} data.
7. The method for isolating faults based on a white box switch according to claim 6, characterized in that: The steps of subscribing to the MONITOR_GROUP_STATE|MGxx status and setting the status of MG|port}; portmgrd determining the port admin value sent to the hardware according to the shutdown status and MG_PORT|port} data configured by the user include: The white box switch is initialized, the program enters the working state, and starts subscribing to the redis STATE_DB database MonitorGroup*; The default state of MG is init, which is the initial state. It initializes the soft data and does not perform other operations. When receiving subscription information and the MONITOR_GROUP_STATE state is down, write the shutdown tag to each port in the list of preset ports, record it in STATE_DB, and set the MG_PORT|port} of the corresponding port to MonitorGroupxx down; After receiving the subscription information, when the MONITOR_GROUP_STATE state is up, set the port's corresponding MG_PORT|{port} to MonitorGroupxx down; if the STATE changes from down to up, the delay process will be executed, and the delay time is the up_delay_time in the configuration, which can effectively prevent network jitter caused by frequent updown ports; if it is from init state to up state, no delay is required; The portmgr process determines the port admin status that is ultimately sent to the hardware. Based on the user configuration and admin status, and the status of each monitoring group in MG_PORT|port}, if one is marked as down, it is sent to the hardware admin as down; if all are up, it is sent to the hardware admin as up.
8. A fault isolation device based on a white box switch, characterized in that: The device comprises: Establish monitoring group module, used to establish monitoring group on white box switch; A configuration data module is used for configuring preset ports and data for the monitoring group and configuring status data of the monitoring group; The monitoring group executes a batch shutdown module on the preset ports, which is used to trigger the monitoring group to execute a batch shutdown on the preset ports to isolate the fault when the white box switch detects its own fault or detects a protocol abnormality; The preset port executes a batch no shutdown module, which is used to execute a batch no shutdown on the preset port after the failure or protocol abnormality is recovered, and the white box switch can continue to provide services.
9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the fault isolation method based on a white box switch as described in any one of claims 1 to 7 are implemented.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the fault isolation method based on a white box switch according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and system for achieving batch management switch through improved openflow protocol
CN104486119A
Fault isolation method and device, switch and storage medium
CN115834517A
White-box switch one-way link fault detection method and system
CN115865742A
Fault isolation method based on white-box switch
CN117857486A
Software defined network whitebox infection detection and isolation
US20210067539A1