Switch port operation and maintenance method, device, system, storage medium and program product

By using Docker containers and Redis for communication in switch port links, automated and refined monitoring and maintenance of switch port links are achieved, solving the problem of low efficiency in fault diagnosis of switch port links that relies on manual operation, and improving fault handling efficiency and network stability.

CN122226730APending Publication Date: 2026-06-16SHANGHAI EVEX INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610383271.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-26
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

In the operation and maintenance of switches based on the cloud-based Software-Defined Open Networking (SONiC) operating system, switch port link fault diagnosis relies on manual operation, which is inefficient, error-prone, and lacks automated closed-loop capability, affecting network stability and efficiency.

Method used

Port diagnostic services are deployed using Docker containers and communicate with the deployment environment components via Redis to monitor the port link status of the switch in real time. Combined with multi-dimensional parameter analysis, automated and refined anomaly diagnosis and maintenance are achieved, including port status, link status, signal quality, bit error rate and traffic monitoring. The monitoring granularity and thresholds are dynamically adjusted to achieve fully automated closed-loop management.

Benefits of technology

It significantly improves the efficiency of switch port operation and maintenance, reduces the rate of human error and the impact of faults, enables rapid fault location and accurate handling, ensures the stability and reliability of network links, and enhances the level of intelligent operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122226730A_ABST
    Figure CN122226730A_ABST
Patent Text Reader

Abstract

The application provides a switch port operation and maintenance method, device, system, storage medium and program product, and relates to the technical field of communication. The method comprises the following steps: in the operation process of a switch port link, a port diagnosis service is called to obtain first operation state parameters of the switch port link, the port diagnosis service is deployed in the form of a Docker container, and communicates with components in a deployment environment through Redis, and the first operation state parameters comprise port state, link state, in-place state, signal quality, bit error rate and port traffic; whether the switch port link has an abnormal risk is determined according to the first operation state parameters, and a determination result is obtained; and if the determination result indicates that the switch port link has an abnormality, the switch port is subjected to abnormal operation and maintenance. The application can improve the operation and maintenance efficiency of the switch port.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to a method, device, system, storage medium and program product for the operation and maintenance of switch ports. Background Technology

[0002] In modern network architectures, switches serve as core communication nodes, and their operational stability directly impacts network stability and performance. During the operation and maintenance of switches based on the Software for Open Networking in the Cloud (SONiC) operating system, hardware engineers and other relevant personnel need to manually monitor the operational status of switch port links and perform manual fault diagnosis when port link failures are detected, in order to ensure the normal operation of the port links.

[0003] This process is highly dependent on manual operation, prone to errors, and inefficient. Summary of the Invention

[0004] This application provides a method, device, system, storage medium, and program product for the operation and maintenance of switch ports, in order to improve the efficiency of switch port operation and maintenance.

[0005] In a first aspect, embodiments of this application provide a method for operating and maintaining a switch port, including:

[0006] During the operation of the switch port link, the port diagnostic service is called to obtain the first operating status parameters of the switch port link. The port diagnostic service is deployed in the form of a Docker container and communicates with the components in the deployment environment through Redis. The first operating status parameters include port status, link status, in-situ status, signal quality, bit error rate, and port traffic.

[0007] Based on the first operating status parameters, determine whether there are any abnormal risks in the switch port links, and obtain the determination result;

[0008] If the results indicate that there is an anomaly in the switch port link, perform abnormal operation and maintenance on the switch port.

[0009] In one possible implementation, the deployment environment is the SONiC operating system, and the port diagnostic service is specifically used for:

[0010] Perform the following operations via inter-component communication:

[0011] Monitor changes in the port status information table in the distributed database of the SONiC operating system to obtain the port status;

[0012] Communicate with the Software Development Kit (SDK) to obtain link status and in-situ status;

[0013] It communicates with the SDK to obtain link signals, and monitors signal quality through eye diagram analysis and pseudo-random binary sequence (PRBS) testing, and monitors port traffic through serpentine flow testing.

[0014] In one possible implementation, it also includes:

[0015] If the result indicates that there is an abnormal risk in the switch port link, monitor the second operating status parameter of the switch port link. The second operating status parameter includes the first operating status parameter, and the parameter types of the second operating status parameter are more than the parameter types of the first operating status parameter.

[0016] Based on the second operating status parameters, determine whether there are any abnormalities in the switch port links;

[0017] If there is an anomaly in the switch port link, perform abnormal operation and maintenance on the switch port;

[0018] If the detected abnormal risk is eliminated, switch to monitoring the first operating status parameter.

[0019] In one possible implementation, abnormal operation and maintenance of switch ports includes:

[0020] The anomaly type was determined through cross-validation;

[0021] Perform abnormal operation and maintenance on switch ports according to the type of abnormality.

[0022] In one possible implementation, abnormal operation and maintenance for switch ports is performed according to the anomaly type, including:

[0023] Based on the fault type, retrieve the corresponding solution from the fault database;

[0024] The switch port links were repaired based on the solution.

[0025] In one possible implementation, based on the first operating state parameters, it is determined whether there is an abnormal risk in the switch port link, and the determination result is obtained, including:

[0026] If the first operating status parameter is greater than the risk threshold corresponding to the first operating status parameter, it is determined that there is an abnormal risk in the switch port link, and a determination result indicating that there is an abnormal risk in the switch port link is obtained.

[0027] If the first operating status parameter is greater than the abnormal threshold corresponding to the first operating status parameter, it is determined that there is an abnormality in the switch port link, and a determination result indicating that there is an abnormality in the switch port link is obtained.

[0028] Among them, the risk threshold and the abnormal threshold are determined by a sliding window mechanism based on historical operating status parameters within a preset time period.

[0029] In one possible implementation, it also includes:

[0030] During the switch port link initialization phase, the switch port configuration begins, and switch port information is obtained.

[0031] Based on switch port information and JSON template, automatically generate configuration files corresponding to the switch ports;

[0032] Configure the switch ports according to the configuration file.

[0033] Secondly, embodiments of this application provide a switch port maintenance device, including:

[0034] The acquisition module is used to call the port diagnostic service to obtain the first operating status parameters of the switch port link during the operation of the switch port link. The port diagnostic service is deployed in the form of a Docker container and communicates with the components in the deployment environment through Redis. The first operating status parameters include port status, link status, in-situ status, signal quality, bit error rate, and port traffic.

[0035] The determination module is used to determine whether there is any abnormal risk in the switch port link based on the first operating status parameters, and to obtain the determination result;

[0036] The operation and maintenance module is used to perform abnormal operation and maintenance on the switch port if the determination result indicates that there is an anomaly in the switch port link.

[0037] Thirdly, embodiments of this application provide a switch port maintenance system, including: a port diagnostic service and switch port maintenance equipment;

[0038] The switch port maintenance equipment interacts with the port diagnostic service to execute switch port maintenance methods such as those described in the first aspect, in order to perform switch port maintenance.

[0039] Fourthly, embodiments of this application provide a switch port maintenance device, including: a memory and a processor;

[0040] The memory stores instructions that the computer executes;

[0041] The processor executes computer execution instructions stored in memory, causing the processor to perform the methods described in the various possible implementations of the first aspect above.

[0042] Fifthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed, are used to implement the methods described in the various possible implementations of the first aspect above.

[0043] Sixthly, embodiments of this application provide a computer program product, including a computer program that, when executed, implements the methods described in the various possible implementations of the first aspect above.

[0044] The switch port operation and maintenance method, device, system, storage medium, and program product provided in this application embodiment, during the operation of the switch port link, calls a port diagnostic service to obtain the first operating status parameters of the switch port link. The port diagnostic service is deployed in the form of a Docker container and communicates with components in the deployment environment through Redis. The first operating status parameters include port status, link status, presence status, signal quality, bit error rate, and port traffic. Based on the first operating status parameters, it is determined whether there is an abnormal risk in the switch port link, and a determination result is obtained. If the determination result indicates that there is an abnormality in the switch port link, abnormal operation and maintenance is performed on the switch port. Among them, the port diagnostic service deployed in the form of a Docker container has good device compatibility. It automatically obtains the first operating parameters of the switch port link and automatically performs abnormal diagnosis and abnormal operation and maintenance on the switch port link, thereby improving the efficiency of switch port operation and maintenance. Attached Figure Description

[0045] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0046] Figure 1 This is a schematic diagram of the switch port operation and maintenance method provided in the embodiments of this application;

[0047] Figure 2 This is a schematic diagram of the switch port operation and maintenance architecture provided in the embodiments of this application;

[0048] Figure 3 This is a schematic diagram of the structure of the switch port maintenance device provided in the embodiments of this application;

[0049] Figure 4 A schematic diagram of a switch port maintenance system provided in this application embodiment;

[0050] Figure 5 This is a schematic diagram of the structure of the switch port maintenance equipment provided in the embodiments of this application.

[0051] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0052] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0053] In related technologies, if an anomaly occurs in a switch port link after the switch device starts up, engineers need to manually check multiple related diagnostic commands such as SERDE link signals, Bit Error Rate (BER), and Pseudo-Random Binary Sequence (PRBS), analyze the output logs item by item, and locate the root cause of the problem. This process is time-consuming and heavily reliant on the engineer's experience. Even if the problem is located (such as polarity error or insufficient pre-emphasis), the configuration file still needs to be manually modified and the port restarted for repeated verification. The lack of an integrated automated closed-loop capability of "detection-diagnosis-repair-verification" seriously affects product debugging efficiency and deployment speed. In addition, the configuration file needs to be written during the switch device startup process, which is prone to errors. When customers raise new requirements, engineers need to manually modify these files, involving the mapping relationship between logical ports and physical ports and front panel ports, port speed, error correction mode, pre-emphasis parameters, and other configurations. The modification process is tedious and prone to introducing errors.

[0054] To address the aforementioned technical issues, the switch port maintenance method provided in this application calls the port diagnostic service to automatically diagnose and repair anomalies in the switch port links, thereby improving the efficiency of switch port maintenance.

[0055] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0056] Figure 1 This is a schematic diagram of the switch port operation and maintenance method provided in an embodiment of this application. Figure 1As shown in the figure, this application provides a method for operating and maintaining a switch port, including:

[0057] S101. During the operation of the switch port link, the port diagnostic service is called to obtain the first operating status parameters of the switch port link. The port diagnostic service is deployed in the form of a Docker container and communicates with the components in the deployment environment through Redis. The first operating status parameters include port status, link status, in-situ status, signal quality, bit error rate, and port traffic.

[0058] Specifically, Docker containers are a lightweight, portable, and self-contained software packaging technology. They package an application and its required code, dependencies, environment variables, configuration files, etc., into a standardized unit for execution in an isolated runtime environment. Compared to traditional virtual machines, Docker containers share the host operating system kernel, resulting in faster startup, lower resource consumption, and higher deployment efficiency. The port diagnostic service is deployed using Docker containers, enabling environment isolation, flexible deployment, convenient version control, and easy start-up, migration, and scaling, while ensuring service independence and stability. Using Redis as middleware to communicate with other components in the environment enables data decoupling, asynchronous interaction, and high-concurrency access, improving system scalability, maintainability, and operational efficiency, and facilitating unified monitoring and troubleshooting.

[0059] By deploying port diagnostic services in Docker containers and leveraging Redis for inter-component communication, automated and real-time monitoring and data collection of switch port link status can be achieved. Compared to traditional manual maintenance, this not only significantly reduces the workload and error rate of manual troubleshooting and enables uninterrupted status monitoring around the clock, but also allows for rapid fault location and quantitative performance analysis by acquiring accurate parameters from multiple dimensions such as port physical status, link connectivity, signal quality, bit error rate, and traffic statistics. This improves maintenance efficiency and problem-solving speed, while providing reliable data support for standardized and intelligent network maintenance management, ensuring the stable and reliable operation of switch port links.

[0060] S102. Based on the first operating status parameters, determine whether there is any abnormal risk in the switch port link and obtain the determination result.

[0061] Based on the acquired first operating status parameters, including key indicators such as port status, link status, presence status, signal quality, bit error rate, and port traffic, these parameters are comprehensively analyzed and compared with thresholds. Through preset diagnostic rules and anomaly judgment logic, it is possible to accurately determine whether there are problems such as physical connection abnormalities, link interruptions, signal degradation, excessive bit error rate, or abnormal traffic fluctuations in the switch port links. This enables automated and refined assessment of the health status of port links, providing a reliable basis for timely detection and location of network faults and ensuring stable link operation.

[0062] S103. If the result indicates that there is an abnormality in the switch port link, perform abnormal operation and maintenance on the switch port.

[0063] If an anomaly is detected in the switch port link, the corresponding anomaly operation and maintenance handling process will be automatically triggered according to the anomaly type and severity. This includes, but is not limited to, automatically restarting the abnormal port, resetting the link connection, adjusting port signal parameters, optimizing traffic scheduling strategies, or pushing detailed alarm information and anomaly location results to the operation and maintenance platform. This enables rapid response, automated processing, and closed-loop management of port link anomalies, effectively reducing the impact of faults on network services and improving the overall stability and reliability of network operation.

[0064] The switch port operation and maintenance method provided in this application deploys the port diagnostic service in the form of a Docker container and uses Redis to realize inter-component communication. Combined with the real-time collection, comprehensive analysis and anomaly judgment of the first operating status parameters of the switch port link, it can realize automated and refined monitoring and anomaly operation and maintenance of the port link operating status. Compared with traditional manual operation and maintenance, it not only greatly improves the efficiency of fault detection and handling, reduces labor costs and error rate, but also realizes accurate fault location and quantitative analysis through multi-dimensional parameters, realizes automated and closed-loop management of anomaly handling, effectively ensures the stable operation of the switch port link, and improves the intelligence level and reliability of the overall network operation and maintenance.

[0065] In one possible implementation, the deployment environment is the SONiC operating system, and the port diagnostic service is specifically used for:

[0066] Perform the following operations via inter-component communication:

[0067] Monitor changes in the port status information table in the distributed database of the SONiC operating system to obtain port status; communicate with the Software Development Kit (SDK) to obtain link status and in-situ status; communicate with the SDK to obtain link signals, and monitor signal quality through eye diagram analysis and pseudo-random binary sequence testing, and monitor port traffic through serpentine flow testing.

[0068] Specifically, the port diagnostic service is deployed as a Docker container and communicates with other SONiC components via Redis. The Docker containers added to the port diagnostic service can monitor port status by sharing state with the distributed database StateDB and listening for changes in the PORT_TABLE table. The SDK obtains the SERDE link status, as well as the presence status of the loopback and optical modules. The SDK output can be used to view the signal strength of each lane, analyze signal quality through eye diagrams, check the bit error rate of data transmission using PRBS's BIR, and perform snake-flow testing to determine if there is packet loss or continuity issues on the switch ports. These tests help determine if the ports are in a normal operating state.

[0069] The switch port operation and maintenance method provided in this application, through the collaborative communication between components, can, on the one hand, monitor changes in the port status information table in the distributed database of the SONiC operating system in real time to accurately obtain the port status; on the other hand, it can obtain the link status and presence status by interacting with the SDK, and at the same time, it can obtain the link signal with the help of the SDK and perform eye diagram analysis and pseudo-random binary sequence testing to achieve fine monitoring of signal quality. Then, it can complete the real-time monitoring of port traffic through the snake flow test method. This integrated data acquisition and analysis mechanism, compared with the traditional decentralized and manual operation and maintenance method, can more comprehensively, accurately and in real time grasp the operating status of switch port links, realize early detection and quick location of anomalies, improve the automation and intelligence level of operation and maintenance, and ensure the stable and efficient operation of network links.

[0070] In one possible implementation, the switch port maintenance method further includes:

[0071] If the result indicates that there is an abnormal risk in the switch port link, monitor the second operating status parameter of the switch port link. The second operating parameter includes the first operating parameter, and the parameter types of the second operating status parameter are more than those of the first operating status parameter. Based on the second operating status parameter, determine whether there is an abnormality in the switch port link. If there is an abnormality in the switch port link, perform abnormal operation and maintenance on the switch port. If the abnormal risk is detected and eliminated, switch the monitoring object to the first operating status parameter.

[0072] Specifically, if the first operating status parameter predicts a potential anomaly risk in the switch port link, the monitoring granularity is automatically increased by switching to and collecting and analyzing the second operating status parameter, which has a richer set of parameters. This allows for in-depth diagnosis of the port link status using more comprehensive and refined monitoring data. Subsequently, based on the second operating status parameter, comprehensive analysis and threshold verification are performed to accurately determine whether there is an actual anomaly in the link. If an anomaly is confirmed, a targeted anomaly operation and maintenance process is automatically triggered, including port reset, link optimization, alarm push, and automated repair. This process continues until the anomaly risk in the port link is completely eliminated. Then, the monitoring object is switched back to the relatively simple first operating status parameter. This achieves fully automated, closed-loop management from risk warning, in-depth monitoring, anomaly determination, proactive operation and maintenance to restoration to normal operation, effectively improving the operational stability and intelligent operation and maintenance level of the switch port link.

[0073] In one implementation, the second operational status parameter serves as a supplement and enhancement to the first parameter, encompassing richer and more granular monitoring indicators. In addition to the basic items already covered by the first parameter, such as port status, link status, in-situ status, signal quality, bit error rate, and port traffic, it also adds hardware operational parameters such as port optical power transmission and reception values, link latency and jitter, packet loss rate, packet error rate, frame error statistics, link negotiation status, port buffer utilization, SerDes link signal eye diagram parameters, temperature, and voltage, as well as in-depth operational indicators such as link retransmission count, error frame count, number of port auto-negotiation failures, and number of link oscillations. This comprehensively covers various operational data from the physical layer, data link layer, to the service layer, providing more comprehensive data support for anomaly detection and in-depth diagnosis.

[0074] The switch port operation and maintenance method provided in this application can achieve refined and intelligent management of switch port links through hierarchical monitoring and dynamic operation and maintenance mechanisms. When the first operating status parameter predicts a potential abnormal risk, the monitoring granularity can be automatically increased. In-depth diagnosis can be carried out by collecting more diverse second operating status parameters. Combined with comprehensive analysis and threshold verification, the actual abnormality can be accurately determined. Then, targeted operation and maintenance processes such as port reset, link optimization, alarm push and automated repair are triggered until the abnormal risk is completely eliminated. Then, the method switches back to the simple and efficient first parameter monitoring mode. This forms a fully automated closed-loop management from risk warning, in-depth monitoring, abnormality determination, proactive operation and maintenance to restoration to normal. This not only greatly improves the accuracy and efficiency of abnormality handling and reduces the cost of manual intervention and the scope of fault impact, but also effectively ensures the continuous and stable operation of switch port links, and significantly improves the intelligence level and reliability of overall network operation and maintenance.

[0075] In one possible implementation, abnormal operation and maintenance of switch ports includes: determining the abnormality type through cross-validation; and performing abnormal operation and maintenance on the switch ports according to the abnormality type.

[0076] After determining the anomalies in the switch port links, a multi-dimensional, multi-source data cross-validation mechanism is used to comprehensively compare and logically verify the identified anomalies, further eliminating interference information and confirming the specific type and cause of the anomaly, ensuring the accuracy and uniqueness of the anomaly judgment. Subsequently, based on the precisely located anomaly type, targeted maintenance operations are performed, including port restart, link reset, parameter optimization, traffic scheduling, alarm reporting, or hardware reset, etc., to achieve precise and differentiated handling of switch port anomalies, avoid blind maintenance, improve the efficiency and effectiveness of fault handling, and ensure that the switch port links are quickly restored to normal operation.

[0077] The switch port operation and maintenance method provided in this application, through multi-dimensional and multi-source data cross-verification of switch port link anomalies, can effectively eliminate interference information, accurately confirm the anomaly type and cause, and significantly improve the accuracy and reliability of anomaly judgment. On this basis, targeted operation and maintenance measures are performed according to the specific anomaly type, including port restart, link reset, parameter tuning, traffic scheduling, alarm reporting and hardware reset, etc. This avoids the drawbacks of traditional blind operation and maintenance, greatly improves the efficiency and effectiveness of fault handling, promotes the rapid recovery of switch port links to normal operation, effectively reduces the impact of anomalies on network services, and ensures the stability and reliability of the overall network operation.

[0078] In one possible implementation, abnormal operation and maintenance of switch ports is performed according to the anomaly type, including: obtaining the solution corresponding to the fault type from the fault database according to the fault type; and performing abnormal repair on the switch port link according to the solution.

[0079] Specifically, problem-solving documents can be compiled into a fault database, and appropriate actions can be matched based on diagnostic results. For example, temporary configurations can be written to the AppDB; if the port's on-state does not match the actual status, the logical and physical correspondence of the port will be modified. When eye diagrams, PRBS, or signal quality are outside the normal range, or a port fails to start, cross-validation is automatically performed to accurately diagnose the fault location. After modifying the configuration, it is determined whether the fault is fixed. If not, the problem is re-analyzed and the cause is re-generated, and a problem record is generated.

[0080] Once a switch port link malfunctions and the fault type is determined, a standardized and refined solution corresponding to that fault type is precisely matched from a pre-built and continuously updated fault database. This database integrates handling procedures, parameter configuration suggestions, operation and maintenance specifications, and historical successful cases for various typical faults, ensuring the relevance and feasibility of the solution. Subsequently, according to the obtained solution, automated fault repair operations are performed on the switch port link, including parameter reset, link renegotiation, port restart, traffic scheduling, hardware status calibration, and other targeted measures. This achieves precise and standardized fault repair, effectively improving the success rate and efficiency of fault repair, shortening service interruption time, and ensuring that the switch port link quickly returns to a stable operating state. At the same time, the accumulated repair experience continuously feeds back into the fault database, forming a closed-loop optimization of operation and maintenance data.

[0081] In one implementation, after executing a single repair plan, the port status is obtained through the SDK, the port signal information, and the outputs such as PRBS and eye diagram are viewed to see if the problem has been recovered. If not, the configuration file is replaced with the previous file, the cause of the problem is re-diagnosed and analyzed, and the repair operation is performed.

[0082] The switch port operation and maintenance method provided in this application accurately matches corresponding standardized and refined solutions from a continuously updated fault database based on the determined fault type. It then uses these solutions to perform automated anomaly repair on switch port links, effectively avoiding blind handling and improving the targeting and success rate of fault repair. Simultaneously, standardized operations such as parameter resetting, link renegotiation, and port restarting quickly restore link stability, shortening service interruption time and ensuring continuous and reliable network operation. Furthermore, the experience accumulated during the repair process can feed back into the fault database, enabling continuous optimization and improvement of the operation and maintenance solution. This forms a closed-loop management model of fault diagnosis, repair, and iteration, further enhancing the overall intelligence and efficiency of operation and maintenance.

[0083] In one possible implementation, determining whether a switch port link has an abnormal risk based on a first operating state parameter, and obtaining a determination result, includes: if the first operating state parameter is greater than the risk threshold corresponding to the first operating state parameter, then determining that the switch port link has an abnormal risk, and obtaining a determination result indicating that the switch port link has an abnormal risk; if the first operating state parameter is greater than the abnormal threshold corresponding to the first operating state parameter, then determining that the switch port link has an abnormality, and obtaining a determination result indicating that the switch port link has an abnormality. The risk threshold and the abnormality threshold are determined using a sliding window mechanism based on historical operating state parameters within a preset time period.

[0084] When the real-time value of the first operating status parameter exceeds its corresponding anomaly threshold, an anomaly is determined to exist in the switch port link. This anomaly threshold is not a fixed value, but rather dynamically calculated using a sliding window mechanism combined with historical operating parameter data collected over a preset time period. This approach not only reflects the actual changing trends of the link operation but also effectively filters out interference from instantaneous fluctuations, improving the accuracy and rationality of anomaly detection. It avoids false alarms or missed alarms caused by unreasonable fixed threshold settings, laying a reliable data foundation for subsequent accurate anomaly maintenance and fault handling. The determination of the risk threshold follows the same principle as the anomaly threshold. When the switch port link is operating normally, if a single first operating status parameter exceeds the risk threshold, an anomaly risk is determined, and more comprehensive monitoring of the switch port link should be conducted.

[0085] The determination of risk and anomaly thresholds relies on a sliding window mechanism combined with historical parameters for dynamic calculation. The specific process is as follows: First, a reasonable time window length and sliding step size are set. Based on the historical data of the first operating state parameters collected within the preset time period, historical data samples are extracted segment by segment according to the sliding window method. Then, the historical parameter data within each window is statistically analyzed to calculate key statistical indicators such as mean, standard deviation, and quantiles. Combined with the normal fluctuation range of network link operation and business needs, a reasonable warning boundary is set. As time goes by, the window continues to slide backward, continuously incorporating new historical data and removing outdated data. The statistical analysis results are dynamically updated and the warning threshold is adjusted synchronously to ensure that the threshold always fits the operating characteristics of the link at different times. This ensures both sensitive detection of anomalies and effective filtering of interference from normal fluctuations, thereby improving the accuracy of anomaly judgment.

[0086] Furthermore, setting a reasonable time window length and sliding step size requires comprehensive consideration of the service fluctuation characteristics of the switch port links, data collection frequency, and anomaly response time. Regarding the time window length, priority should be given to selecting a duration covering the entire service cycle. Too short a window may result in insufficient sample size and an inability to reflect real-time operational patterns, while too long a window may cause historical data lag and difficulty in adapting to real-time changes in link status. Typically, a fixed period of 5 minutes, 15 minutes, or 1 hour can be selected as the basic window based on the alternation of peak and off-peak periods, while reserving flexibility for adjustment. For links with drastic fluctuations, the window length can be appropriately shortened, while for links with stable operation, it can be appropriately extended to improve data stability. The sliding step size needs to balance data update efficiency and computational resource consumption. An excessively large step size will lead to delayed threshold updates and untimely anomaly detection, while an excessively small step size will increase the burden of repetitive calculations and reduce processing efficiency. Generally, it is set to 1 / 4 to 1 / 2 of the window length, ensuring both timely dynamic adjustment of thresholds and reducing redundant calculations. This ensures that the anomaly threshold and risk threshold always closely match the real-time operational characteristics of the link, improving the accuracy of anomaly detection and the effectiveness of maintenance and repair.

[0087] The switch port operation and maintenance method provided in this application replaces the traditional fixed threshold with dynamically calculated risk thresholds and anomaly thresholds. By using a sliding window mechanism combined with historical parameters for real-time adjustment, the warning standard always aligns with the actual operating trend of the switch port link. This effectively filters out interference caused by instantaneous fluctuations and significantly improves the accuracy and rationality of anomaly judgment. It reduces false alarms and missed alarms caused by unreasonable threshold settings at the root, providing solid data support for subsequent accurate anomaly operation and maintenance, rapid fault location, and efficient problem handling, thus ensuring the stability and reliability of network link operation.

[0088] In one possible implementation, the switch port operation and maintenance method further includes: during the switch port link initialization phase, responding to the start of switch port configuration and obtaining switch port information; automatically generating a configuration file corresponding to the switch port based on the switch port information and a JSON template; and configuring the switch port according to the configuration file.

[0089] During the switch port link initialization phase, the system first responds to the switch port configuration start command, actively acquiring basic port information, including core parameters such as speed, laneswap, polarity, and port mapping. Then, based on the acquired port information and a pre-defined JSON configuration template, it automatically parses the template structure and fills in the corresponding parameters, quickly generating a standardized configuration file adapted to the current port scenario. Finally, following the instructions and parameters in the configuration file, it performs fine-grained port configuration operations on the switch port, including link activation, parameter validation, service binding, and permission configuration, completing the entire configuration process for port link initialization.

[0090] The switch port operation and maintenance method provided in this application embodiment can realize automated, standardized and refined management of switch port configuration. By actively acquiring and parsing the core parameters of the port, automatically generating the adaptation configuration file, and executing the full-process configuration implementation, it greatly reduces the tedious operation and error risk of manual configuration, improves configuration efficiency and consistency, ensures the rapid and stable initialization of port links, and facilitates subsequent operation and maintenance and unified policy management.

[0091] Figure 2 This is a schematic diagram of the switch port operation and maintenance architecture provided in an embodiment of this application. Figure 2 As shown, the core consists of a port diagnostic service, an SDK interaction container, a switch status service, a configuration database, a status database, an application database, a switch chip, and a front-end interface. All components interact with each other via Redis. Both the port diagnostic service and the switch status service communicate with the configuration database and status database through Redis. The SDK interaction container also interfaces with the switch chip, forming a complete closed loop from anomaly diagnosis, configuration management, status monitoring to front-end display.

[0092] Figure 3 This is a schematic diagram of the structure of the switch port maintenance device provided in the embodiments of this application, as shown below. Figure 3 The embodiment of this application provides a switch port maintenance device 30, including:

[0093] The acquisition module 301 is used to call the port diagnostic service to obtain the first operating status parameters of the switch port link during the operation of the switch port link. The port diagnostic service is deployed in the form of a Docker container and communicates with the components in the deployment environment through Redis. The first operating status parameters include port status, link status, in-situ status, signal quality, bit error rate and port traffic.

[0094] The determination module 302 is used to determine whether there is any abnormal risk in the switch port link based on the first operating status parameters, and to obtain the determination result;

[0095] The operation and maintenance module 303 is used to perform abnormal operation and maintenance on the switch port if the determination result indicates that there is an abnormality in the switch port link.

[0096] In one possible implementation, the deployment environment is the SONiC operating system, and the port diagnostic service is specifically used for:

[0097] Perform the following operations via inter-component communication:

[0098] Monitor changes in the port status information table in the distributed database of the SONiC operating system to obtain the port status;

[0099] Communicate with the SDK to obtain link status and in-situ status;

[0100] It communicates with the SDK to obtain link signals, and monitors signal quality through eye diagram analysis and pseudo-random binary sequence (PRBS) testing, and monitors port traffic through serpentine flow testing.

[0101] In one possible implementation, the operation and maintenance module 303 is further configured to: if the determination result indicates that there is an abnormal risk in the switch port link, monitor the second operating status parameter of the switch port link, wherein the parameter types of the second operating status parameter are more than the parameter types of the first operating status parameter;

[0102] Based on the second operating status parameter, determine whether there is an anomaly in the switch port link; if there is an anomaly in the switch port link, perform abnormal operation and maintenance on the switch port; if the detected anomaly risk is eliminated, switch the monitoring object to the first operating status parameter.

[0103] In one possible implementation, the operation and maintenance module 303 is specifically used to: determine the anomaly type through cross-validation; and perform anomaly operation and maintenance on the switch port according to the anomaly type.

[0104] In one possible implementation, the operation and maintenance module 303 is specifically used to: obtain the solution corresponding to the fault type from the fault database according to the fault type; and perform abnormal repair on the switch port link according to the solution.

[0105] In one possible implementation, the determining module 302 is specifically used to: if the first operating state parameter is greater than the risk threshold corresponding to the first operating state parameter, determine that there is an abnormal risk in the switch port link, and obtain a determination result indicating that there is an abnormal risk in the switch port link; if the first operating state parameter is greater than the abnormal threshold corresponding to the first operating state parameter, determine that there is an abnormality in the switch port link, and obtain a determination result indicating that there is an abnormality in the switch port link. The risk threshold and the abnormal threshold are determined using a sliding window mechanism based on historical operating state parameters within a preset time period.

[0106] In one possible implementation, a configuration module is also included, which is used to: obtain switch port information in response to the start of switch port configuration during the switch port link initialization phase; automatically generate a configuration file corresponding to the switch port based on the switch port information and JSON template; and configure the switch port according to the configuration file.

[0107] The switch port maintenance device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0108] Figure 4 This is a schematic diagram of a switch port maintenance system provided in an embodiment of this application. Figure 4 As shown, this application embodiment provides a switch port operation and maintenance system 40, including: a port diagnostic service 401 and a switch port operation and maintenance device 402;

[0109] The switch port maintenance device 402 interacts with the port diagnostic service 401 to execute the switch port maintenance method as described in the above embodiment, so as to realize the port maintenance of the switch device.

[0110] Figure 5 This is a schematic diagram of the structure of the switch port maintenance equipment provided in the embodiments of this application, as shown below. Figure 5As shown, the switch port maintenance device 50 provided in this embodiment includes at least one processor 501 and a memory 502. Optionally, the switch port maintenance device 50 further includes a communication interface 503. The processor 501, memory 502, and communication interface 503 are connected via a communication bus 505.

[0111] In a specific implementation, at least one processor 501 executes computer execution instructions stored in memory 502, causing at least one processor 501 to perform the above-described method.

[0112] The specific implementation process of processor 501 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0113] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0114] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0115] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0116] This application also provides a computer program product, including a computer program that, when executed, implements the above-described method.

[0117] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed, implement the above-described method.

[0118] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0119] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an application-specific integrated circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0120] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0121] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0122] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0123] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0124] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0125] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A method for operating and maintaining a switch port, characterized in that, include: During the operation of the switch port link, the port diagnostic service is called to obtain the first operating status parameters of the switch port link. The port diagnostic service is deployed in the form of a Docker container and communicates with the components in the deployment environment through Redis. The first operating status parameters include port status, link status, in-situ status, signal quality, bit error rate, and port traffic. Based on the first operating status parameters, determine whether there is any abnormal risk in the switch port link, and obtain the determination result; If the determination result indicates that there is an anomaly in the switch port link, abnormal operation and maintenance shall be performed on the switch port.

2. The switch port operation and maintenance method according to claim 1, characterized in that, The deployment environment is based on the SONiC operating system, a software-defined open network within the cloud. The port diagnostic service is specifically used for: Perform the following operations via inter-component communication: Monitor changes in the port status information table in the distributed database of the SONiC operating system to obtain the port status; Communicate with the Software Development Kit (SDK) to obtain link status and in-situ status; It communicates with the Software Development Kit (SDK) to obtain link signals, and monitors signal quality through eye diagram analysis and pseudo-random binary sequence (PRBS) testing, and monitors port traffic through serpentine flow testing.

3. The switch port operation and maintenance method according to claim 1, characterized in that, Also includes: If the determination result indicates that there is an abnormal risk in the switch port link, monitor the second operating status parameter of the switch port link. The second operating status parameter includes the first operating status parameter, and the parameter types of the second operating status parameter are more than the parameter types of the first operating status parameter. Based on the second operating status parameter, determine whether there is an anomaly in the switch port link; If there is an anomaly in the link of the switch port, perform abnormal operation and maintenance on the switch port; If the detected abnormal risk is eliminated, switch to monitoring the first operating status parameter.

4. The switch port operation and maintenance method according to any one of claims 1 to 3, characterized in that, The abnormal operation and maintenance of the switch port includes: The anomaly type was determined through cross-validation; Perform abnormal operation and maintenance on the switch port according to the abnormality type.

5. The switch port maintenance method according to claim 4, characterized in that, The abnormal operation and maintenance for the switch port based on the abnormality type includes: Based on the fault type, retrieve the corresponding solution from the fault database; The switch port link was repaired according to the solution.

6. The switch port operation and maintenance method according to any one of claims 1 to 3, characterized in that, The step of determining whether there is an abnormal risk in the switch port link based on the first operating status parameter, and obtaining the determination result, includes: If the first operating status parameter is greater than the risk threshold corresponding to the first operating status parameter, then it is determined that there is an abnormal risk in the switch port link, and a determination result indicating that there is an abnormal risk in the switch port link is obtained. If the first operating status parameter is greater than the abnormal threshold corresponding to the first operating status parameter, it is determined that there is an abnormality in the switch port link, and a determination result indicating that there is an abnormality in the switch port link is obtained. The risk threshold and the anomaly threshold are determined using a sliding window mechanism based on historical operating status parameters within a preset time period.

7. The switch port operation and maintenance method according to any one of claims 1 to 3, characterized in that, Also includes: During the switch port link initialization phase, the switch port configuration begins, and switch port information is obtained. Based on switch port information and JSON template, automatically generate configuration files corresponding to the switch ports; Configure the switch ports according to the configuration file.

8. A switch port maintenance device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1 to 7.

9. A switch port maintenance system, characterized in that, include: Port diagnostic services and switch port maintenance equipment; The switch port maintenance device interacts with the port diagnostic service to execute the switch port maintenance method as described in any one of claims 1 to 7, so as to realize the maintenance of the switch port.

10. A computer-readable storage medium / computer program product, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed, are used to implement the method as described in any one of claims 1 to 7; The computer program product includes a computer program that, when executed, implements the method according to any one of claims 1 to 7.