Method for testing reliability of switch
By regulating environmental data and traffic in the switch test room and combining the fault prediction model, the problem that switch reliability testing in the prior art cannot fully simulate the actual environment is solved, and efficient reliability testing and preventive maintenance of the switch are achieved.
Patent Information
- Application Number
- CN202510145367.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-10
AI Technical Summary
Existing switch reliability testing methods cannot fully simulate complex environmental conditions and dynamic load changes in actual operation, resulting in lag in fault discovery, affecting maintenance efficiency and response speed.
By setting up a switch testing room that can regulate environmental data and traffic, determine the preliminary test plan according to the switch's working specifications, restore the switch to its initial state, record baseline data, adjust traffic and environmental data, establish a fault prediction model, predict future fault types and locations, and issue a warning to notify the maintenance personnel.
The performance test of the switch in complex environments is realized, the authenticity and accuracy of the test is improved, the failure type and location can be predicted in advance, and the reliability and maintenance efficiency of the switch are significantly improved.
Smart Images

Figure CN120034473A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of network equipment testing, in particular to a method for testing the reliability of a switch. Background Art
[0002] In the field of network communications, switches are key nodes that connect different network devices, and their reliability directly affects the stability and performance of the entire network system. With the rapid development of information technology and the increasing demand for data transmission by enterprises, the reliability and stability of switches have become important topics in research and application.
[0003] Traditional switch reliability testing is mainly conducted under static conditions, that is, fixed parameters such as temperature and humidity are set in a laboratory environment, and specific traffic patterns are injected manually to evaluate the basic performance of the switch. However, such testing methods cannot fully simulate the various complex environmental conditions and dynamic load changes that may be encountered in actual operation. In addition, early testing methods mostly rely on post-analysis and lack a prediction mechanism, resulting in delayed fault discovery, which affects maintenance efficiency and response speed. Summary of the invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method for testing the reliability of a switch to solve the problem that the lack of a prediction mechanism leads to delayed fault discovery, affecting maintenance efficiency and response speed.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a method for testing the reliability of a switch, comprising:
[0008] Set up a switch test room that can regulate environmental data and traffic, and determine the preliminary test plan for switch performance based on the switch's operating specifications;
[0009] Restore the switch to its initial state and record the performance data of the switch in normal state as baseline data;
[0010] Adjust the switch traffic and environment data according to the switch performance preliminary test plan, record the switch performance index data, and compare it with the baseline data to obtain the preliminary performance analysis results of the switch;
[0011] Integrate the switch's performance indicator data and preliminary performance analysis results into test data and perform preprocessing;
[0012] According to the preprocessed test data, the performance characteristics are extracted, and a switch fault prediction model is established based on the extracted performance characteristics to obtain the type of fault that will occur in the future of the switch, and the location of the switch fault is further determined according to the type of fault that will occur in the future of the switch;
[0013] Based on the location of the switch fault, a warning is issued and maintenance personnel are notified to take corresponding measures in advance.
[0014] As a preferred solution of the method for testing the reliability of a switch of the present invention, a switch test room capable of regulating environmental data and traffic is set up, and a preliminary test plan for the performance of the switch is determined according to the working specification of the switch, which specifically includes the following steps:
[0015] Choose a completely closed test room as the location for testing the switch and install equipment that can independently control environmental data;
[0016] The environmental data includes temperature, humidity, vibration and electromagnetic interference;
[0017] Set the time series change curves of different environmental data through LabVIEW;
[0018] Determine the switch operating specifications through switch manufacturers and industry standards, and determine the preliminary test plan for switch performance based on the operating specifications;
[0019] The preliminary test plan includes specific ranges of flow and environmental data, as well as the duration of each environmental condition and its cycle pattern;
[0020] The cycle mode refers to a mode in which the switch is continuously cycled through high and low temperature alternating cycles, damp heat cycles, vibration tests, and EMI tests by adjusting environmental data.
[0021] As a preferred solution of the method for testing the reliability of a switch of the present invention, the switch is restored to an initial state and the performance data of the switch in a normal state is recorded as baseline data, which specifically includes the following steps:
[0022] Execute the factory reset command through the switch's management interface;
[0023] Run the switch under normal environment data and record the switch performance indicators after running under normal environment data through NetFlow and SNMP;
[0024] The performance indicators of the switch are CPU utilization, memory occupancy, throughput, packet loss rate and delay;
[0025] The switch performance indicators running under normal environment data are sorted and saved by timestamp to form a baseline data set and saved in CSV format.
[0026] As a preferred solution of the method for testing the reliability of a switch of the present invention, the flow and environmental data of the switch are adjusted according to the preliminary test plan of the switch performance, the performance index data of the switch is recorded, and compared with the baseline data to obtain the preliminary performance analysis result of the switch, which specifically includes the following steps:
[0027] Based on the preliminary switch performance test plan, use a device that can independently control the environmental data to adjust the switch's environmental data, and execute the loop mode until the loop mode ends;
[0028] According to the preliminary switch performance test plan, use a traffic tester to inject traffic into the switch and record the difference between the actual sending rate and receiving rate at each stage;
[0029] Use NetFlow and SNMP to collect switch environment data and traffic adjusted performance indicators, and compare them with the baseline data set item by item to obtain the difference between the adjusted performance indicators and the baseline data set;
[0030] Set performance thresholds based on historical test data. When the difference between each adjusted performance indicator and the baseline data set is greater than the performance threshold, mark this performance indicator as abnormal.
[0031] Use Plotly to draw a performance graph for each adjusted performance indicator, and distinguish the performance indicators marked as abnormal in red, which is the preliminary performance analysis result of the switch.
[0032] As a preferred solution of the method for testing the reliability of a switch of the present invention, the performance index data and the preliminary performance analysis results of the switch are integrated into test data and preprocessed, which specifically includes the following steps:
[0033] Combine the adjusted performance indicators of switch environment data and traffic collected by NetFlow and SNMP with the performance graph to form test data;
[0034] Based on the test data, statistical methods are used to identify and remove data points that deviate from the normal range and fill in missing values;
[0035] Standardize the test data.
[0036] As a preferred solution of the method for testing the reliability of a switch of the present invention, the method comprises the following steps: extracting performance characteristics according to the preprocessed test data, establishing a switch fault prediction model based on the extracted performance characteristics, obtaining the type of fault that will occur in the future of the switch, and further determining the location of the switch fault according to the type of fault that will occur in the future of the switch,
[0037] Based on the preprocessed test data, information gain analysis is used to extract the CPU utilization, memory occupancy, throughput, packet loss rate and delay of the test data as switch features;
[0038] Define the switch fault type based on switch characteristics;
[0039] For each switch feature, calculate its mean and standard deviation in the baseline dataset;
[0040] Based on the mean and standard deviation of each switch feature, an information filtering function is introduced to obtain the degree of deviation of the performance feature from the normal state;
[0041] Calculate the weighted sum of each fault type of the switch, and apply the softmax function and the deviation value of the performance characteristics from the normal state to obtain the probability distribution of future faults of the switch, which is expressed as:
[0042]
[0043] Among them, Z j represents the probability of the jth fault type occurring in the future, m represents the total number of fault types, n represents the total number of performance characteristics, and w ij represents the weight factor of the i-th performance characteristic for the j-th fault type, F i represents the value of the i-th performance feature, φ() represents the information filtering function, and k represents the index of the traversed fault type;
[0044] According to the probability distribution of future faults of the switch, the fault type with the highest probability is taken as the fault type that will occur in the future of the switch;
[0045] Use diagnostic tools and troubleshooting techniques to locate the fault based on the type of future switch failures that occur.
[0046] As a preferred solution of the method for testing the reliability of a switch of the present invention, the method introduces an information filtering function based on the mean and standard deviation of each switch feature to obtain the deviation degree value of the performance feature from the normal state, which specifically includes the following steps:
[0047] According to the mean and standard deviation of each switch feature, the deviation of the current performance feature from its mean is calculated, and based on the calculated deviation, the exponential function and standard deviation are used for further calculation to obtain the deviation degree of the performance feature from the normal state, which is expressed as follows:
[0048]
[0049] Among them, φ(F i) represents the deviation value of the value of the i-th performance characteristic after being processed by the information filtering function,
[0050] μ i represents the mean value of the i-th performance characteristic under normal conditions, σ i represents the standard deviation of the ith performance characteristic.
[0051] As a preferred solution of the method for testing the reliability of a switch of the present invention, a warning is issued based on the location of the switch fault, and maintenance personnel are notified to take corresponding measures in advance, which specifically includes the following steps:
[0052] Create a test report that reflects the reliability of the switch based on the location of the switch fault;
[0053] The test report includes the switch fault type, fault location, affected performance indicators, performance indicator changes of the switch under different traffic and environmental data, and adjustment suggestions;
[0054] Send the test report to the maintenance personnel as a maintenance reference.
[0055] In a second aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the method for testing switch reliability as described in the first aspect of the present invention is implemented.
[0056] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the method for testing switch reliability as described in the first aspect of the present invention.
[0057] The beneficial effects of the present invention are as follows: by selecting a completely closed test room and installing equipment that independently controls environmental data, various extreme working conditions that may be encountered in actual operation can be accurately simulated, which not only improves the authenticity of the test, but also makes the performance of the switch closer to the actual application scenario; in addition, by establishing a switch fault prediction model, accurate prediction of future switch fault types is achieved, and maintenance personnel can be guided to take targeted measures to prevent the occurrence of potential problems, significantly improving the reliability and maintenance efficiency of the switch. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0059] Figure 1 This is a flow chart of the method for testing switch reliability in Example 1.
[0060] Figure 2 This is a schematic diagram of the switch failure prediction result obtained in Example 1. DETAILED DESCRIPTION
[0061] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0062] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0063] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0064] Example 1, reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, provides a method for testing the reliability of a switch, comprising the following steps:
[0065] S1. Set up a switch test room that can control environmental data and traffic, and determine the preliminary test plan for switch performance based on the switch's operating specifications.
[0066] The specific steps include:
[0067] S1.1. Select a completely enclosed and well-soundproof test room to ensure that external environmental factors do not affect the test results. Install environmental control equipment that can independently control temperature, humidity, vibration, and electromagnetic interference (EMI) in the test room. For example, use precision heaters to accurately adjust the indoor temperature; use humidifiers and dehumidifiers to maintain a specific humidity level; install a vibration table to simulate mechanical vibrations in actual work; deploy shielding boxes or EMI generators to introduce or eliminate electromagnetic interference.
[0068] Further explanation: By selecting a completely closed test room and installing equipment to independently control environmental data, a high degree of controllability and consistency of the switch test environment is achieved. This not only ensures the consistency of each test condition, but also avoids interference from external factors on the test results, thereby improving the accuracy and reliability of the test results.
[0069] S1.2. By connecting LabVIEW with the interface of the environmental control equipment, automatic control of temperature, humidity, vibration and EMI environmental data can be achieved.
[0070] Carefully read the technical manuals and maintenance manuals provided by the switch manufacturer to understand the specific working parameters and technical requirements of the switch, and combine them with the relevant standards issued by the IEC organization to ensure that the test plan complies with industry specifications.
[0071] Develop a detailed preliminary test plan based on the switch's specific operating parameters, technical requirements and industry specifications. The preliminary test plan includes flow testing, environmental testing and cycle mode;
[0072] Traffic testing refers to setting different traffic modes (such as constant traffic, pulse traffic, random traffic) to evaluate the processing capacity and stability of the switch under network load.
[0073] Environmental testing is about defining specific ranges for temperature, humidity, vibration, and EMI, as well as the duration and cycling patterns of each environmental condition.
[0074] The cycle mode refers to the continuous cycle mode of designing high and low temperature alternating cycles (such as once every 4 hours), damp heat cycles (such as once every 8 hours), vibration tests (such as once every 2 hours) and EMI tests (such as once every hour) to simulate various extreme conditions that may be encountered in long-term operation.
[0075] For example, set the temperature range: according to the working specifications of the switch (such as commercial-grade equipment is usually 0℃ to 40℃), determine the specific test temperature range of -10℃ to 60℃. Humidity range: set the relative humidity range to 20% to 90% to cover various conditions from dry to humid. Vibration frequency: set the vibration frequency range to 10Hz to 500Hz to simulate mechanical vibrations of different intensities. Electromagnetic interference (EMI) intensity: set the EMI intensity range to 80dBμV to 120dBμV, covering common electromagnetic interference levels. Inject different types of Layer 2 code streams into the switch, including but not limited to data packets with frame lengths of 64B, 128B, 256B, etc. Set the initial sending rate to 1Gbps and gradually increase it to the maximum design capacity (such as 10Gbps) to evaluate the performance of the switch under different loads. Continue testing for at least 30 minutes at each rate level to ensure that the data fully reflects the performance of the device.
[0076] High and low temperature alternating cycle: In the high temperature stage, the temperature is maintained at 60℃ for 4 hours. In the low temperature stage, the temperature drops to -10℃ for 4 hours. Number of cycles: A total of 10 complete high and low temperature alternating cycles are performed.
[0077] It is further explained that by specifying the specific range of traffic and environmental data, as well as the duration and cycle mode of each environmental condition, the standardization and repeatability of the test process are ensured. This method not only improves the reliability of the test results, but also provides strong support for evaluating the long-term performance of the switch in a complex and dynamic environment. In particular, the design of the cycle mode can more realistically reproduce the actual conditions of the switch in long-term operation, providing an important basis for evaluating its reliability and durability.
[0078] S2. Restore the switch to its initial state and record the performance data of the switch in the normal state as the baseline data.
[0079] The specific steps include:
[0080] Use the switch's built-in web management interface and select "Restore Factory Settings" or a similar option according to the switch manufacturer's instructions. By executing the factory reset command, you ensure that each test is performed under the same and known initial conditions, eliminating the result deviation caused by different configurations.
[0081] Adjust the temperature, humidity, vibration and electromagnetic interference in the test room to the standard range, for example, the temperature is maintained at 25℃±2℃, the humidity is controlled at 40%-60%, and there is no obvious vibration and electromagnetic interference.
[0082] NetFlow is a widely used traffic monitoring technology. SNMP is mainly used to manage and monitor the status and performance of network devices. NetFlow and SNMP are used to collect the CPU utilization, memory occupancy, throughput, packet loss rate and latency of the switch. The switch is allowed to run continuously for a period of time (such as 24 hours) in a normal environment to collect sufficient performance data.
[0083] Use NetFlow and SNMP to collect switch performance indicator data regularly (such as once a minute) to ensure data continuity and integrity. Add accurate timestamps to each set of performance indicators to ensure that the time sequence of the data is clear. Arrange all collected performance indicator data into a table format, with each row representing a data record at a time point and each column corresponding to a performance indicator (such as CPU utilization, memory occupancy, etc.). Export the sorted data into a CSV format file to facilitate subsequent data processing and analysis.
[0084] By sorting the performance indicator data by timestamp and saving it in CSV format, a structured baseline data set is formed. This data set is not only easy to manage and analyze, but also provides a solid foundation for subsequent performance comparison and fault prediction.
[0085] S3. Adjust the traffic and environmental data of the switch according to the preliminary switch performance test plan, record the performance indicator data of the switch, and compare it with the baseline data to obtain the preliminary performance analysis results of the switch.
[0086] The specific steps include:
[0087] S3.1. Perform the test task according to the preliminary test plan. First, set a collection period (for example, 5 minutes), duration (each load level is maintained for at least 15 minutes), and total test time (for example, 15 minutes for each scenario, a total of 45 minutes); when starting the test task, start with low traffic, gradually increase to medium traffic, and finally reach high traffic. For environmental conditions, simulate different environmental conditions (such as high temperature, high humidity, and strong electromagnetic interference) and record their impact on switch performance. Each time you collect switch performance indicators, record the current time to ensure that each set of performance indicators has an accurate timestamp.
[0088] Load the baseline data set under normal environment established before, compare the currently collected performance indicators with the baseline data set item by item, calculate the difference, and for each performance indicator, calculate its change range at different time points.
[0089] It is further explained that by comparing with the baseline dataset, the performance changes of the switch under different environmental conditions can be accurately evaluated.
[0090] S3.2. Based on the historical test data, calculate the mean and standard deviation of each performance indicator, and set the performance threshold range according to the mean and standard deviation. For each performance indicator, if the difference between the currently collected performance indicator and the baseline data set exceeds the set performance threshold, it will be marked as abnormal.
[0091] Plotly was selected as the visualization tool due to its powerful interactivity and rich chart types. Histograms were drawn to compare performance indicators at different time points, and performance indicators marked as abnormal were highlighted in red for quick identification.
[0092] It is further explained that the performance data is displayed in a graphical manner to make the analysis results more intuitive and easy to understand.
[0093] S4. Integrate the performance indicator data and preliminary performance analysis results of the switch into test data and perform preprocessing.
[0094] The specific steps include:
[0095] The performance indicator data of the switch collected by NetFlow and SNMP after the execution of the preliminary test plan are associated and merged with the performance graph according to the timestamps to form the test data.
[0096] Further explanation: By merging the data collected by NetFlow and SNMP with the performance graphs, a complete test data set containing all relevant information is formed. This not only improves the integrity of the data, but also provides rich information support for subsequent analysis.
[0097] Based on the test data, a statistical method is used to calculate the Z score (standardized score) of each performance indicator to identify data points that deviate from the normal range. The LOF algorithm is used to identify data points with low local density, which may be outliers. For data points marked as abnormal, they are directly removed from the data set to prevent them from interfering with subsequent analysis.
[0098] For a small number of missing values, you can use the mean or median of the column to fill in the missing values.
[0099] Each performance indicator is converted into a Z score and standardized so that its mean is 0 and its standard deviation is 1. The standardization eliminates the impact of the dimensions between different performance indicators, making the comparison between the indicators more fair and reasonable.
[0100] S5. Extract performance features according to the preprocessed test data, and establish a switch fault prediction model based on the extracted performance features to obtain the type of fault that will occur in the future of the switch, and further determine the location of the switch fault according to the type of fault that will occur in the future of the switch.
[0101] The specific steps include:
[0102] Based on the preprocessed test data, information gain analysis is used to extract the CPU utilization, memory usage, throughput, packet loss rate, and latency of the test data as switch features for storage; information gain is a measure of the amount of information a feature provides when distinguishing different categories. For the performance prediction task in this example, it is used to measure the contribution of each performance indicator to distinguishing normal from abnormal states.
[0103] Define switch fault types based on switch characteristics (e.g., hardware fault, software fault, configuration problem, external environment impact);
[0104] Hardware failures such as CPU overload, insufficient memory, network card failure, power failure;
[0105] Software failures such as operating system or firmware problems, application errors;
[0106] Configuration issues such as incorrect routing configuration, incorrect VLAN configuration, and improper QoS settings;
[0107] External environmental influences such as excessive temperature and electromagnetic interference (EMI).
[0108] Using a baseline data set representing the normal operation of the switch, the mean and standard deviation of each switch feature are calculated, and the expression is:
[0109]
[0110] Among them, μ i represents the mean of the i-th performance feature in the baseline data set, n represents the total number of performance features, and is also the number of calculations. represents the i-th performance characteristic in the b-th calculation, σ i represents the standard deviation of the i-th performance characteristic;
[0111] Based on the mean and standard deviation of each switch feature, an information filtering function is introduced to calculate the deviation between the current performance feature and its mean. Based on the calculated deviation, an exponential function and standard deviation are used for further calculation to obtain the degree of deviation between the performance feature and the normal state. The expression is:
[0112]
[0113] Among them, φ(F i ) represents the deviation value of the value of the i-th performance feature after being processed by the information filtering function, μ i represents the mean value of the i-th performance characteristic under normal conditions, σ i represents the standard deviation of the ith performance characteristic.
[0114] It is further explained that the degree of deviation of the performance characteristics relative to the baseline data set is quantified in numerical form to facilitate comparison and further analysis.
[0115] Calculate the weighted sum of each fault type j and apply the softmax function and the deviation value φ(F i ), and the probability distribution of future switch failures is obtained, which is expressed as:
[0116]
[0117] Among them, Z j represents the probability of the jth fault type occurring in the future, m represents the total number of fault types, n represents the total number of performance characteristics, and w ij represents the weight factor of the i-th performance characteristic for the j-th fault type, F i represents the value of the i-th performance feature, φ() represents the information filtering function, and k represents the index of the traversed fault type;
[0118] According to the probability distribution Z of future failure of the switch j , the highest probability of failure is taken as the type of failure that will occur in the future of the switch; the probability distribution of failures that will occur in the future is Z j The value range is [0,1], and the sum of all probabilities is 1. Combined with special diagnostic tools and techniques (such as network packet capture analysis, hardware detection, etc.), and logical elimination methods to narrow the scope of the fault and ultimately accurately locate the specific location of the fault.
[0119] For example, the probabilities of three types of failures are [0.344, 0.336, 0.320], among which the probability of the first type of failure is the highest. Its corresponding failure type is insufficient memory. First, use the built-in management interface of the switch to view the current memory usage, use a network packet capture tool (such as Wireshark) to capture traffic, and analyze whether there are abnormally large data packets or continuous data transmission, which may cause memory buffer overflow. Further check whether the physical memory module is working properly to ensure that there is no hardware failure. At this point, it is comprehensively determined that the specific cause of insufficient memory is the physical memory module, and the inspection found that one of the memory sticks has a hardware problem.
[0120] S6. Based on the location of the switch fault, a warning is issued and maintenance personnel are notified to take corresponding measures in advance.
[0121] The specific steps include:
[0122] Based on the determined switch fault location, create a switch test report that includes the switch fault type, fault location, affected performance indicators, changes in switch performance indicators under different traffic and environmental data, and adjustment suggestions.
[0123] For example, the content of the test report is: switch fault type: insufficient memory; fault location: a memory bar in the physical memory module; the affected performance indicators are: CPU utilization, normal state: average 50%, standard deviation 10%, in fault state: up to 85%; memory occupancy, normal state: average 60%, standard deviation 15%, in fault state: continuously over 90%. In low-traffic environment, CPU utilization: slightly fluctuates, remains at around 60%; memory occupancy: close to 90%, but not completely exhausted. In medium-traffic environment, CPU utilization: significantly increases to 75%, memory occupancy: continuously close to 100%. In high-traffic environment, CPU utilization: close to 100%, memory occupancy: reaches 100%, and memory allocation fails. In high-temperature environment (>40℃), CPU utilization: due to heat dissipation problems, CPU frequency is reduced, resulting in increased utilization, memory occupancy: no significant change. The adjustment suggestion is to immediately replace the problematic memory bar to ensure that the hardware works properly, and consider adding additional memory modules to cope with higher load requirements in the future.
[0124] Send the complete test report to the maintainer for them to take corrective actions.
[0125] This embodiment also provides a computer device, which is suitable for the method of testing switch reliability, including: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to implement the method of testing switch reliability proposed in the above embodiment.
[0126] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.
[0127] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the method for testing the reliability of a switch proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Red-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0128] In summary, the present invention can accurately simulate various extreme working conditions that may be encountered in actual operation by: selecting a completely closed test room and installing equipment for independently controlling environmental data, which not only improves the authenticity of the test, but also makes the performance of the switch closer to the actual application scenario; in addition, by establishing a switch fault prediction model, it can achieve accurate prediction of future switch fault types, and can also guide maintenance personnel to take targeted measures to prevent the occurrence of potential problems, significantly improving the reliability and maintenance efficiency of the switch.
[0129] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for testing the reliability of a switch, characterized in that: include, Set up a switch test room that can regulate environmental data and traffic, and determine the preliminary test plan for switch performance based on the switch's operating specifications; Restore the switch to its initial state and record the performance data of the switch in normal state as baseline data; Adjust the switch traffic and environment data according to the switch performance preliminary test plan, record the switch performance index data, and compare it with the baseline data to obtain the preliminary performance analysis results of the switch; Integrate the switch's performance indicator data and preliminary performance analysis results into test data and perform preprocessing; According to the preprocessed test data, the performance characteristics are extracted, and a switch fault prediction model is established based on the extracted performance characteristics to obtain the type of fault that will occur in the future of the switch, and the location of the switch fault is further determined according to the type of fault that will occur in the future of the switch; Based on the location of the switch fault, a warning is issued and maintenance personnel are notified to take corresponding measures in advance.
2. The method for testing the reliability of a switch according to claim 1, wherein: Set up a switch test room that can control environmental data and traffic, and determine the preliminary test plan for switch performance based on the switch's working specifications, including the following steps: Choose a completely closed test room as the location for testing the switch and install equipment that can independently control environmental data; The environmental data includes temperature, humidity, vibration and electromagnetic interference; Set the time series change curves of different environmental data through LabVIEW; Determine the switch operating specifications through switch manufacturers and industry standards, and determine the preliminary test plan for switch performance based on the operating specifications; The preliminary test plan includes specific ranges of flow and environmental data, as well as the duration of each environmental condition and its cycle pattern; The cycle mode refers to a mode in which the switch is continuously cycled through high and low temperature alternating cycles, damp heat cycles, vibration tests, and EMI tests by adjusting environmental data.
3. The method for testing the reliability of a switch according to claim 2, wherein: Restore the switch to its initial state and record the performance data of the switch in normal state as baseline data. The specific steps include: Execute the factory reset command through the switch's management interface; Run the switch under normal environment data and record the switch performance indicators after running under normal environment data through NetFlow and SNMP; The performance indicators of the switch are CPU utilization, memory occupancy, throughput, packet loss rate and delay; The switch performance indicators running under normal environment data are sorted and saved by timestamp to form a baseline data set and saved in CSV format.
4. The method for testing the reliability of a switch according to claim 3, wherein: According to the preliminary switch performance test plan, adjust the switch traffic and environmental data, record the switch performance index data, and compare it with the baseline data to obtain the preliminary performance analysis results of the switch. The specific steps include: Based on the preliminary switch performance test plan, use a device that can independently control the environmental data to adjust the switch's environmental data, and execute the loop mode until the loop mode ends; According to the preliminary switch performance test plan, use a traffic tester to inject traffic into the switch and record the difference between the actual sending rate and receiving rate at each stage; Use NetFlow and SNMP to collect switch environment data and traffic adjusted performance indicators, and compare them with the baseline data set item by item to obtain the difference between the adjusted performance indicators and the baseline data set; Set performance thresholds based on historical test data. When the difference between each performance indicator after adjustment and the baseline data set is greater than the performance threshold, mark this performance indicator as abnormal. Use Plotly to draw a performance graph for each adjusted performance indicator, and distinguish the performance indicators marked as abnormal in red, which is the preliminary performance analysis result of the switch.
5. The method for testing switch reliability as claimed in claim 4, characterized in that: The performance index data and preliminary performance analysis results of the switch are integrated into test data and preprocessed, which specifically includes the following steps: Combine the adjusted performance indicators of switch environment data and traffic collected by NetFlow and SNMP with the performance graph to form test data; Based on the test data, statistical methods are used to identify and remove data points that deviate from the normal range and fill in missing values; Standardize the test data.
6. The method for testing switch reliability as claimed in claim 5, characterized in that: According to the preprocessed test data, the performance characteristics are extracted, and a switch fault prediction model is established based on the extracted performance characteristics to obtain the type of fault that will occur in the future of the switch. According to the type of fault that will occur in the future of the switch, the location of the switch fault is further determined, which specifically includes the following steps: Based on the preprocessed test data, information gain analysis is used to extract the CPU utilization, memory occupancy, throughput, packet loss rate and delay of the test data as switch features; Define the switch fault type based on switch characteristics; For each switch feature, calculate its mean and standard deviation in the baseline dataset; Based on the mean and standard deviation of each switch feature, an information filtering function is introduced to obtain the degree of deviation of the performance feature from the normal state; Calculate the weighted sum of each fault type of the switch, and apply the softmax function and the deviation value of the performance characteristics from the normal state to obtain the probability distribution of future faults of the switch, which is expressed as: Among them, Z j represents the probability of the jth fault type occurring in the future, m represents the total number of fault types, n represents the total number of performance characteristics, and w ij represents the weight factor of the i-th performance characteristic for the j-th fault type, F i represents the value of the i-th performance feature, φ() represents the information filtering function, and k represents the index of the traversed fault type; According to the probability distribution of future faults of the switch, the fault type with the highest probability is taken as the fault type that will occur in the future of the switch; Use diagnostic tools and troubleshooting techniques to locate the fault based on the type of future switch failures that occur.
7. The method for testing the reliability of a switch according to claim 6, characterized in that: The method introduces an information filtering function based on the mean and standard deviation of each switch feature to obtain a value of the degree of deviation between the performance feature and the normal state, specifically including the following steps: According to the mean and standard deviation of each switch feature, the deviation of the current performance feature from its mean is calculated, and based on the calculated deviation, the exponential function and standard deviation are used for further calculation to obtain the deviation degree of the performance feature from the normal state, which is expressed as follows: Among them, φ(F i ) represents the deviation value of the value of the i-th performance feature after being processed by the information filtering function, μ i represents the mean value of the i-th performance characteristic under normal conditions, σ i represents the standard deviation of the ith performance characteristic.
8. The method for testing the reliability of a switch according to claim 7, characterized in that: Based on the location of the switch fault, a warning is issued and maintenance personnel are notified to take corresponding measures in advance, including the following steps: Create a test report that reflects the reliability of the switch based on the location of the switch fault; The test report includes the switch fault type, fault location, affected performance indicators, performance indicator changes of the switch under different traffic and environmental data, and adjustment suggestions; Send the test report to the maintenance personnel as a maintenance reference.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for testing switch reliability according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for testing switch reliability according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Automated test apparatus and test method of switch
CN106789423A
Automatic test method and system for switch
CN115208787A
Network switch detection system
CN117478540A
Switch fault diagnosis method and system
CN118337741A
Fault detection visualization processing system and method for industrial switch
CN119254613A