Server system stability test system and method
Through the multi-dimensional data collection and analysis of the SWC data interaction unit and the SOST system, the shortcomings of existing server restart stability tests have been solved, automation, centralized management and real-time monitoring have been achieved, and the comprehensiveness, accuracy and efficiency of server stability tests have been improved.
Patent Information
- Application Number
- CN202510832617.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-26
AI Technical Summary
Existing server restart stability tests are unable to comprehensively evaluate key data before and after multiple server restarts, lack in-depth operational status monitoring, the inspection process is not detailed enough, and there is a lack of unified management and centralized viewing capabilities, resulting in low test efficiency and insufficient accuracy.
Adopting SWC data interaction unit, SOST log unit, SOST server information collection unit, SOST stability test unit and SOST stability test self-starting unit, through multi-dimensional data collection and analysis, it realizes automation, centralized management and real-time monitoring of the server, including multiple comparative analysis of memory utilization, CPU utilization, disk I/O, log content and thread status, and supports CS distributed architecture design and visual interface.
It has achieved improvements in the comprehensiveness, accuracy and efficiency of server stability testing, reduced manual operation delays and errors, improved the accuracy of test results and management convenience, and reduced labor costs and operational risks.
Smart Images

Figure CN120704964A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of server stability testing, and relates to a server system stability testing system and method. Background Art
[0002] As server application scenarios become increasingly complex and the number of servers continues to increase, the requirements for server stability testing are becoming increasingly stringent. However, current server restart stability testing has significant shortcomings, mainly reflected in the following aspects. These problems make it difficult for existing testing methods to fully evaluate server stability:
[0003] Insufficient comparative information coverage: Server restart stability testing typically focuses only on the results of a single restart, lacking comprehensive comparison and trend analysis of key data before and after multiple restarts. This critical data includes, but is not limited to, memory usage, CPU load, log information, and thread status. Without this comprehensive comparison and trend analysis, it's difficult to effectively capture the patterns and changing trends of issues, potentially missing potential stability risks. For example, some potential issues may only become apparent after multiple restarts, but traditional testing methods cannot detect them in a timely manner.
[0004] Incomplete information acquisition: Existing testing methods primarily focus on basic operating parameters, such as whether the system boots normally and basic performance indicators. However, they lack visibility into deeper operational status, such as hardware health, dependencies between microservices, and network packet loss and latency, making it difficult to accurately locate and resolve issues.
[0005] Inadequate inspection process: During the post-reboot verification phase, testing often focuses solely on whether the system successfully boots up. Deeper checks on the system's health, such as thread deadlocks, resource leaks, and abnormal file handle usage, are often overlooked. This can easily mask potential long-term risks and lead to inaccurate assessments of the server's actual stability. For example, a thread deadlock may not affect system startup in the short term, but could cause a system crash over the long term.
[0006] Lack of unified management and centralized viewing capabilities: Current stability testing requires testers to remotely log in to each server one by one to check its operating status. This is manageable with a small number of servers, but with a large number of servers or a distributed cluster environment, testing and management efficiency is extremely low. Testers are unable to centrally view the status and test results of each server, which not only increases manual workload but also significantly increases operational complexity and the risk of error. For example, in a distributed cluster environment, testers need to spend a lot of time and effort switching between multiple machines, which can easily lead to missing important information or making operational errors.
[0007] In summary, existing restart-type stability tests cannot meet the requirements of modern complex server environments for comprehensiveness, accuracy, efficiency, and centralized management, and are in urgent need of improvement to enhance the effectiveness, automation, and operability of the tests. Summary of the Invention
[0008] The purpose of the present invention is to solve the problem that the restart type stability test in the prior art cannot meet the comprehensiveness, accuracy, efficiency and centralized management requirements of the modern complex server environment, and to provide a server system stability testing system and method.
[0009] In order to achieve the above object, the present invention adopts the following technical solutions:
[0010] A server system stability testing system, comprising:
[0011] SWC data interaction unit, used for communication between client and server, including SOSTClientPort service management, SOSTClientWeb service management, as well as data interfaces for stability information data, remote test stop, remote test start, test log download, and stability history viewing;
[0012] The SOST log unit provides complete log services for SOST operation, recording all executed commands or error messages during SOST operation into files;
[0013] The SOST server information collection unit runs together with the SWC data interaction unit in the swc-manager.server service, collects the server stability test type, CPU, and MEM information in real time and records the results in a json file, and transmits the data to the server through the SWC data interaction unit;
[0014] In the SOST stability test unit, the user selects the stability test type, checks the status of the server before the stability test begins to see if it is capable of conducting the stability test, and then starts the stability test;
[0015] SOST stability test auto-start unit, used to execute the SOST auto-start program after each server restart;
[0016] The SOST log comparison unit evaluates whether the current stability test type passes by comparing the results of the first-round device information with the current-round device information.
[0017] The system performs data comparison and analysis before and after multiple restarts, specifically through the following steps:
[0018] Memory utilization is collected by using the free command in Linux to match the current memory used and free memory capacity information, and the current memory usage is calculated by using used / free*100;
[0019] CPU utilization is collected using a time difference algorithm to obtain CPU utilization information. Specifically, the actual working time ratio is obtained by dividing (the change in the total CPU running time - the difference between the two CPU idle times) / the total time difference * 100.
[0020] Disk I / O collection: Use the iostat command to match the %util information of each hard disk and return the disk IO utilization rate;
[0021] Log content analysis: Using keyword matching principles, we analyze the number of occurrences of the error / fail / unknown keywords in the dmesg and messages log files, as well as the severity of the errors, to assess whether there are any issues with the current system stability.
[0022] Thread status analysis: Analyze the thread status by checking the driver loading status in lsmod after each boot and the delay in command return information after running the specified command;
[0023] Identify potential resource bottlenecks due to system performance fluctuations. Capture CPU / MEM information within a specified timeframe and assess whether excessive CPU core / MEM usage occurs when the server is idle or under stress. If multiple abnormal CPU core / MEM usage occurs within 10 seconds, it indicates a resource bottleneck.
[0024] Early warning mechanism: when the system performance fluctuates abnormally, an alarm message will be sent to the executing user's email address via the SMTP protocol.
[0025] The system achieves comprehensive information collection and in-depth monitoring through the following methods:
[0026] Hardware health status collection: When performing stability testing on a server, run the first run to collect detailed hardware information from the current server using lspci / lsmem / lscpu / dmidecode / ipmitool. After the server is restarted and the system is re-entered, collect the server hardware information using the same method and compare them.
[0027] Network delay and packet loss monitoring: When the server undergoes a network stress test, the detailed ifconfig / ethtool information of all the server's network ports is detected, and whether there is packet loss or error is captured and recorded in real time. Ping tests are performed on the SOST server using the stress network port to collect real-time delay information. The server's stability is evaluated by analyzing the packet loss / error packet / Ping data.
[0028] Disk I / O performance monitoring: Use the iostat command to match the %util information of each hard disk and return the disk I / O utilization. When the hard disk is under stress testing, the hard disk utilization information is collected in real time. When the utilization fluctuates by 10%, a disk I / O performance alarm is triggered.
[0029] Memory and CPU usage monitoring: When performing stability testing on a server, CPU / MEM utilization information is collected within a set time period. When the utilization fluctuates by 15%, an abnormal CPU / MEM utilization alarm is triggered.
[0030] The data association analysis algorithm uses different methods to obtain certain data during the stability test. When the method for obtaining information is not unique, a correlation comparison analysis is performed on multiple aspects of the data to ensure the consistency of the same information obtained from all aspects. When the information obtained by a certain method is different from the information obtained by other methods, the system re-acquisition method is triggered and the abnormal information acquisition method is executed again to re-acquire the data. If the data is still inconsistent, the user is informed of the inconsistency between the interface for obtaining information and the data returned by the interface.
[0031] Identify problems through cross-level data correlation analysis. When a network port loss problem occurs during the server stability test, first check the hardware netinfo information. If a physical device is lost, feedback will be given on the network port loss caused by the loss of the hardware network card. If the physical device is not lost, but the network port is lost, check whether the driver status of the network card in the system is abnormal and whether the link status of the network port is normal. If normal, automatically configure the network port IP and test it with the SOST server through a ping request to check whether the problem is caused by a network layer problem.
[0032] The system adopts CS distributed architecture design to achieve unified management and real-time status monitoring of server clusters. The specific implementation method is as follows:
[0033] After the user installs SOST on the Client, the SOST Server will proactively discover new SOST Client devices on the network. After SOST is installed on the Client, the TCP port will be automatically opened. The SOST Server will scan the survival status of specific ports in the LAN. If it finds that the port device is not in the device list, it will be added to the updated device list.
[0034] When a client device is added to the device list for the first time, the server proactively sends a request to the client to collect detailed server information. After receiving the request, the client executes the relevant information collection program on the server to obtain the current server device information and operating status. The server then sends a test request to the client. After receiving the request, the client marks the flag to start collection and executes the hardinfo collection program to collect information about the server system's CPU, MEM, HDD, CPU usage, MEM usage, IO usage, and net usage. Finally, the SOST analyzes the results and records the data in a json file.
[0035] After data collection is completed, the flag is set to "collect completed". The server requests the client to complete the collection and request the result json data. The data is saved to the server for subsequent data display.
[0036] The server sends data update requests to all active clients at set intervals. When the server performs a restart stability test, the client indicates that the stability test has started. The JSON file records the current stability test type, number of laps, results, start time, and end time. When the server detects that the client is performing the stability test normally, it stores the stability test JSON data file locally.
[0037] The SOST Server supports remote execution of stability tests, stopping stability tests, downloading logs, and executing commands. The Server sends a specific post request to the Client. After receiving the request, the Client determines the commands to be executed and the returned information based on different flags.
[0038] The specific process of executing the stability test is as follows:
[0039] After the user starts the SOST process, they select the stability type to be tested and confirm the test. SOST first checks and configures the reboot user self-start method in the Linux system, and then starts the server detailed information collection module, collecting detailed information about the server's hardware and software and recording it in the result folder.
[0040] After the information collection is completed, execute the command to restart the server. After the restart, enter the system again. After the Linux system automatically logs in, it automatically starts the SOST process for the next stability test.
[0041] Data transmission: SOST Server periodically obtains the stability test status of the current server from all clients. Every time SOST restarts and enters the system, the server collects detailed information about the stability test of the current client.
[0042] Report generation: When the user stops the stability test in the Linux system or through the dashboard, the client executes the result_html() method to analyze the stability test results of the server, determine whether the current server stability test result is passed, and synchronously record the key server information in the stability test report. If the stability test result is a failure, all failure information is synchronously recorded in the stability test report.
[0043] The test types supported by the system include:
[0044] Standard reboot test, reboot, systemctl reboot, shutdown - r;
[0045] Power management operation test, power cycle, power reset, power failure recovery test AC Lost;
[0046] BMC remote power management test, including BMC Remote Power Reset, BMC Remote Power Cycle, BMC remote bmc reset warm, BMC remote bmc reset cold, and BMC remote raw 0x06 0x02.
[0047] Other tests include a user-defined reboot command test, a multi-model test that runs four types of reboot, power cycle, power reset, and aclost tests in sequence, a BMC Sensor / SDR test that collects Sensor. / SDR information from the server's BMC and converts it to CSV format, and a BMC Quickly Create Sel Log and BMC Quickly Create Audit log.
[0048] The system also includes a visual interface for displaying the test process, result monitoring and problem diagnosis information, and provides real-time data display, automated report generation, task management and abnormal warning.
[0049] The core function implementation method in the visual interface includes:
[0050] One-click stop test: When the test needs to be stopped, the background specifies the swc_stop_test.py program to pass in the system IP parameter, and continuously sends stop test requests to the Client IP. When the server stability is restarted and the swc service is running, swc receives the stop test request and modifies the stop_flag bit to 1. When the sost main program continues to perform the stability test, it checks the swc_flag bit. If it is 1, the test is stopped.
[0051] Download test files. When downloading test files, the client code backs up the files generated by the current stability test type to the / tmp directory and packages them into a tar file. The backed-up tar test file is downloaded via HTTP. After the download is complete, the client actively deletes the backup file and the tar result file.
[0052] Add the SOST client IP address. When the new client cannot be automatically identified by the server, customize the added ClientIP. The server will capture new client devices in the network in real time and record them in the device list file. It will check whether the added IP is alive and is a SOST client. If it meets the conditions, it will be added to the device list file. The server will then regularly collect client data.
[0053] The core function implementation method in the visual interface also includes:
[0054] Delete useless lines. When the user deletes the specified client data, the server first deletes the client IP information from the device list file, and then deletes the file information recorded by the client on the server.
[0055] Project filtering / test result filtering / test type filtering is implemented based on the HTML code on the server front end. After classifying all projects / results / test types, a drop-down menu option is provided for users to filter specific data.
[0056] A server system stability testing method comprises the following steps:
[0057] Data interaction, which realizes the communication connection between the client and the server through the SWC data interaction unit. The SWC data interaction unit includes SOSTClientPort service management, SOSTClientWeb service management, and data interfaces for providing stability information data, remotely stopping the test, remotely starting the test, downloading the test log, and viewing the stability history;
[0058] Logging: Use the SOST log unit to provide complete logging services for SOST operation, and record all executed commands or error messages during SOST operation into files;
[0059] Server information collection: The SOST server information collection unit and the SWC data interaction unit run together in the swc-manager.server service to collect server stability test type, CPU, and MEM information in real time, and record the collection results in a json file. The data is then transmitted to the server through the SWC data interaction unit;
[0060] Stability test execution, in the SOST stability test unit, the user selects the stability test type. This unit checks whether the status assessment of the server before the stability test begins is suitable for stability testing. If the assessment passes, the stability test begins;
[0061] Self-start, passing the SOST stability test. The self-start unit executes the SOST self-start program after each server restart;
[0062] Log comparison evaluation, using the SOST log comparison unit, compares the results of the first-round device information with the current-round device information to evaluate whether the current stability test type has passed.
[0063] Compared with the prior art, the present invention has the following beneficial effects:
[0064] The server system stability test system in the present invention provides a variety of data interfaces through the SWC data interaction unit, covering functions such as stability information data, remote control, log download and history viewing. It can obtain various types of information in the server stability test process in an all-round and multi-dimensional manner, ensuring that the test data is comprehensive and accurate, providing a solid foundation for accurately evaluating server stability and avoiding the evaluation deviation caused by incomplete information acquisition in traditional testing methods. The SOST server information collection unit collects key information such as server stability test type, CPU, MEM, etc. in real time and records it in a json file. This enables testers to promptly grasp the status changes of the server during the test process and have a clearer understanding of server performance bottlenecks and potential problems. Compared with traditional manual monitoring methods, it greatly improves the accuracy of server stability evaluation.
[0065] The system is equipped with a SOST stability test self-starting unit, which can automatically execute the SOST self-starting program after each server restart without manual intervention, realizing the automated startup of the test process, reducing delays and errors caused by manual operations, improving test efficiency, and ensuring that the server can quickly enter the stability test state after restarting and detect potential problems in a timely manner.
[0066] After the user selects a test type, the SOST stability test unit automatically checks the server status and assesses whether the conditions for stability testing are met. If the assessment passes, the test automatically begins. This simplifies the test process, reduces the professional knowledge and operational skills required of testers, and reduces test interruptions or errors caused by human judgment errors.
[0067] The SOST log unit provides comprehensive logging services for SOST operations, recording all executed commands and error messages in files. Testers can quickly locate and analyze problems that arise during testing by viewing log files. This significantly reduces problem location time and improves problem-solving efficiency compared to traditional manual troubleshooting.
[0068] The SOST log comparison unit automatically evaluates whether the current stability test type passes by comparing the first-round device information with the current-round device information. This allows for rapid conclusions and intuitive feedback to testers, eliminating the tedious and error-prone manual analysis of large amounts of data and improving the efficiency and accuracy of test result analysis.
[0069] The SWC data exchange unit provides multiple data interfaces, allowing testers to remotely control and manage the test process, such as stopping and starting a test remotely. It also facilitates access to test logs and historical stability data. This makes test management more flexible and efficient, allowing testers to monitor and adjust the test process in real time without having to be on-site, improving the convenience and efficiency of test management.
[0070] All types of data collected and recorded by the system can be transmitted to the server through the SWC data interaction unit, facilitating data sharing and collaborative analysis among team members. Personnel in different positions can obtain corresponding test data according to their own needs and jointly participate in the analysis and resolution of server stability issues, thereby improving team collaboration efficiency and promoting the optimization and improvement of server performance. The automation and intelligent features of this system reduce the need for manual operations and lower labor costs. At the same time, since server stability issues can be discovered in a timely manner, risks such as business interruption and data loss caused by server failures are avoided, the economic losses and reputation damage caused by server failures are reduced, and a large amount of potential costs are saved for the enterprise. Through comprehensive and accurate testing, potential defects in server design and configuration can be discovered in advance, providing a basis for server optimization and upgrades, avoiding the cost increase and business impact caused by problems discovered after large-scale server deployment, and further reducing the operational risks of the enterprise. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0072] Figure 1 is a block diagram of the stability testing system of the present invention;
[0073] Figure 2 Schematic diagram of the overall architecture of the SOST system of the present invention;
[0074] Figure 3 Schematic diagram of the visualization interface of the present invention. DETAILED DESCRIPTION
[0075] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0076] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.
[0077] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0078] The present invention is described in further detail below with reference to the accompanying drawings:
[0079] See also Figure 1 , is a block diagram of a server system stability test system in the present invention, including:
[0080] SWC data interaction unit, used for communication between client and server, including SOSTClientPort service management, SOSTClientWeb service management, as well as data interfaces for stability information data, remote test stop, remote test start, test log download, and stability history viewing;
[0081] The SOST log unit provides complete log services for SOST operation, recording all executed commands or error messages during SOST operation into files;
[0082] The SOST server information collection unit runs together with the SWC data interaction unit in the swc-manager.server service, collects the server stability test type, CPU, and MEM information in real time and records the results in a json file, and transmits the data to the server through the SWC data interaction unit;
[0083] In the SOST stability test unit, the user selects the stability test type, checks the status of the server before the stability test begins to see if it is capable of conducting the stability test, and starts the stability test.
[0084] The specific process of stability testing is as follows:
[0085] After the user starts the SOST process, they select the stability type to be tested and confirm the test. SOST first checks and configures the reboot user self-start method in the Linux system, and then starts the server detailed information collection module, collecting detailed information about the server's hardware and software and recording it in the result folder.
[0086] After the information collection is completed, execute the command to restart the server. After the restart, enter the system again. After the Linux system automatically logs in, it automatically starts the SOST process for the next stability test.
[0087] Data transmission: SOST Server periodically obtains the stability test status of the current server from all clients. Every time SOST restarts and enters the system, the server collects detailed information about the stability test of the current client.
[0088] Report generation: When the user stops the stability test in the Linux system or through the dashboard, the client executes the result_html() method to analyze the stability test results of the server, determine whether the current server stability test result is passed, and synchronously record the key server information in the stability test report. If the stability test result is a failure, all failure information is synchronously recorded in the stability test report.
[0089] SOST stability test auto-start unit, used to execute the SOST auto-start program after each server restart;
[0090] The SOST log comparison unit evaluates whether the current stability test type passes by comparing the results of the first-round device information with the current-round device information.
[0091] A visual interface used to display the test process, result monitoring, and problem diagnosis information, providing real-time data display, automated report generation, task management, and exception warnings.
[0092] The core functions of the visual interface are as follows:
[0093] One-click stop test: When the test needs to be stopped, the background specifies the swc_stop_test.py program to pass in the system IP parameter, and continuously sends stop test requests to the Client IP. When the server stability is restarted and the swc service is running, swc receives the stop test request and modifies the stop_flag bit to 1. When the sost main program continues to perform the stability test, it checks the swc_flag bit. If it is 1, the test is stopped.
[0094] Download test files. When downloading test files, the client code backs up the files generated by the current stability test type to the / tmp directory and packages them into a tar file. The backed-up tar test file is downloaded via HTTP. After the download is complete, the client actively deletes the backup file and the tar result file.
[0095] Add the SOST client IP address. When the new client cannot be automatically identified by the server, customize the added ClientIP. The server will capture new client devices in the network in real time and record them in the device list file. It will check whether the added IP is alive and is a SOST client. If it meets the conditions, it will be added to the device list file. The server will then regularly collect client data.
[0096] Delete useless lines. When the user deletes the specified client data, the server first deletes the client IP information from the device list file, and then deletes the file information recorded by the client on the server.
[0097] Project filtering / test result filtering / test type filtering is implemented based on the HTML code on the server front end. After classifying all projects / results / test types, a drop-down menu option is provided for users to filter specific data.
[0098] The system performs data comparison and analysis before and after multiple restarts, specifically through the following steps:
[0099] Memory utilization is collected by using the free command in Linux to match the current memory used and free memory capacity information, and the current memory usage is calculated by using used / free*100;
[0100] CPU utilization is collected using a time difference algorithm to obtain CPU utilization information. Specifically, the actual working time ratio is obtained by dividing (the change in the total CPU running time - the difference between the two CPU idle times) / the total time difference * 100.
[0101] Disk I / O collection: Use the iostat command to match the %util information of each hard disk and return the disk IO utilization rate;
[0102] Log content analysis: Using keyword matching principles, we analyze the number of occurrences of the error / fail / unknown keywords in the dmesg and messages log files, as well as the severity of the errors, to assess whether there are any issues with the current system stability.
[0103] Thread status analysis: Analyze the thread status by checking the driver loading status in lsmod after each boot and the delay in command return information after running the specified command;
[0104] Identify potential resource bottlenecks due to system performance fluctuations. Capture CPU / MEM information within a specified timeframe and assess whether excessive CPU core / MEM usage occurs when the server is idle or under stress. If multiple abnormal CPU core / MEM usage occurs within 10 seconds, it indicates a resource bottleneck.
[0105] Early warning mechanism: when the system performance fluctuates abnormally, an alarm message will be sent to the executing user's email address via the SMTP protocol.
[0106] Comprehensive information collection and in-depth monitoring are achieved through the following methods:
[0107] Hardware health status collection: When performing stability testing on a server, run the first run to collect detailed hardware information from the current server using lspci / lsmem / lscpu / dmidecode / ipmitool. After the server is restarted and the system is re-entered, collect the server hardware information using the same method and compare them.
[0108] Network delay and packet loss monitoring: When the server undergoes a network stress test, the detailed ifconfig / ethtool information of all the server's network ports is detected, and whether there is packet loss or error is captured and recorded in real time. Ping tests are performed on the SOST server using the stress network port to collect real-time delay information. The server's stability is evaluated by analyzing the packet loss / error packet / Ping data.
[0109] Disk I / O performance monitoring: Use the iostat command to match the %util information of each hard disk and return the disk I / O utilization. When the hard disk is under stress testing, the hard disk utilization information is collected in real time. When the utilization fluctuates by 10%, a disk I / O performance alarm is triggered.
[0110] Memory and CPU usage monitoring: When performing stability testing on a server, CPU / MEM utilization information is collected within a set time period. When the utilization fluctuates by 15%, an abnormal CPU / MEM utilization alarm is triggered.
[0111] The data association analysis algorithm uses different methods to obtain certain data during the stability test. When the method for obtaining information is not unique, a correlation comparison analysis is performed on multiple aspects of the data to ensure the consistency of the same information obtained from all aspects. When the information obtained by a certain method is different from the information obtained by other methods, the system re-acquisition method is triggered and the abnormal information acquisition method is executed again to re-acquire the data. If the data is still inconsistent, the user is informed of the inconsistency between the interface for obtaining information and the data returned by the interface.
[0112] Identify problems through cross-level data correlation analysis. When a network port loss problem occurs during the server stability test, first check the hardware netinfo information. If a physical device is lost, feedback will be given on the network port loss caused by the loss of the hardware network card. If the physical device is not lost, but the network port is lost, check whether the driver status of the network card in the system is abnormal and whether the link status of the network port is normal. If normal, automatically configure the network port IP and test it with the SOST server through a ping request to check whether the problem is caused by a network layer problem.
[0113] The system adopts CS distributed architecture design to achieve unified management of server clusters and real-time status monitoring. The specific implementation methods are as follows:
[0114] After the user installs SOST on the Client, the SOST Server will proactively discover new SOST Client devices on the network. After SOST is installed on the Client, the TCP port will be automatically opened. The SOST Server will scan the survival status of specific ports in the LAN. If it finds that the port device is not in the device list, it will be added to the updated device list.
[0115] When a client device is added to the device list for the first time, the server proactively sends a request to the client to collect detailed server information. After receiving the request, the client executes the relevant information collection program on the server to obtain the current server device information and operating status. The server then sends a test request to the client. After receiving the request, the client marks the flag to start collection and executes the hardinfo collection program to collect information about the server system's CPU, MEM, HDD, CPU usage, MEM usage, IO usage, and net usage. Finally, the SOST analyzes the results and records the data in a json file.
[0116] After data collection is completed, the flag is set to complete. The server requests the client to complete the collection and request the result json data, which is then saved to the server for subsequent data display.
[0117] The server sends data update requests to all active clients at set intervals. When the server performs a restart stability test, the client indicates that the stability test has started. The JSON file records the current stability test type, number of laps, results, start time, and end time. When the server detects that the client is performing the stability test normally, it stores the stability test JSON data file locally.
[0118] The SOST Server supports remote execution of stability tests, stopping stability tests, downloading logs, and executing commands. The Server sends a specific post request to the Client. After receiving the request, the Client determines the commands to be executed and the returned information based on different flags.
[0119] The test types supported by the system include:
[0120] Standard reboot test, reboot, systemctl reboot, shutdown - r;
[0121] Power management operation test, power cycle, power reset, power failure recovery test AC Lost;
[0122] BMC remote power management test, including BMC Remote Power Reset, BMC Remote Power Cycle, BMC remote bmc reset warm, BMC remote bmc reset cold, and BMC remote raw 0x06 0x02.
[0123] Other tests include: user-defined reboot command test, multi-model test (reboot, power cycle, power reset, and aclost) that runs in sequence, BMC Sensor / SDR (collects Sensor. / SDR information from the server's BMC and converts it to CSV format), BMC Quickly Create Sel Log, and BMC Quickly Create Audit log.
[0124] An embodiment of the present invention is a method for testing the stability of a server system, comprising the following steps:
[0125] Data interaction, which realizes the communication connection between the client and the server through the SWC data interaction unit. The SWC data interaction unit includes SOSTClientPort service management, SOSTClientWeb service management, and data interfaces for providing stability information data, remotely stopping the test, remotely starting the test, downloading the test log, and viewing the stability history;
[0126] Logging: Use the SOST log unit to provide complete logging services for SOST operation, and record all executed commands or error messages during SOST operation into files;
[0127] Server information collection: The SOST server information collection unit and the SWC data interaction unit run together in the swc-manager.server service to collect server stability test type, CPU, and MEM information in real time, and record the collection results in a json file. The data is then transmitted to the server through the SWC data interaction unit;
[0128] Stability test execution, in the SOST stability test unit, the user selects the stability test type. This unit checks whether the status assessment of the server before the stability test begins is suitable for stability testing. If the assessment passes, the stability test begins;
[0129] Self-start, passing the SOST stability test. The self-start unit executes the SOST self-start program after each server restart;
[0130] Log comparison evaluation, using the SOST log comparison unit, compares the results of the first-round device information with the current-round device information to evaluate whether the current stability test type has passed.
[0131] The overall architecture of SOST is divided into: SWC data interaction system, SOST log system, SOST server information collection system, SOST stability test system, SOST stability test self-starting system, and SOST log comparison system.
[0132] SWC data interaction system: SWC data interaction system is used for the communication connection system between client and server. The system includes SOSTClientPort service management, SOSTClientWeb service management, stability information data / remote stop test / remote start test / download test log / view stability history and other interfaces.
[0133] SOST log system: The SOST log system provides a complete set of log services for SOST operation. All executed commands or error messages during SOST operation are recorded in files through the log system.
[0134] SOST server information collection system: The SOST server information collection module and the SWC data interaction system run together in the swc-manager.server service. The SOST server information collection system mainly collects stability test type / CPU / MEM and other information in the system in real time and records the results in a json file, and transmits the data to the server through the SWC data interaction system.
[0135] SOST stability test system: The SOST main program allows users to select the stability test type, check the status of the server before starting the stability test to see if it is capable of conducting the stability test, and start the stability test.
[0136] SOST stability test self-starting system: SOST stability test self-starting system is the SOST self-starting program executed every time the server is restarted.
[0137] SOST log comparison system: The SOST log comparison system evaluates whether the current stability test type passes by comparing the results of the first-round device information with the current-round device information.
[0138] This precise server stability testing method successfully overcomes the shortcomings of traditional testing methods. Compared to traditional decentralized manual operations, SOST achieves centralized management, automated data collection, and efficient real-time monitoring. Through this system, users can not only easily obtain comprehensive performance metrics but also use intelligent algorithms to analyze and predict potential issues in real time. Most importantly, SOST is equipped with a comprehensive stability dashboard, allowing testing processes, result monitoring, and problem diagnosis to be conducted on a unified platform.
[0139] The process of data flow in the system: The SOST Server will periodically send data acquisition requests to all clients. When the client responds to the request, it will feed back the JSON data to the server via HTTP and record it in a file. The dashboard displays information such as the current server stability test status by reading JSON data files from different devices.
[0140] The dashboard not only displays real-time data but also supports automated report generation, task management, and anomaly alerts, helping users quickly identify key issues within massive amounts of data and respond promptly. Overall, the SOST system significantly improves testing efficiency, enhances risk prediction capabilities, and effectively reduces manual intervention, ensuring efficient server operation and long-term stability.
[0141] 1. Comparative analysis of data before and after multiple restarts
[0142] The SOST system comprehensively collects and compares key performance indicators of the server before and after multiple restarts, such as memory usage, CPU usage, disk I / O, log content, thread status, etc.
[0143] Specific methods for collecting data:
[0144] Memory utilization: Use the Linux free command to match the current memory used and free memory capacity information, and calculate the current memory utilization by using used / free * 100.
[0145] CPU utilization: CPU utilization is obtained using a time difference-based algorithm. The calculation method is: (change in total CPU run time - difference between two CPU idle times) / total time difference * 100 to obtain the actual working time percentage, which is the CPU utilization.
[0146] Disk I / O: Use the iostat command to match the %util information of each hard disk and return the disk IO utilization.
[0147] Log content: SOST uses keyword matching principles to analyze the number of occurrences of keywords such as error / fail / unknown in the dmesg and messages log files, as well as the severity of the errors, to assess whether there are any issues with the current system stability.
[0148] Thread status: SOST analyzes thread status by checking the driver loading status in lsmod after each boot and the delay in command return information after running the specified command.
[0149] It can deeply explore the changing patterns of various indicators during system operation. Through real-time monitoring and analysis of these data, SOST can quickly identify fluctuations in system performance, potential resource bottlenecks and other abnormal situations, predict possible problem trends in advance, and promptly notify relevant personnel through an intelligent early warning mechanism.
[0150] Specific algorithm for potential resource bottleneck anomalies in system performance fluctuations: In the case of system performance fluctuations, by capturing CPU / MEM information within a specified time, evaluate whether the CPU core / MEM memory usage is too high when the server is in a static or stressed state. If the CPU core / MEM memory usage appears multiple times within 10 seconds, then the SOST assessment system currently has a resource bottleneck anomaly.
[0151] Early Warning Mechanism: SOST features a comprehensive email delivery system. When system performance fluctuates abnormally, an SMTP protocol sends an alert to the executing user's email address, notifying relevant personnel. Powerful data analysis and early warning capabilities significantly improve testing accuracy, not only accurately identifying potential system risks but also helping users take effective countermeasures before issues occur, thereby avoiding system crashes or significant performance degradation and safeguarding the long-term stability of the server.
[0152] 2. Comprehensive information collection and in-depth monitoring
[0153] The SOST system provides comprehensive in-depth monitoring from the hardware layer to the application layer, covering all important aspects of the server, including core indicators such as hardware health status, network latency and packet loss, disk I / O performance, memory and CPU usage, and dependencies between microservices.
[0154] Specific definitions and collection methods of core indicators:
[0155] Hardware health status: During the server stability test, detailed hardware information such as lspci / lsmem / lscpu / dmidecode / ipmitool will be collected during the first run. When the server sends a restart command and waits to re-enter the system, the server hardware information will be collected in the same way. After the collection is complete, the current information will be compared with the first run information to check the server hardware monitoring status.
[0156] Network latency and packet loss: When a server undergoes a network stress test, SOST automatically monitors detailed ifconfig / ethtool information for all network ports, capturing and recording any packet loss or errors in real time. For latency checks, SOST uses the stress network port to perform a ping test on the SOST server to collect real-time latency information. SOST evaluates server stability by analyzing packet loss, error packets, and ping data.
[0157] Disk I / O performance: SOST uses the iostat command to match the %util information for each hard drive and returns disk I / O utilization. During a hard drive stress test, SOST collects disk utilization information in real time and triggers a disk I / O performance alarm if utilization fluctuates by 10%.
[0158] Memory and CPU usage:
[0159] Memory utilization: Use the Linux free command to match the current memory used and free memory capacity information, and calculate the current memory utilization by using used / free * 100.
[0160] CPU utilization: CPU utilization is obtained using a time difference-based algorithm. The calculation method is: (change in total CPU run time - difference between two CPU idle times) / total time difference * 100 to obtain the actual working time percentage, which is the CPU utilization.
[0161] When SOST performs stability testing on a server, it collects CPU / MEM utilization information over a period of time. If the utilization fluctuates by 15%, SOST triggers an abnormal CPU / MEM usage alarm.
[0162] Through the meticulous collection and real-time monitoring of these multi-dimensional data, the SOST system can comprehensively and accurately evaluate the performance of the server and conduct in-depth analysis of potential bottlenecks and anomalies at all levels.
[0163] When it comes to information collection, SOST prioritizes not only comprehensive data but also multi-dimensional data. Through cross-layer monitoring, the system correlates various performance indicators and reveals the complex interactions between hardware, networks, and applications.
[0164] Data association analysis algorithm and cross-layer data analysis and identification method:
[0165] Data association analysis algorithm: During the stability test, SOST may use different methods to obtain a certain data. For example, to obtain memory information, the first method uses dmidecode -t memory to obtain detailed memory information. The second method captures detailed memory information on the BMC page. This creates an association between the data. When the method for obtaining information is not unique, SOST will perform association comparison and analysis on multiple aspects of data to ensure the consistency of the same information obtained from all aspects. When the information obtained by one method is different from the information obtained by other methods, the SOST re-acquisition method will be triggered, and the abnormal information acquisition method will be executed again to re-acquire the data. If the data is still inconsistent, the user will be informed of the inconsistency between the information acquisition interface and the data returned by the interface.
[0166] Cross-layer data correlation analysis identifies issues: Network port loss can occur during server stability testing. To address this issue, SOST proposes a cross-layer data correlation analysis method. When a logical loss of an application-layer network port occurs, SOST checks the hardware's netinfo information to see if the physical device is missing. If a physical device is missing, SOST will indicate that the network port loss is caused by a hardware network card. If the physical device is not missing, but the network port is lost, SOST will check the system driver for abnormalities and the link status of the network port. If there are no issues, SOST will automatically configure the network port IP address and test it with the SOST server via a ping request to determine if the issue is caused by a network layer issue.
[0167] Through hardware device inspections, system driver checks, and configured network port IP ping tests, SOST analyzes and locates network port loss issues. For example, hardware-layer failures can cause network latency, impacting microservice response speeds, while application-layer anomalies can exacerbate hardware resource consumption. Through this comprehensive, cross-layer data collection, SOST can quickly identify and accurately locate the root cause of issues during testing, providing solid data support for server optimization. This in-depth monitoring and precise problem location not only improve testing efficiency and accuracy, but also provide testers with critical decision-making insights, helping them quickly identify the source of server performance bottlenecks and implement effective optimization measures to ensure long-term system stability and efficient operation.
[0168] 3. Unified management and centralized viewing capabilities
[0169] The SOST system adopts an advanced distributed architecture. The SOST distributed architecture design and specific implementation methods are as follows: SOST adopts a CS distributed architecture design, where C stands for Client and S stands for Server. The CS mode stands for client / server mode. The SOST Server is a public server, and the SOST Client represents the client that has installed the SOST tool. After the user installs SOST on the Client, the SOST Server will proactively discover new SOST Client devices on the network. Specific acquisition method: After SOST is installed on the Client, a TCP port will be automatically opened. The SOST Server will scan the survival status of specific ports in the LAN. When it is found that the device on the port is not in the device list, it will automatically add and update the device list. When the Client device is added to the device list for the first time, the Server will proactively send a request to the Client to collect detailed information about the server. After receiving the request, the Client will execute the relevant information collection program on the server to obtain the current device information and operating status of the server. Specific method: The server sends a test request to the client. After receiving the request, the client marks the flag to start data collection and executes the hardinfo collection program to collect information about the CPU, MEM, HDD, CPU usage, MEM usage, IO usage, and net usage in the server system. Finally, SOST analyzes the results and records the data in a JSON file.
[0170] After data collection is complete, the collection completion flag is set. The server then requests the client to retrieve the result JSON data and saves it to the server for subsequent data display. The server sends a data update request to all active clients every three seconds. When the server performs a restart-type stability test, the client indicates that the stability test has begun. The JSON file records the current stability test type, number of laps, results, start time, and end time. When the server detects that the client is performing the stability test successfully, it stores the stability test JSON data file locally and displays the data through the dashboard front-end code, allowing users to view the stability test status.
[0171] The SOST Server supports remote execution of stability test / stop stability test / download log / execute command. The Server sends a specific post request to the Client. After receiving the request, the Client determines the command to be executed and the returned information based on different flags.
[0172] It supports unified management and real-time status monitoring of large-scale server clusters. This architecture eliminates the need for testers to manually log in to each server one by one. Instead, they can monitor all servers comprehensively through a centralized stability dashboard. Whether viewing the real-time status of a single server, test results, or analyzing historical data, testers can efficiently operate from the same platform, significantly saving time and effort.
[0173] The stability dashboard integrates multiple powerful features, including visual charts, status comparisons, anomaly alerts, and problem trend analysis. Through visual charts, testers can intuitively understand the operating status, performance indicators, and test progress of each server. The status comparison function allows users to conduct intuitive comparative analysis at different time points or between different servers, allowing for quick identification of performance fluctuations or potential issues. In addition, the system integrates anomaly alerts and problem trend analysis functions, which automatically trigger warnings based on set thresholds, reminding testers to take timely measures to prevent problems from expanding or affecting the stability of the entire cluster. The introduction of a unified management platform not only significantly improves testing efficiency and reduces manual intervention, but also makes the entire testing process more efficient, accurate, and easy to operate. Through real-time monitoring and intelligent analysis, SOST helps testers fully control the status of the server cluster, ensuring that the system can maintain stable and efficient operation over the long term.
[0174] 4. Complete stability dashboard
[0175] The SOST system utilizes a CS (Client-Server) architecture, providing users with an efficient and stable restart-related stability testing and management experience through the collaborative work of the client and server. This architecture, combined with a distributed design concept, addresses the performance, scalability, and management efficiency shortcomings of existing testing systems. The SOST system's stability dashboard provides a powerful, intuitive, and efficient set of functional modules, allowing users to easily manage and monitor all test tasks, server status, and related data. The following are the core functions of the stability dashboard:
[0176] 1. Supports one-click stop test. Users can conveniently stop the ongoing test task with one click through the dashboard interface. During the execution of the test task, if you need to pause or terminate the test, the user only needs to click the "Stop Test" button to immediately interrupt the task, preventing the test from continuing to consume resources or affecting system stability. This function is particularly suitable for situations where abnormalities occur during the test or immediate intervention is required, ensuring the flexibility and controllability of the test process. Function implementation: When the user clicks one-click stop test on the dashboard, the background will specify the swc_stop_test.py program to pass in the system IP parameters, and will continue to send stop test requests to the Client IP. When the server stability restarts and enters the system and the swc service is running, swc receives the stop test request and modifies the stop_flag bit to 1. When the sost main program continues to perform stability testing, it will check swc_flag. If it is 1, then stop the test.
[0177] 2. Supports downloading test files. The dashboard allows users to download all files related to the test task, such as test logs, data reports, and test configuration files. Users can obtain detailed data from the test process directly from the dashboard, which facilitates subsequent analysis, auditing, or report generation. This function ensures that users can export historical test data or logs and process them offline, making it easier to archive or share information with other teams. Function implementation: When the user clicks to download the test file on the dashboard, the client code will back up the files generated by the current stability test type to the / tmp directory and package them into a tar file. The backed-up tar test file is downloaded via HTTP. When the download is complete, the client actively deletes the backup file and the tar result file.
[0178] 3. Supports manual addition of SOST client IP addresses. Users can manually add the IP addresses of SOST clients to facilitate connection and data synchronization between the test system and the client. This function is suitable for some special test scenarios, especially in a dynamically changing environment, where users may need to update the client list regularly. After adding, the client will be integrated into the system to support further test tasks and monitoring management. Function implementation: When a new Client cannot be automatically identified by the Server, the user can customize the addition of ClientIP for monitoring. Implementation method: The Server will capture new Client devices in the network in real time and record them in the device list file. When the user actively adds SOSTClientIP, the SOST Server will actively check whether the added IP is alive and whether it is a SOST Client. If these two conditions are met, it will be added to the device list file. Subsequently, the Server will regularly collect Client data.
[0179] 4. Support for manual deletion of unused rows. The dashboard provides the ability to manually delete unused rows (such as invalid test data or outdated test records) to help users maintain a clean and efficient interface. Users can clean up irrelevant or unused test tasks, records, or configuration rows as needed, thereby improving the dashboard's usability and data accuracy. This ensures that the dashboard remains maintainable even after long-term operation and prevents disorganized test records from impacting analysis. Functional Implementation: Because client data is recorded in the server as files, when a user deletes the specified client data, the server first deletes the client IP information from the device list file and then deletes the client's file information recorded on the server.
[0180] 5. Supports project filtering / test result filtering / test type filtering. The dashboard provides a variety of filtering functions. Users can filter and view specific data based on projects, test results or test types. Project filtering: allows users to view relevant test data based on different projects, which facilitates cross-project comparison and analysis. Test result filtering: users can quickly filter out passed or failed test tasks, helping to focus on abnormal tests and avoid redundant data interference. Test type filtering: users can filter according to different types of tests, such as load testing, performance testing, stress testing, etc., to accurately view different types of test tasks and results. The multi-dimensional filtering function greatly improves the user's operating efficiency and data analysis capabilities. Function implementation: Project / result / test type filtering is all implemented based on the HTML code of the Server front end. After classifying all projects / results / test types, a drop-down menu option is provided for users to filter specific data.
[0181] 6. Supports usage status detection of more test tools. The dashboard not only supports the built-in test tools of SOST, but can also be expanded to support usage status detection of other third-party test tools. Users can monitor the status of each test tool in real time through the dashboard to understand whether it is operating normally, whether there are errors or excessive resource usage. This function provides great convenience for environments where multiple tools are used together, ensuring the efficient operation of all test tools and timely discovery and resolution of potential problems. Function implementation: SOST not only detects its own stability test status, but also detects the execution status of server stress testing tools including iperf / fio / netperf / amdsst / ptu / stress / stream. After the client starts the swc service, it will monitor in real time whether these tools are executed in the current system. If the execution status of a tool appears in the process, the current tool execution status will be returned to the server. The server will record and display the server's current stress test.
[0182] 5. Automation and efficient execution. The SOST system integrates powerful automated test execution capabilities, including the specific process of automated test execution, task scheduling, test execution, data feedback, and report generation. Specific implementation method: Automated test execution specific process: SOST is a server stability testing tool that currently supports 18 types of stability tests:
[0183] 1. Reboot test: The reboot test simulates a user issuing a reboot command in the system to restart the server.
[0184] 2. Power cycle test: The power cycle test simulates the power on and off of the server.
[0185] 3. Power reset: The power reset test simulates the server sending a forced restart test under the system.
[0186] 4. ACLost, ACLost simulates a sudden power outage on the server during operation. After a period of time, the power is restored and the server automatically starts up and enters the system test.
[0187] 5. The systemctl reboot test is the same as the reboot test and satisfies the test that some Linux systems require systemctlreboot to reboot.
[0188] 6. init 6, the init6 command in the Linux system will execute the shutdown script to restart the system normally.
[0189] 7. The poweroff -r test simulates the server being in the system and issues the poweroff -r command to restart the system.
[0190] 8. shutdown -r, shutdown -r test simulates the server in the system and issues shutdown -r to restart the test.
[0191] 9. BMC Sensor / SDR: By using the IPMI KCS channel under the system, the Sensor. / SDR information in the server BMC is continuously collected. Finally, SOST will process the collected data and convert it into CSV format for user viewing.
[0192] 10. BMC Quickly Create Sel Log: Test the BMC to quickly generate a sel log. By continuously issuing sel commands to the BMC, the BMC will record the generated sel logs.
[0193] 11. BMC Quickly Create Audit log: This test allows for quick generation of audit logs by sending incorrect request data to the BMC login interface, causing the BMC to record a login fail condition.
[0194] 12. Other test: SOST provides a test for sending a user-defined reboot command. After the user enters the command to reboot the system, SOST will perform a stability test based on the command provided by the user.
[0195] 13. Mulit Test, SOST performs multi-mode testing. When the user selects MulitMode, the user will be asked to enter the number of stability test cycles to be performed. Currently, four types of tests are supported: reboot / power cycle / power reset / aclost, which can be run in sequence.
[0196] 14. BMC remote power reset: Use the IPMI LAN channel under the system to download the power reset test to the server BMC.
[0197] 15. BMC remote power cycle, by using the IPMI LAN channel under the system, download the power cycle test of the server BMC.
[0198] 16. BMC remote type bmc reset warm, by using the IPMI LAN channel under the system, the server BMC under the overheating restart test.
[0199] 17. BMC remote type bmc reset cold, by using the IPMI LAN channel under the system, cold restart test of the server BMC.
[0200] 18. BMC remote class raw 0x06 0x02, by using the IPMI LAN channel under the system, issuing a raw command to the server BMC to restart the BMC.
[0201] The automated testing process: After the user starts the sost process, selects the stability test type, and confirms the test, SOST will first check and configure the reboot method for the Linux system. It then begins collecting detailed server information, logging detailed hardware and software information into a results folder. Once this information is collected, the server reboot command will be executed to reboot the server. After the reboot, the user will re-enter the system. After the Linux system automatically logs in, it will automatically start the sost process for the next stability test.
[0202] Data feedback: SOST Server will periodically obtain the stability test status of the current server from all clients. Every time SOST restarts and enters the system, the server can collect detailed information about the current client stability test.
[0203] Report Generation: When a user stops a stability test in Linux or through the dashboard, the client executes the result_html() method to analyze the server's stability test results, determine whether the current server stability test result is Pass, and synchronize key server information into the stability test report. If the stability test result is FAIL, all FAIL information is synchronized into the stability test report. Predefined scripts and test scenarios allow users to quickly configure and execute complex restart test environments, significantly reducing the workload of manual configuration and operation. Through automation, SOST can efficiently execute large-scale server restart tests, avoiding errors and inconsistencies caused by manual intervention, and significantly improving test coverage and execution speed.
[0204] During execution, test data is automatically transmitted back to the stability dashboard, updating the status and results of all test tasks in real time. This automated data transmission mechanism not only improves data collection efficiency but also ensures unified management and real-time monitoring of all test information. Through this closed-loop optimization process, the SOST system continuously tracks the execution status and results of each test, automatically generates detailed reports, and uses intelligent analysis to identify potential system issues. The combination of automation and closed-loop processes makes the entire testing process more efficient, accurate, and seamless, significantly reducing the need for manual intervention and promptly capturing any anomalies, providing real-time, reliable data support for system optimization and risk management. This innovative capability of SOST not only improves testing efficiency but also ensures the long-term stable operation of the server.
[0205] 6. Support multiple test types
[0206] The SOST system offers a wide variety of test types to meet the needs of server stability assessment in various scenarios. These tests range from basic system reboots to complex remote power management operations, helping users comprehensively verify server stability and reliability.
[0207] First, SOST supports standard reboot tests, such as reboot, systemctl reboot, and shutdown -r. These tests simulate various reboot operations to verify the system's ability to recover and operate stably during the reboot process. Furthermore, the system provides a variety of power management operation tests, including power cycle, power reset, and power outage recovery tests (such as AC Lost). These tests assess the server's ability to restart after a power outage, ensuring rapid recovery from abnormal power outages.
[0208] At a higher level, SOST also supports remote power management testing using the BMC (Baseboard Management Controller). Users can use remote commands to perform operations such as power resets, power cycles, and hardware resets to evaluate the stability and recovery capabilities of the system under remote management. For example, by executing BMC Remote Power Reset and BMC Remote Power Cycle, you can test the power recovery and restart of the server under remote control.
[0209] SOST also offers the ability to perform long-term system operation tests, timed operations, and other types of stability tests. These tests simulate server performance under extended loads or specific operating conditions, helping users proactively identify potential system risks and performance bottlenecks.
[0210] By supporting such a wide range of tests, the SOST system provides users with comprehensive server stability assessments, ensuring efficient and reliable system operation in a variety of operating environments and under extreme conditions. These tests are not only useful for routine system maintenance but also effectively support troubleshooting and performance optimization, helping enterprises achieve long-term server stability.
[0211] 7. Support refined test configuration
[0212] The SOST system supports comprehensive and refined test configuration, allowing users to flexibly adjust test parameters to meet the stability testing needs of various complex scenarios. This rich configuration option allows users to have detailed control over key aspects of the testing process, achieving more accurate and efficient testing.
[0213] In the test configuration (Test_Config), SOST allows users to customize the test result storage path (Result_path), the number of dedicated and shared LANs (Dedicated_lan_num and Share_lan_num), the maximum restart time (Maximum_restart_time), and the wait time for each phase (such as start_wait_time, aclost_wait_time, and end_wait_time). These options enable users to adjust the test process and resource allocation based on actual test requirements.
[0214] The log filtering (dmesg_filterate) module provides in-depth filtering of kernel logs, enabling detection and analysis of critical errors such as I / O errors, USB device anomalies, hardware failures, PCIe bus errors, and call tracing. By configuring these filtering conditions, the system automatically captures exception information and triggers alarms or terminates the test based on severity, improving problem location efficiency.
[0215] The Multimodal Stability test supports counting and switching between multiple test modes, including AC Loss, Reboot, Power Cycle, and Power Reset. Users can flexibly combine multimodal test scenarios by configuring the number of executions for each test.
[0216] System Information Control (SysInfo_control) provides real-time monitoring and verification of server hardware status, such as PSU type detection (psu_get_type), IPMI sensor check (ipmi_sensor_check), and OS IP type acquisition (osip_get_type), helping users fully understand the hardware and software operating environment.
[0217] The BMC Survival Configuration (BMC_Survival_Config) module allows users to configure key parameters for remote management, such as the BMC IP address, user name, password, and target IP address of the system, and supports targeted testing of the BMC's survivability and responsiveness.
[0218] The SOST system also supports SMTP email notifications and temporary test configuration modules. Through SMTP settings, the system automatically sends email notifications when a test fails. Users can configure information such as the server address, port, sender and recipient email addresses for instant alerts. Temporary test configuration (Test_tmp) allows users to quickly specify the type, number of tests, directory, and run command for a single test, making it suitable for quick verification and short-term testing.
[0219] Finally, the SOST system's version management and update module (sost) provides a variety of control options, including enabling the web console (Web_Console), debugging flags (debug_flags), and log cycle counts (clear_log_circles). This refined configuration mechanism ensures that the SOST system can operate efficiently under different testing requirements and provides users with an optimal testing experience.
[0220] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A server system stability testing system, characterized in that: include: SWC data interaction unit, used for communication between client and server, including SOSTClientPort service management, SOSTClientWeb service management, as well as data interfaces for stability information data, remote test stop, remote test start, test log download, and stability history viewing; The SOST log unit provides complete log services for SOST operation, recording all executed commands or error messages during SOST operation into files; The SOST server information collection unit runs together with the SWC data interaction unit in the swc-manager.server service, collects the server stability test type, CPU, and MEM information in real time and records the results in a json file, and transmits the data to the server through the SWC data interaction unit; In the SOST stability test unit, the user selects the stability test type, checks the status of the server before the stability test begins to see if it is capable of conducting the stability test, and then starts the stability test; SOST stability test auto-start unit, used to execute the SOST auto-start program after each server restart; The SOST log comparison unit evaluates whether the current stability test type passes by comparing the results of the first-round device information with the current-round device information.
2. A server system stability testing system according to claim 1, characterized in that: The system performs data comparison and analysis before and after multiple restarts, specifically through the following steps: Memory utilization is collected by using the free command in Linux to match the current memory used and free memory capacity information, and the current memory usage is calculated by using used / free*100; CPU utilization is collected using a time difference algorithm to obtain CPU utilization information. Specifically, the actual working time ratio is obtained by dividing (the change in the total CPU running time - the difference between the two CPU idle times) / the total time difference * 100. Disk I / O collection: Use the iostat command to match the %util information of each hard disk and return the disk IO utilization rate; Log content analysis: Using keyword matching principles, we analyze the number of occurrences of the error / fail / unknown keywords in the dmesg and messages log files, as well as the severity of the errors, to assess whether there are any issues with the current system stability. Thread status analysis: Analyze the thread status by checking the driver loading status in lsmod after each boot and the delay in command return information after running the specified command; Identify potential resource bottlenecks due to system performance fluctuations. Capture CPU / MEM information within a specified timeframe and assess whether excessive CPU core / MEM usage occurs when the server is idle or under stress. If multiple abnormal CPU core / MEM usage occurs within 10 seconds, it indicates a resource bottleneck. Early warning mechanism: when the system performance fluctuates abnormally, an alarm message will be sent to the executing user's email address via the SMTP protocol.
3. A server system stability testing system as claimed in claim 1, characterized in that: The system achieves comprehensive information collection and in-depth monitoring through the following methods: Hardware health status collection: When performing stability testing on a server, run the first run to collect detailed hardware information from the current server using lspci / lsmem / lscpu / dmidecode / ipmitool. After the server is restarted and the system is re-entered, collect the server hardware information using the same method and compare them. Network delay and packet loss monitoring: When the server undergoes a network stress test, the detailed ifconfig / ethtool information of all the server's network ports is detected, and whether there is packet loss or error is captured and recorded in real time. Ping tests are performed on the SOST server using the stress network port to collect real-time delay information. The server's stability is evaluated by analyzing the packet loss / error packet / Ping data. Disk I / O performance monitoring: Use the iostat command to match the %util information of each hard disk and return the disk I / O utilization. When the hard disk is under stress testing, the hard disk utilization information is collected in real time. When the utilization fluctuates by 10%, a disk I / O performance alarm is triggered. Memory and CPU usage monitoring: When performing stability testing on a server, CPU / MEM utilization information is collected within a set time period. When the utilization fluctuates by 15%, an abnormal CPU / MEM utilization alarm is triggered. The data association analysis algorithm uses different methods to obtain certain data during the stability test. When the method for obtaining information is not unique, a correlation comparison analysis is performed on multiple aspects of the data to ensure the consistency of the same information obtained from all aspects. When the information obtained by a certain method is different from the information obtained by other methods, the system re-acquisition method is triggered and the abnormal information acquisition method is executed again to re-acquire the data. If the data is still inconsistent, the user is informed of the inconsistency between the interface for obtaining information and the data returned by the interface. Cross-level data correlation analysis identifies problems. When a network port loss occurs during server stability testing, the hardware netinfo information is checked first. If a physical device is lost, feedback is provided regarding the network port loss caused by the hardware network card loss. If the physical device is not lost but the network port is lost, check whether the driver status of the network card in the system is abnormal and whether the link status of the network port is normal. If normal, automatically configure the network port IP and test it with the SOST server through a ping request to check whether the problem is caused by a network layer problem.
4. A server system stability testing system as claimed in claim 1, characterized in that: The system adopts CS distributed architecture design to achieve unified management and real-time status monitoring of server clusters. The specific implementation method is as follows: After the user installs SOST on the Client, the SOST Server will proactively discover new SOST Client devices on the network. After SOST is installed on the Client, the TCP port will be automatically opened. The SOST Server will scan the survival status of specific ports in the LAN. If it finds that the port device is not in the device list, it will be added to the updated device list. When a client device is added to the device list for the first time, the server proactively sends a request to the client to collect detailed server information. After receiving the request, the client executes the relevant information collection program on the server to obtain the current server device information and operating status. The server then sends a test request to the client. After receiving the request, the client marks the flag to start collection and executes the hardinfo collection program to collect information about the server system's CPU, MEM, HDD, CPU usage, MEM usage, iousage, and net usgae. Finally, the SOST analyzes the results and records the data in a json file. After data collection is completed, the flag is set to "collect completed". The server requests the client to complete the collection and request the result json data. The data is saved to the server for subsequent data display. The server sends data update requests to all active clients at set intervals. When the server performs a restart stability test, the client indicates that the stability test has started. The JSON file records the current stability test type, number of laps, results, start time, and end time. When the server detects that the client is performing the stability test normally, it stores the stability test JSON data file locally. The SOST Server supports remote execution of stability tests, stopping stability tests, downloading logs, and executing commands. The Server sends a specific post request to the Client. After receiving the request, the Client determines the commands to be executed and the returned information based on different flags.
5. A server system stability testing system as claimed in claim 1, characterized in that: The specific process of executing the stability test is as follows: After the user starts the SOST process, they select the stability type to be tested and confirm the test. SOST first checks and configures the reboot user self-start method in the Linux system, and then starts the server detailed information collection module, collecting detailed information about the server's hardware and software and recording it in the result folder. After the information collection is completed, execute the command to restart the server. After the restart, enter the system again. After the Linux system automatically logs in, it automatically starts the SOST process for the next stability test. Data transmission: SOST Server periodically obtains the stability test status of the current server from all clients. Every time SOST restarts and enters the system, the server collects detailed information about the stability test of the current client. Report generation: When the user stops the stability test in the Linux system or through the dashboard, the client executes the result_html() method to analyze the stability test results of the server, determine whether the current server stability test result is passed, and synchronously record the key server information in the stability test report. If the stability test result is a failure, all failure information is synchronously recorded in the stability test report.
6. A server system stability testing system as claimed in claim 1, characterized in that: The test types supported by the system include: Standard reboot test, reboot, systemctl reboot, shutdown - r; Power management operation test, power cycle, power reset, power failure recovery test ACLost; BMC remote power management test, including BMC Remote Power Reset, BMC Remote Power Cycle, BMC remote bmc reset warm, BMC remote bmc reset cold, and BMC remote raw 0x06 0x02. Other tests include a user-defined reboot command test, a multi-model test that runs four types of reboot, power cycle, power reset, and aclost tests in sequence, a BMC Sensor / SDR test that collects Sensor. / SDR information from the server's BMC and converts it to CSV format, and a BMC Quickly Create Sel Log and BMC Quickly Create Audit log.
7. A server system stability testing system as claimed in claim 1, characterized in that: The system also includes a visual interface for displaying the test process, result monitoring and problem diagnosis information, and provides real-time data display, automated report generation, task management and abnormal warning.
8. A server system stability testing system as claimed in claim 7, characterized in that: The core function implementation method in the visual interface includes: One-click stop test: When the test needs to be stopped, the background specifies the swc_stop_test.py program to pass in the system IP parameter, and continuously sends stop test requests to the Client IP. When the server stability is restarted and the swc service is running, swc receives the stop test request and modifies the stop_flag bit to 1. When the sost main program continues to perform the stability test, it checks the swc_flag bit. If it is 1, the test is stopped. Download test files. When downloading test files, the client code backs up the files generated by the current stability test type to the / tmp directory and packages them into a tar file. The backed-up tar test file is downloaded via HTTP. After the download is complete, the client actively deletes the backup file and the tar result file. Add the SOST client IP address. When the new client cannot be automatically identified by the server, customize the added ClientIP. The server will capture new client devices in the network in real time and record them in the device list file. It will check whether the added IP is alive and is a SOST client. If it meets the conditions, it will be added to the device list file. The server will then regularly collect client data.
9. A server system stability testing system as claimed in claim 8, characterized in that: The core function implementation method in the visual interface also includes: Delete useless lines. When the user deletes the specified client data, the server first deletes the client IP information from the device list file, and then deletes the file information recorded by the client on the server. Project filtering / test result filtering / test type filtering is implemented based on the HTML code on the server front end. After classifying all projects / results / test types, a drop-down menu option is provided for users to filter specific data.
10. A server system stability testing method, characterized in that: The following steps are involved: Data interaction, which realizes the communication connection between the client and the server through the SWC data interaction unit. The SWC data interaction unit includes SOSTClientPort service management, SOSTClientWeb service management, and data interfaces for providing stability information data, remotely stopping the test, remotely starting the test, downloading the test log, and viewing the stability history; Logging: Use the SOST log unit to provide complete logging services for SOST operation, and record all executed commands or error messages during SOST operation into files; Server information collection: The SOST server information collection unit and the SWC data interaction unit run together in the swc-manager.server service to collect server stability test type, CPU, and MEM information in real time, and record the collection results in a json file. The data is then transmitted to the server through the SWC data interaction unit; Stability test execution, in the SOST stability test unit, the user selects the stability test type. This unit checks whether the status assessment of the server before the stability test begins is suitable for stability testing. If the assessment passes, the stability test begins; Self-start, passing the SOST stability test. The self-start unit executes the SOST self-start program after each server restart; Log comparison evaluation, using the SOST log comparison unit, compares the results of the first-round device information with the current-round device information to evaluate whether the current stability test type has passed.