Intelligent pressure testing method for server hardware, computer device and storage device
Intelligent stress testing of server hardware is performed through the SSH protocol and multi-threading technology, which solves the problems of low efficiency, incomplete monitoring, non-intuitive results and insufficient data security in existing methods. It realizes efficient and automated server hardware testing and provides comprehensive real-time monitoring and self-healing capabilities.
Patent Information
- Application Number
- CN202511285269.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2025-10-17
AI Technical Summary
Existing server hardware stress testing methods are inefficient, unable to monitor key operating indicators in real time, and the test results are not intuitive. Data security is insufficient and there is a lack of automated processing mechanisms, which increases operation and maintenance costs and risks.
It executes tasks in parallel through the SSH protocol, uses multi-threading technology to stress test multiple devices, collects and encrypts key operating indicators, generates visual test results, and automatically restarts the device in abnormal situations to achieve self-healing capabilities.
It improves test efficiency, achieves comprehensive real-time monitoring and data security, reduces manual intervention, generates intuitive test results, and improves system reliability and response speed.
Smart Images

Figure CN120803878A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of server testing, in particular to a server hardware intelligent stress testing method, a computer device and a storage medium. BACKGROUND
[0002] With the rapid development of information technology, servers are increasingly widely used in various fields, and their hardware performance and stability are crucial to the reliable operation of the system. Stress testing of server hardware is an important means to evaluate its running ability under high load and extreme conditions, which can help to identify potential problems, optimize system configuration and ensure stable and efficient operation of servers in actual use.
[0003] However, the existing server hardware stress testing methods have many shortcomings. On the one hand, traditional stress testing usually requires manual operation, which is inefficient and difficult to meet the testing needs of large-scale server clusters. On the other hand, the monitoring of server hardware during the testing process is not comprehensive enough, and it is difficult to accurately locate the problem without real-time access to key operating indicators. In addition, the presentation of test results is relatively simple, lacking in intuitiveness and readability, making it difficult for users to quickly understand and analyze.
[0004] In terms of data security, the existing methods do not adequately protect the key operating indicator data collected during the testing process, which can easily lead to the leakage of sensitive information. At the same time, there is a lack of automated handling mechanism for abnormal situations found during testing, which requires manual intervention, increasing the operation and maintenance cost and testing risk. SUMMARY
[0005] The present application provides a server hardware intelligent stress testing method, a computer device and a storage medium, which aims to at least solve one of the technical problems existing in the prior art.
[0006] The technical solution of the present application is a server hardware intelligent stress testing method, which includes: Obtaining the configuration file of the server hardware, the configuration file of the server hardware including the basic information of a plurality of devices to be managed; Executing a preset task command on the plurality of devices to be managed through the SSH protocol, and using multi-threading technology to execute the same task on the plurality of devices to be managed in parallel; Generating a stress testing command, and applying a stress testing load to the plurality of devices to be managed in parallel according to the stress testing command; Collecting system information of the plurality of devices to be managed in parallel to obtain key operating indicators corresponding to each device to be managed; Encrypting the key operating indicators and storing them in a device-specific log file; Traverse the key operation indexes corresponding to the plurality of to-be-managed devices, judge whether there is an abnormal key operation index, if there is an abnormal key operation index, restart the to-be-managed device corresponding to the abnormal key operation index through an HTTP POST request; Collect the hardware information of the plurality of to-be-managed devices and store in a directory dedicated to the device, analyze the hardware information of the plurality of to-be-managed devices, and generate a visual server hardware intelligent stress test result.
[0007] According to some embodiments of the application, the preset task command is executed on the plurality of to-be-managed devices through the SSH protocol, and the same task is executed on the plurality of to-be-managed devices in parallel using a multi-threading technology, including: Load data from the configuration file of the server hardware using the json.load() method, and extract the devices variable field therein as a device list, the device list containing a plurality of to-be-managed device result dictionaries, each of which represents information of a to-be-managed device; Initialize a thread list and the result dictionary, traverse the device list, create an independent sub-thread for each to-be-managed device, call a first function to generate the preset task command, establish an SSH connection with the target management device through paramiko.SSHClient, use a multi-threading technology to execute the same task on a plurality of target management devices in parallel, and return a result dictionary corresponding to the target management device; Wait for all sub-threads to execute the task, and end the task execution operation.
[0008] According to some embodiments of the application, the stress test command is generated, and the stress test load is applied to the plurality of to-be-managed devices in parallel according to the stress test command, including: Generate a stress test command for the specific hardware type of the to-be-managed device; Call a second function to return a command string, use the stress tool to perform stress testing, and simulate a high-load scenario; Apply the stress test command to all to-be-managed devices in parallel through a third function call to implement the stress test load applied to the plurality of to-be-managed devices in parallel.
[0009] According to some embodiments of the application, the system information of the plurality of to-be-managed devices is collected in parallel to obtain the key operation index corresponding to each to-be-managed device, including: Collect the system information of a single to-be-managed device through an SSH command, and distribute the system information collection task to all to-be-managed devices through a fourth function call to collect the task in parallel; A command string is returned by the fifth function, combining the top command and the curl command to obtain the key operating indicators of each of the to-be-managed devices, and returning the results.
[0010] According to some embodiments of the application, the encryption processing of the key operating indicators and storage in the device-specific log file include: The data of the key operating indicators is encrypted using a preset encryption algorithm, and is saved in a device-specific log file named after the to-be-managed device; The encryption processing results of all to-be-managed devices are traversed, and the encrypted log files of all to-be-managed devices are stored.
[0011] According to some embodiments of the application, the collection of the hardware information of the plurality of to-be-managed devices and storage in the device-specific directory include: A command string is returned by calling the sixth function, combining the lspci, lsblk, and nvme list commands to collect the hardware information of the plurality of to-be-managed devices; A hardware log file is created for each to-be-managed device, and the collected hardware information of each to-be-managed device is stored in the corresponding hardware-specific log file.
[0012] According to some embodiments of the application, the parsing of the hardware information of the plurality of to-be-managed devices and the generation of the visual server hardware intelligent stress test results include: The encrypted key operating indicator data is extracted from the device-specific log file, and the encrypted key operating indicator data is decrypted to obtain decrypted key operating indicator data; The hardware information of the to-be-managed device is extracted from the hardware-specific log file; The visual server hardware intelligent stress test results are generated according to the decrypted key operating indicator data and the hardware information of the to-be-managed device.
[0013] According to some embodiments of the application, the generation of the visual server hardware intelligent stress test results according to the decrypted key operating indicator data and the hardware information of the to-be-managed device includes: A report directory is created, and an HTML template is defined to generate the structure and style of the test result report; The device list is traversed, and the seventh function is called for each to-be-managed device in the device list, the chart path is added to the HTML template, and the placeholder in the HTML template is replaced; According to the decrypted data of the key operation index and the hardware information of the to-be-managed device, a test result table and a chart of each to-be-managed device are dynamically generated based on the HTML template; The test result table and the chart of each to-be-managed device are saved as an HTML file, and a file path is output, so that a visual server hardware intelligent stress test result is obtained.
[0014] The technical scheme of the present application also relates to a computer device, comprising a memory and a processor, wherein the processor executes a computer program stored in the memory to implement the above method.
[0015] The technical scheme of the present application also relates to a computer readable storage medium, which stores computer program instructions, wherein the computer program instructions are executed by a processor to implement the above method.
[0016] The server hardware intelligent stress test method, the computer device and the storage medium provided by the embodiment of the present application have at least one of the following advantages or beneficial effects: the configuration file of the server hardware is obtained, the configuration file of the server hardware includes the basic information of a plurality of to-be-managed devices, the configuration of the server hardware can be centrally managed, the device list can be conveniently dynamically expanded or modified, different stress loads can be allocated according to the configuration differences, the test precision is improved, a preset task command is executed on the plurality of to-be-managed devices through an SSH protocol, the SSH protocol realizes agentless remote control, the same task is simultaneously issued to the plurality of to-be-managed devices by using a multi-thread parallel technology, multi-device parallel operation is realized, the task efficiency is significantly improved, the test period is greatly shortened, and manual intervention is reduced. According to the configuration file, a stress test command is dynamically generated, the stress test command is issued to all the to-be-managed devices in parallel, extreme scenarios such as CPU full load, memory leakage, disk IO full, and network congestion are simulated, stress test load is realized, and the stability problem of hardware under extreme load can be quickly exposed. The system information of the plurality of to-be-managed devices is collected in parallel to obtain the corresponding key operation indicators of each to-be-managed device, the monitoring capability of the system of the to-be-managed device is provided, the corresponding key operation indicators of each to-be-managed device are obtained, and it is ensured that the running state of each to-be-managed device can be grasped in real time. The key operation indicators are encrypted and stored in a device-specific log file, sensitive information in the log file is prevented from being leaked, and the security of the log file is enhanced. The key operation indicators corresponding to the plurality of to-be-managed devices are traversed, it is judged whether there is an abnormal key operation indicator, if there is an abnormal key operation indicator, an HTTP POST request is automatically triggered to restart the to-be-managed device corresponding to the abnormal key operation indicator, the abnormal device is avoided from affecting the overall cluster performance, the self-healing capability is realized, the automatic abnormal detection and processing mechanism is realized, the need for manual intervention is reduced, and the reliability and response speed of the system are improved. The hardware information of the plurality of to-be-managed devices is collected and stored in a device-specific directory, the hardware information of the plurality of to-be-managed devices is parsed, and a visual server hardware intelligent stress test result is generated, and the visualization degree of the test result is improved.
[0017] In addition, additional aspects and advantages of the present application will be described in part below, will become apparent from the following description, or will be learned by practice of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 is a general flowchart of a server hardware intelligent stress test method provided by the embodiment of the present application; Figure 2 is a detailed flowchart of step S200 in the server hardware intelligent stress test method provided by the embodiment of the present application; Figure 3is a detail flow chart of step S300 in the server hardware intelligent stress test method provided by the embodiment of the present application; Figure 4 is a detail flow chart of step S300 in the server hardware intelligent stress test method provided by the embodiment of the present application; Figure 5 is a detail flow chart of step S500 in the server hardware intelligent stress test method provided by the embodiment of the present application; Figure 6 is a first detail flow chart of step S700 in the server hardware intelligent stress test method provided by the embodiment of the present application; Figure 7 is a second detail flow chart of step S700 in the server hardware intelligent stress test method provided by the embodiment of the present application; Figure 8 is a display diagram of the visual stress test result of the first to-be-managed device provided by the embodiment of the present application; Figure 9 is a display diagram of the visual stress test result of the second to-be-managed device provided by the embodiment of the present application; Figure 10 is a detail flow chart of step S750 in the server hardware intelligent stress test method provided by the embodiment of the present application. DETAILED DESCRIPTION
[0019] The concept, specific structure and generated technical effects of the present application will be described clearly and completely in combination with embodiments and drawings, so as to fully understand the purpose, scheme and effects of the present application.
[0020] It should be noted that, unless otherwise specified, when a certain feature is referred to as being "fixed" or "connected" to another feature, it can be directly fixed or connected to the other feature, or indirectly fixed or connected to the other feature. The singular forms "a", "an" and "the" used in this text are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise defined, all technical and scientific terms used in this text have the same meaning as understood by those skilled in the art. The terms used in the specification of this text are only used to describe specific embodiments, and are not intended to limit the present application. The term "and / or" used in this text includes any combination of one or more related listed items.
[0021] It should be understood that, although the terms first, second, third, etc. can be employed in describing various elements in the application, the elements should not be limited to these terms. These terms are only used to distinguish one type of element from another. For example, a first element could also be termed a second element, and, similarly, a second element could also be termed a first element, without departing from the scope of the application. The use of any and all examples, or exemplary language (e.g., "such as", "for instance", etc.) provided herein is intended merely to better illuminate embodiments of the application and does not pose a limitation on the scope of the application unless otherwise claimed.
[0022] With the rapid development of information technology, servers are increasingly widely used in various fields, and their hardware performance and stability are crucial for reliable operation of the system. Pressure testing of server hardware is an important means to evaluate its running ability under high load and extreme conditions, which can help to find potential problems in advance, optimize system configuration, and ensure that the server can run stably and efficiently in actual use.
[0023] However, the existing server hardware pressure testing method has many shortcomings. On the one hand, traditional pressure testing usually needs manual operation, which is inefficient and difficult to meet the testing needs of large-scale server clusters. On the other hand, the monitoring of server hardware during the testing process is not comprehensive, and it is difficult to accurately locate the problem. In addition, the presentation form of the test results is relatively single, lacking intuitiveness and readability, which is not convenient for users to quickly understand and analyze.
[0024] In terms of data security, the existing method lacks protection for the key running indicator data collected during the testing process, which can easily lead to sensitive information leakage. At the same time, for the abnormal situations found during the testing, there is a lack of automatic processing mechanism, which requires manual intervention, increasing the operation and maintenance cost and testing risk.
[0025] Therefore, the embodiments of the present application provide a server hardware intelligent pressure testing method, a computer device and a storage medium, which can efficiently and automatically test the server hardware, comprehensively monitor the key running indicators, protect the data security, and present the test results in a visual way, to meet the needs of modern server hardware testing.
[0026] Please refer to Figures 1 to 10 The server hardware intelligent pressure testing method, the computer device and the storage medium provided by the embodiments of the present application are further described.
[0027] Please refer to Figure 1 As shown in the figure, Figure 1 is a general flowchart of a server hardware intelligent pressure testing method provided by the embodiments of the present application, which includes but is not limited to steps S100 to S700, including: S100: Obtain a configuration file of server hardware, the configuration file of server hardware including basic information of a plurality of to-be-managed devices; S200: Execute a preset task command on the plurality of to-be-managed devices through an SSH protocol, and perform the same task on the plurality of to-be-managed devices in parallel by using a multithreading technology; S300: Generate a stress test command, and apply a stress test load to the plurality of to-be-managed devices in parallel according to the stress test command; S400: Collect system information of the plurality of to-be-managed devices in parallel to obtain a corresponding key operation index of each to-be-managed device; S500: Perform encryption processing on the key operation index, and store the key operation index in a device-specific log file; S600: Traverse the key operation index corresponding to the plurality of to-be-managed devices, determine whether there is an abnormal key operation index, and if there is an abnormal key operation index, restart the to-be-managed device corresponding to the abnormal key operation index through an HTTP POST request; S700: Collect hardware information of the plurality of to-be-managed devices and store the hardware information in a device-specific directory, analyze the hardware information of the plurality of to-be-managed devices, and generate a visual server hardware intelligent stress test result.
[0028] In some embodiments of the present application, the server hardware intelligent stress test method comprises: obtaining a configuration file of server hardware by reading a JSON format file devices.json, the configuration file of server hardware including basic information of a plurality of to-be-managed devices, wherein the basic information of the to-be-managed devices includes IP addresses, usernames, passwords and the like of the to-be-managed devices, and the basic information of the to-be-managed devices is stored in a device list, and the basic information of the to-be-managed devices is the basis for subsequent operations, for connecting to each device and performing tasks. By storing the basic information of the to-be-managed devices in the configuration file of server hardware, the configuration of server hardware can be centrally managed, the device list can be dynamically expanded or modified, different stress loads can be allocated according to configuration differences, test accuracy can be improved, and a unified data source is provided to ensure that subsequent operations can be performed based on consistent information.
[0029] The preset task command is executed on the plurality of to-be-managed devices through an SSH protocol, the SSH protocol realizes agentless remote control, is compatible with Linux / Unix systems, and has high security; the same task is simultaneously issued to the plurality of to-be-managed devices by using a multithreading parallel technology, multi-device parallel operation is realized, task efficiency is significantly improved, test period is greatly shortened, manual intervention is reduced, and automation level is improved; in addition, the operation of each to-be-managed device is independently run by the multithreading mechanism, and does not interfere with each other.
[0030] The stress test command is dynamically generated according to the configuration file, and the stress test command is issued to all devices to be managed in parallel, simulating extreme scenarios such as CPU full load, memory leakage, disk IO full, network congestion, etc., realizing stress test load, and quickly exposing the stability problems (such as overheating, reboot, frequency reduction) of hardware under extreme load; providing limit value reference (such as maximum QPS, IO throughput) for capacity planning. In addition, the stress test command of the embodiment of the present application supports flexible stress test scenarios, and users can adjust the stress test command parameters (such as load type, duration) to adapt to different test requirements.
[0031] The system information of a plurality of devices to be managed is collected in parallel to obtain the corresponding key running indicators of each device to be managed, including collecting CPU temperature, memory usage, disk information, network packet loss rate and other indicators, forming a fine-grained performance baseline, facilitating comparison of abnormalities, and providing data support for subsequent fault location, such as when a device to be managed server leaks memory, it can be traced back to a specific time point. Collecting system information of a plurality of devices to be managed provides comprehensive monitoring capability of the system of the device to be managed, obtains the corresponding key running indicators of each device to be managed, and ensures that the running state of each device to be managed can be grasped in real time; the parallel collection mechanism improves the efficiency of information collection.
[0032] The key running indicators are encrypted and stored in a log file dedicated to the device, and symmetric encryption (such as AES) or certificate encryption (such as RSA) is used to protect the log to prevent sensitive information such as hardware model, vulnerability indicators, etc. in the log file from being leaked. Data encryption enhances the security of the log file, prevents test data from being tampered with, and ensures auditability. By storing the key running indicators in a dedicated log according to device directory, it is convenient for classification management and analysis.
[0033] The key running indicators corresponding to a plurality of devices to be managed are traversed to determine whether there are abnormal key running indicators, and if there are abnormal key running indicators, an HTTP POST request is automatically triggered to restart the device to be managed corresponding to the abnormal key running indicators, avoiding the influence of abnormal devices on the overall cluster performance, realizing self-healing capability, automatic abnormal detection and processing mechanism, reducing the need for manual intervention, and improving the reliability and response speed of the system.
[0034] The hardware information of a plurality of devices to be managed is collected, such as collecting BIOS version, firmware version, hard disk power-on times and other static information, and stored in a directory dedicated to the device, the hardware information of a plurality of devices to be managed is parsed, and a visual server hardware intelligent stress test result is generated, and a chart or PDF report is generated combining the test result and the hardware information, which can intuitively display the hardware intelligent stress test result of each device to be managed, and improve the visualization degree of the test result.
[0035] Therefore, the server hardware intelligent stress test method provided by the embodiments of the present application integrates the whole process automation of the centralized management of the to-be-managed devices, stress test, monitoring collection, exception handling, encrypted storage, and visual report, constructs a complete closed-loop test process of the server hardware intelligent stress test, realizes end-to-end automation through a unified framework, and performs outstandingly in the combination of multi-device concurrent control and abnormal self-healing. The dynamically scalable parallel execution engine realizes lightweight parallel SSH control, a multi-dimensional hardware stress test matrix, and covers four dimensions of computing / storage / network / acceleration card. The key operation indexes collected are stored in the device-specific log files after being encrypted, which not only ensures the security of the data but also facilitates subsequent classified management and analysis, and realizes efficient data processing and storage. The performance indexes of the to-be-managed devices are clearly displayed in an intuitive way through the generated report and chart, which facilitates users to view and analyze the test results and helps to better understand and evaluate the performance of the devices.
[0036] It should be noted that in step S300 of the embodiments of the present application, the stress test command is generated, and stress test load is applied to the plurality of to-be-managed devices in parallel according to the stress test command. The stress test command includes but is not limited to a CPU stress test command, a memory stress test command, a disk I / O stress test command, and a network stress test command, which realizes comprehensive stress test on the performance of the server hardware and ensures the normal operation of the server.
[0037] In one embodiment, the memory stress test includes: The memory stress test is realized by using memtester, which is a tool for testing the stability of memory. The memory stress test is realized in the following way: def generate_memory_load(device): Generate memory stress on the device (use memtester) return "memtester 320G 50" # Test 320GB memory for 50 cycles In one embodiment, the disk I / O stress test includes: The disk I / O stress test is realized by using fio, which is a powerful tool for testing disk I / O performance. The disk I / O stress test is realized in the following way: def generate_disk_io_load(device): Generate disk I / O stress on the device (use fio) return ("fio --name=iotest --ioengine=libaio --rw=randwrite --bs=4k --direct=1 " "--size=1G --numjobs=4 --runtime=60 --group_reporting") In one embodiment, the network stress test includes: The network stress test is implemented using iperf3, which is a network performance testing tool that supports multiple protocols such as TCP, UDP, etc., and implements the network stress test in the following way: def generate_network_load(device, server_ip="192.168.1.1"): """Generate network stress on the device (using iperf3)""" return f"iperf3 -c {server_ip} -t 600" # Client connects to the specified server IP for 600 seconds3.4 GPU stress test - use gpu-burn, gpu-burn is a simple stress testing tool designed for NVIDIA GPU.
[0038] def generate_gpu_load(device): """Generate GPU stress on the device (using gpu-burn)""" return "gpu-burn 6000" # Run 6000 seconds of GPU compute-intensive tasks Through CPU stress testing, memory stress testing, disk I / O stress testing, network stress testing, and other aspects of hardware stress testing on the performance of the server hardware, the normal operation of the server is ensured.
[0039] Referring to Figure 2 , Figure 2 is a detailed flowchart of step S200 in the server hardware intelligent stress testing method provided by the embodiments of the present application, and step S200 includes but is not limited to steps S210 to S230, specifically, S210: Use the json.load() method to load data from the configuration file of the server hardware, and extract the devices variable field therein as a device list, which contains a plurality of result dictionaries of the to-be-managed devices, and each result dictionary represents the information of a to-be-managed device; S220: initialize the thread list and the result dictionary, traverse the device list, create an independent sub-thread for each device to be managed, call the first function to generate a preset task command, establish an SSH connection with the target management device through paramiko.SSHClient, use multi-threading technology to perform the same task on several target management devices in parallel, and return the result dictionary corresponding to the target management device; S230: wait for all sub-threads to complete the task, end the task execution operation.
[0040] In some embodiments of the application, the preset task command is executed on several devices to be managed through the SSH protocol, and the same task is executed on several devices to be managed in parallel using multi-threading technology, including: The json.load() method is used to load data from the configuration file of the server hardware, and the json.load() method is used to read JSON data from the configuration file object and parse it into a dictionary or list data structure that can be operated by Python, thereby facilitating the extraction and use of configuration information such as the devices variable field, achieving efficient reading and parsing of the configuration file, enabling the program to flexibly obtain and process configuration information, and providing basic data support for subsequent device management operations. The devices variable field in the configuration file is extracted as a device list, which contains the result dictionary of several devices to be managed, and each result dictionary represents the information of a device to be managed, clearly defining the scope and related information of the devices to be managed, providing a clear object collection for subsequent independent operations on each device to be managed, and facilitating the program to manage and assign tasks to each device to be managed.
[0041] The thread list is used for storing the sub-threads created for each device to be managed, facilitating management and control of the threads; the result dictionary is used for storing the results after the execution of each device task, and the thread list and the result dictionary are initialized, providing data structure support for the execution of multi-threaded tasks and the collection of results, so that the program can effectively organize and track the running of multiple threads and the execution results of tasks. By traversing the device list, an independent sub-thread is created for each device to be managed, realizing simultaneous operation of multiple devices to be managed, improving the efficiency of task execution and avoiding the time-consuming problem caused by sequential execution of tasks for each device to be managed. The first function is called, which is run_on_device(device, task_func), and the preset task command is generated for each device to be managed, ensuring that the tasks executed by each device to be managed are consistent and standardized, and targeted task commands can be generated according to the specific information of the device to be managed. The SSH connection with the target management device is established through paramiko.SSHClient, so that the program can manage and operate the remote device to be managed through the network, breaking through the limitation of local operation and realizing convenient management of remote devices. By using multi-threading technology, multiple sub-threads can run simultaneously, and each sub-thread is responsible for executing the same task on a target management device, greatly improving the efficiency of task execution and enabling the management operation of multiple devices to be managed to be completed in a short time. After each sub-thread completes the task, the execution result is stored in the result dictionary corresponding to the target management device, facilitating the collection and subsequent processing of the task execution results, so that the program can clearly understand the execution of each device task.
[0042] Wait for all sub-threads to complete the task, end the task execution operation, and ensure the complete execution of the task by waiting for all sub-threads to complete, avoiding incomplete or incorrect results caused by incomplete tasks of some threads, and ensuring the reliability and integrity of task execution.
[0043] In some embodiments of the present application, the preset task command is executed on several devices to be managed through the SSH protocol, and the host key policy is set during the parallel execution of the same task on several devices to be managed using multi-threading technology. It can be understood that the SSH connection needs to verify the host key of the target management device, and here the AutoAddPolicy policy is set, specifically, Establish an SSH connection client.connect( hostname=device["ip"], username=device["username"], password=device["password"] ) Call the connect method to establish an SSH connection with the target management device.
[0044] Parameter Description: hostname: The IP address of the device, obtained from the device dictionary.
[0045] Username: Username for logging into the device, obtained from the device dictionary.
[0046] password: The password for logging into the device, obtained from the device dictionary.
[0047] If the connection fails (for example, due to network problems or authentication failure), the connect method will throw an exception.
[0048] Execute the task command and call the exec_command method to execute the incoming task command (task_func) on the target management device and return stdout for receiving the normal output of the command, or stderr for receiving the error output of the command.
[0049] After that, get the command output, read the contents of the standard output stream, and decode it into a string. The read() method returns a byte stream, so you need to call .decode() to convert it to a string format. Finally, after the task is completed, close the SSH connection, release resources, and return the command output to the caller.
[0050] The above process realizes the parallel operation of multiple devices to be managed, significantly improving task efficiency. Through the thread mechanism, the operation of each device runs independently without interfering with each other.
[0051] Reference Figure 3 As shown, Figure 3 This is a detailed flow chart of step S300 in the server hardware intelligent stress testing method provided by an embodiment of the present invention. Step S300 includes but is not limited to steps S310 to S330. Specifically, S310: Generate a stress test command for a specific hardware type of the device to be managed; S320: Call the second function to return a command string, and use the stress tool to perform stress testing to simulate a high-load scenario; S330: calling the stress test command through the third function, and applying the stress test command to all the devices to be managed in parallel, so as to achieve parallel application of stress test loads to several devices to be managed.
[0052] In some embodiments of the present application, the stress test command is generated, and the stress test load is applied to the plurality of managed devices in parallel according to the stress test command, which includes: generating a targeted stress test command according to the specific hardware type (such as CPU, memory, disk I / O, etc.) of the managed device to simulate a high load scenario. This step can accurately perform stress testing on the specific hardware resources of the target device, facilitating the discovery of performance bottlenecks and potential problems of the hardware under high load, and providing a basis for system optimization and stability evaluation.
[0053] By calling the second function, generate_load, a command string for stress testing is generated, and the stress tool is used to simulate a high load scenario. Various stress test commands can be generated according to different devices and test requirements to meet various test scenarios. The stress tool can simulate high load situations for multiple resources such as CPU, memory, and disk I / O, making the test results closer to the actual running environment. In addition, the stress tool provides a simple command line interface, making it easy to operate and integrate into automated test scripts. Then, The stress test command generated by the third function call, apply_load_on_devices, is applied to all managed devices using multi-threading or parallel processing technology, implementing parallel application of stress test load. By executing stress test commands in parallel, multiple managed devices can be tested simultaneously, greatly reducing test time and allowing for efficient use of test resources in a short period of time to quickly obtain test results for a large number of managed devices.
[0054] Referring to Figure 4 , Figure 4 is a detailed flowchart of step S300 in the server hardware intelligent stress testing method provided by the embodiments of the present application. Step S400 includes but is not limited to steps S410 to S420, specifically, S410: Collect system information of a single managed device through an SSH command, and distribute system information collection tasks to all managed devices through a fourth function call parallel collection task. S420: Return a command string through a fifth function, combine the top command and the curl command to obtain key performance indicators of each managed device, and return the results.
[0055] In some embodiments of the present application, collecting system information of several to-be-managed devices in parallel to obtain the corresponding key operating indicators of each to-be-managed device includes: remotely connecting to a single to-be-managed device through an SSH command and executing a command to collect its system information (such as CPU usage, memory usage, disk space, etc.). SSH is used to access and operate remote devices without direct contact with the to-be-managed device, improving the convenience and flexibility of operation. Real-time acquisition of system information of to-be-managed devices provides the latest data for subsequent monitoring and analysis. The SSH protocol provides an encrypted communication channel to ensure the security of data transmission. Through the fourth function, collect_system_info, the fourth function is used to call and collect parallel tasks, and the system information collection task is distributed to all to-be-managed devices to achieve parallel collection. Parallel collection tasks can simultaneously collect system information of multiple to-be-managed devices, enabling the acquisition of system information of all devices in a short period of time, facilitating rapid evaluation of the running state of the entire system.
[0056] A command string is generated by the fifth function, collect_system_info, which combines the top command and the curl command to obtain the key operating indicators of each to-be-managed device and returns the results. The returned results contain multiple key indicators, providing rich data support for system performance analysis and troubleshooting.
[0057] The top command can display the resource occupation of each process in the system in real time, and the curl command can be used to obtain the status or data of network services. The combination of the two can comprehensively obtain the key operating indicators of the device, real-time acquisition of device operating indicators, and provide comprehensive system monitoring capabilities to ensure real-time monitoring of device operating status, facilitating timely detection of abnormal conditions and processing.
[0058] In some embodiments of the present application, the key operating indicators of the to-be-managed device include but are not limited to: CPU: Load (system load in the past 1 minute, 5 minutes, 15 minutes), system state (percentage of time spent by CPU in system state).
[0059] Memory: Total memory (total physical memory size of the system), used memory (amount of physical memory used).
[0060] PDU: Power state (power-on state), voltage (input or output voltage), current (current intensity through PDU port), energy consumption data (real-time power consumption or cumulative electricity consumption).
[0061] PCI device: Device type (such as network card, video card, etc.), device ID (used to identify specific hardware model).
[0062] Block device: device name (such as / dev / sda), device size (disk capacity).
[0063] NVMe device: device name (such as / dev / nvme0n1) and health status (such as remaining lifespan).
[0064] Real-time acquisition of operating indicators of managed devices provides comprehensive system monitoring capabilities, ensuring real-time understanding of the operating status of devices, facilitating timely detection and handling of abnormal situations. Multiple key indicators provide rich data support for system performance analysis and troubleshooting.
[0065] Reference Figure 5 As shown, Figure 5 This is a detailed flow chart of step S500 in the server hardware intelligent stress testing method provided by an embodiment of the present invention. Step S500 includes but is not limited to steps S510 to S520. Specifically, S510: Encrypt the key operating indicator data using a preset encryption algorithm and save it in a device-specific log file named after the device to be managed; S520: traverse the encryption processing results of all the devices to be managed, and store the encrypted log files for all the devices to be managed.
[0066] In some embodiments of the present invention, encrypting the key operating indicators and storing them in a device-specific log file includes: encrypting the collected key operating indicator data using a preset encryption algorithm, and even if the encrypted key operating indicator data is illegally obtained, the content cannot be directly read, thereby protecting the confidentiality of the key operating indicator data. The encryption algorithm has tamper-proof properties and can ensure the integrity of the data during storage and transmission. The encrypted key operating indicator data is saved in a log file named after the device to be managed. By naming the log file after the device to be managed, the log file of a specific device can be quickly located, which is convenient for management and query. The operating indicator data of each device to be managed is stored independently, avoiding confusion and interference between data. When it is necessary to analyze the operating status of a certain device, the log file of the device can be directly read, which improves the efficiency and convenience of analysis.
[0067] The system traverses the encryption results of all managed devices and stores the encrypted log files for them. This traversal and storage process automates the processing of log files for all devices, reducing manual intervention and improving efficiency. It ensures that all log files for managed devices are correctly stored to avoid omissions. The stored log files can be used for subsequent log analysis, auditing, and backup operations, providing support for system operation and management.
[0068] The combined use of the above technical features can effectively protect the security, integrity and confidentiality of key operating indicator data. At the same time, through reasonable naming and storage methods, it facilitates the management and analysis of log files. It is particularly suitable for scenarios with high data security requirements and can ensure the security and reliability of data during collection, processing and storage.
[0069] Reference Figure 6 As shown, Figure 6 This is a first detailed flow chart of step S700 in the server hardware intelligent stress testing method provided by an embodiment of the present invention. Step S700 includes but is not limited to steps S710 to S720. Specifically, S710: Calling the sixth function to return a command string, combining the lspci, lsblk, and nvme list commands to collect hardware information of several devices to be managed; S720: Create a hardware log file for each device to be managed, and store the collected hardware information of each device to be managed in the corresponding hardware-specific log file.
[0070] In some embodiments of the present invention, collecting hardware information of several devices to be managed and storing it in a device-specific directory includes: calling the sixth function to return a command string, the sixth function is collect_hardware_info. The sixth function is used in combination with the lspci, lsblk and nvme list commands to collect hardware information of the devices to be managed. The lspci command is used to list all PCI devices in the system, including graphics cards, network cards, storage controllers, etc. The lsblk command is used to list information of all block devices, including hard disks, partitions, mount points, etc. The nvme list command is used to list information of all NVMe devices, and is particularly suitable for solid-state drives (SSDs). Combining these three commands, the hardware information of the device can be comprehensively collected, covering all aspects from general PCI devices to storage devices. The sixth function can generate different command strings as needed, and flexibly adjust the scope of the collected hardware information. By generating command strings, these commands can be easily integrated into automated scripts to realize batch hardware information collection for multiple devices.
[0071] A hardware log file is created for each managed device. Each device's hardware information is stored independently, avoiding data confusion and interference, facilitating subsequent analysis and management. Creating a separate log file for each device allows for quick location of specific device hardware information, improving management efficiency. When analyzing a device's hardware configuration, the device's hardware log file can be directly read, allowing for quick problem identification and troubleshooting. Storing hardware information in log files ensures data persistence, facilitating subsequent historical data comparison and auditing.
[0072] The combined use of these technical features enables efficient and comprehensive collection of hardware information for multiple managed devices. Through appropriate storage methods, this data can be isolated, managed, and analyzed easily. This makes it suitable for scenarios requiring hardware information collection and management for a large number of devices, such as data center equipment management and maintenance. This approach allows for rapid understanding of each device's hardware configuration, timely identification of hardware issues, and support for stable system operation.
[0073] Reference Figure 7 As shown, Figure 7 This is a second detailed flow chart of step S700 in the server hardware intelligent stress testing method provided by an embodiment of the present invention. Step S700 includes but is not limited to steps S730 to S740. Specifically, S730: Extracting the encrypted key operating indicator data from the device-specific log file, and decrypting the encrypted key operating indicator data to obtain decrypted key operating indicator data; S740: Extracting hardware information of the device to be managed from the hardware-specific log file; S750: Generates visualized server hardware intelligent stress test results based on the decrypted data of key operating indicators and the hardware information of the devices to be managed.
[0074] In some embodiments of the present invention, a method for generating visualized server hardware intelligent stress test results includes: extracting encrypted key operating indicator data stored in device-specific log files by reading them; decrypting the encrypted data according to the encryption algorithm (e.g., Advanced Encryption Standard (AES)) and key, wherein the decryption process involves determining a decryption function corresponding to the encryption algorithm; and decrypting the encrypted data using the same key as used for encryption, thereby obtaining the original key operating indicator data. Hardware information for the managed devices is extracted from the hardware-specific log files, and the hardware-specific log files corresponding to each managed device are located. Hardware information, including CPU model, memory capacity, storage device type, and so on, is parsed from the log files.
[0075] Integrate the decrypted key performance indicator data with the hardware information to form a complete dataset. Use visualization tools such as the HTML report generated by JMeter or custom visualization scripts to convert the integrated data into intuitive charts or reports. Refer to Figure 8 and Figure 9 , Figure 8 is a display chart of the first to-be-managed device's visualized stress test result (first server hardware intelligent stress test result) provided by the embodiments of the present application, Figure 9 is a display chart of the second to-be-managed device's visualized stress test result (second server hardware intelligent stress test result) provided by the embodiments of the present application, Figure 8 and Figure 9 respectively show the real-time data and historical trends of the key indicators such as CPU usage, memory occupancy, disk I / O, network throughput, and GPU usage of the first to-be-managed device and the second to-be-managed device, realizing key performance indicator visualization. The hardware configuration information such as CPU core number, memory capacity, and hard disk type is displayed in the form of tables or charts, realizing hardware information visualization.
[0076] Through the above technical features, the encrypted data can be safely processed and the hardware information can be effectively utilized, and finally the visualized server hardware intelligent stress test result is generated. This not only improves the security and confidentiality of data, but also enhances the understanding and analysis ability of the test result through visualization means, providing strong support for the optimization and management of server hardware.
[0077] as shown in Figure 10 , Figure 10 is a detailed flowchart of step S750 in the server hardware intelligent stress test method provided by the embodiments of the present application, and step S750 includes but is not limited to steps S751 to S754, specifically, S751: Create a report directory and define an HTML template to generate the structure and style of the test result report; S752: Traverse the device list, call the seventh function for each to-be-managed device in the device list, add the chart path to the HTML template, and replace the placeholder in the HTML template; S753: Based on the decrypted data of the key performance indicators and the hardware information of the to-be-managed device, dynamically generate the test result table and chart of each to-be-managed device based on the HTML template; S754: Save the test result table and chart of each to-be-managed device as an HTML file and output the file path to obtain the visualized server hardware intelligent stress test result.
[0078] In some embodiments of the application, the server hardware intelligent stress test result visualization is generated according to the decrypted key performance indicator data and the hardware information of the device to be managed, including: creating a special directory for storing the generated test result report file, facilitating management and searching. The creation of the report directory enables the centralized storage of all generated HTML files, facilitating subsequent management and searching, defining an HTML template for the structure and style of the test result report, and the HTML template ensures the consistency of the format and style of all test result reports, facilitating user reading and understanding. Through the templating method, the style of the report can be conveniently unified and maintained, improving maintainability.
[0079] The device list is traversed, and the corresponding test result of each device to be managed is generated. A seventh function, generate_chart, is called to generate a test result chart for each device and obtain the path of the chart. The chart path is added to the HTML template and replaces the placeholder in the template to generate complete HTML content. Through the traversal of the device list and the calling of the function, the automation of the test result generation of multiple devices to be managed is realized. According to the actual test result of each device, the HTML content is dynamically generated to ensure the accuracy and real-time performance of the report. Through the placeholder replacement mechanism, the test results of different devices to be managed can be flexibly inserted into the template to adapt to different test scenarios.
[0080] The decrypted key performance indicator data and hardware information are combined to generate a test result table and chart for each device to be managed. The HTML template is used as the basis to dynamically fill in the data and generate a complete test result report. The key performance indicators and hardware information are combined to generate a comprehensive test result report, providing rich data support for users; through the generation of charts, complex data is displayed in an intuitive way, facilitating quick understanding and analysis by users; the report is dynamically generated according to the actual data of each device to be managed, ensuring the accuracy and real-time performance of the report.
[0081] The generated test result table and chart are saved as an HTML file, and the test result is saved as an HTML file, realizing the persistent storage of data and facilitating subsequent viewing and analysis; the path of the output HTML file is output, facilitating users to quickly locate and access the report, and improving user experience. The generated HTML file contains charts and tables, which displays the server hardware intelligent stress test result in an intuitive way.
[0082] Through the above technical means, the test result report of each to-be-managed device can be efficiently generated, the report is saved in an HTML format, and a table of key operation indexes, a table of hardware information, and corresponding charts are included. Such a scheme not only improves the automation degree of report generation, but also enhances the readability and practicability of the report through visual means, and provides comprehensive support for intelligent stress testing of server hardware.
[0083] It should be appreciated that the method steps in the embodiments of the present application can be realized or implemented by computer hardware, a combination of hardware and software, or through computer instructions stored in a non-transitory computer readable memory. The method can use standard programming techniques. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with a computer system. However, if necessary, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. Furthermore, for this purpose, the program can run on a programmed special integrated circuit.
[0084] Further, the operations of the processes described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The processes described herein (or variations and / or combinations thereof) can be performed under the control of one or more computer systems configured with executable instructions (e.g., executable instructions for a computer program or applications, one or more computer programs or applications, etc.) and can be implemented as code (e.g., executable instructions, one or more computer programs or applications, etc.) executing collectively on one or more processors, by hardware or combinations thereof. The computer programs include a plurality of instructions executable by one or more processors.
[0085] Further, the methods can be implemented in any suitable type of computing platform operably connected to, including but not limited to, a personal computer, a mini-computer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Aspects of the present application can be implemented in machine readable code stored on a non-transitory storage medium or device, whether removable or integrated to the computing platform, such as a hard disk, an optical read and / or write storage medium, RAM, ROM, etc., such that it is readable by a programmable computer and, when the storage medium or device is read by the computer, is usable to configure and operate the computer to perform the processes described herein. Further, the machine readable code, or portions thereof, can be transmitted over a wired or wireless network. The present application described herein includes these and other different types of non-transitory computer readable storage media when such media include instructions or programs that implement the steps described above in conjunction with a microprocessor or other data processor. The present application can also include the computer itself when programmed in accordance with the methods and techniques described herein.
[0086] A computer program can be applied to input data to perform the functions described herein, thereby transforming the input data to generate output data that is stored to non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the application, the transformed data represents a physical and tangible object, including a particular visual depiction of a physical and tangible object produced on a display.
[0087] The above description is only preferred embodiments of the present application, the present application is not limited to the above-described embodiments, as long as the same means to achieve the technical effects of the present application, any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application, should be included in the scope of protection of the present application. The technical solutions and / or embodiments within the scope of protection of the present application can have various modifications and changes.
Claims
1. Server hardware intelligent stress testing method, characterized in that: include: Obtaining a configuration file of the server hardware, wherein the configuration file of the server hardware includes basic information of several devices to be managed; Executing preset task commands on the plurality of devices to be managed through the SSH protocol, and using multi-threading technology to execute the same task on the plurality of devices to be managed in parallel; Generate a stress test command, and apply a stress test load to the plurality of devices to be managed in parallel according to the stress test command; Collecting system information of the plurality of devices to be managed in parallel to obtain key operating indicators corresponding to each device to be managed; Encrypting the key operating indicators and storing them in a device-specific log file; Traversing the key operating indicators corresponding to the plurality of devices to be managed, determining whether there are any abnormal key operating indicators, and if there are any abnormal key operating indicators, restarting the device to be managed corresponding to the abnormal key operating indicator through an HTTP POST request; The hardware information of the plurality of devices to be managed is collected and stored in a device-specific directory, the hardware information of the plurality of devices to be managed is analyzed, and a visual server hardware intelligent stress test result is generated.
2. The server hardware intelligent stress testing method according to claim 1, characterized in that: The step of executing preset task commands on the plurality of devices to be managed through the SSH protocol and executing the same task on the plurality of devices to be managed in parallel using a multi-threading technology includes: Use the json.load() method to load data from the configuration file of the server hardware and extract the devices variable field as a device list. The device list contains result dictionaries of several devices to be managed, each of which represents information about a device to be managed. Initialize the thread list and the result dictionary, traverse the device list, create an independent child thread for each device to be managed, call the first function to generate the preset task command, establish an SSH connection with the target management device through paramiko.SSHClient, use multi-threading technology to execute the same task on multiple target management devices in parallel, and return the result dictionary corresponding to the target management device; Wait for all child threads to complete the task and end the task execution operation.
3. The server hardware intelligent stress testing method according to claim 1, characterized in that: Generating a stress test command, and applying a stress test load to the plurality of devices to be managed in parallel according to the stress test command, includes: Generate a stress test command for a specific hardware type of the device to be managed; Calling the second function returns a command string, and using the stress tool to perform stress testing to simulate a high-load scenario; The stress test command is called by a third function, and the stress test command is applied in parallel to all the devices to be managed, so as to achieve the parallel application of stress test loads to the devices to be managed.
4. The server hardware intelligent stress testing method according to claim 1, characterized in that: The parallel collection of system information of the plurality of devices to be managed to obtain key operating indicators corresponding to each device to be managed includes: Collect the system information of a single device to be managed through the SSH command, call the parallel collection task through the fourth function, and distribute the system information collection task to all devices to be managed; A command string is returned through the fifth function, and the key operating indicators of each device to be managed are obtained by combining the top command and the curl command, and the results are returned.
5. The server hardware intelligent stress testing method according to claim 2, characterized in that: The encryption of the key operating indicators and storage in a device-specific log file includes: Encrypting the key operating indicator data using a preset encryption algorithm and saving the data into a device-specific log file named after the device to be managed; The encryption processing results of all devices to be managed are traversed, and the encrypted log files are stored for all devices to be managed.
6. The server hardware intelligent stress testing method according to claim 5, characterized in that: The collecting of hardware information of the plurality of devices to be managed and storing the information in a device-specific directory includes: Calling the sixth function to return a command string, combining the lspci, lsblk, and nvme list commands to collect hardware information of the plurality of devices to be managed; Create a hardware log file for each device to be managed, and store the collected hardware information of each device to be managed in the corresponding hardware-specific log file.
7. The server hardware intelligent stress testing method according to claim 6, characterized in that: The step of analyzing the hardware information of the plurality of devices to be managed and generating a visual server hardware intelligent stress test result includes: Extracting encrypted key operating indicator data from the device-specific log file, and decrypting the encrypted key operating indicator data to obtain decrypted key operating indicator data; Extracting hardware information of the device to be managed from the hardware-specific log file; A visualized server hardware intelligent stress test result is generated based on the decrypted data of the key operating indicators and the hardware information of the device to be managed.
8. The server hardware intelligent stress testing method according to claim 7, characterized in that: Generating a visualized server hardware intelligent stress test result based on the decrypted data of the key operating indicators and the hardware information of the device to be managed includes: Create a report directory and define HTML templates to generate the structure and style of the test result report; Traversing the device list, calling a seventh function for each device to be managed in the device list, adding a chart path to the HTML template, and replacing a placeholder in the HTML template; Dynamically generate a test result table and chart for each device to be managed based on the HTML template according to the decrypted data of the key operating indicators and the hardware information of the device to be managed; Save the test result table and chart of each device to be managed as an HTML file and output the file path to obtain a visual server hardware intelligent stress test result.
9. A computer device comprising a memory and a processor, characterized in that: The method according to any one of claims 1 to 8 is implemented when the processor executes the computer program stored in the memory.
10. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Cloud server performance data collection method and device based on python
CN108763042A
Method and device for displaying performance of OpenPOWER server
CN109933476A
Implementation method and device of batch execution framework based on LINUX platform
CN111221510A
Server pressure test inspection method and system, electronic equipment and storage medium
CN116627747A
Computer service pressure test system and method
CN116627799A