Test method and device, electronic equipment and computer readable storage medium
Through the master control program, the test progress and status of multiple test nodes is solved, and the problem of low efficiency and poor reliability in the combination test of dual-node server and intelligent network card is realized, automation and collaborative testing are improved, and testing efficiency and reliability are improved.
Patent Information
- Application Number
- CN202510667800.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-26
AI Technical Summary
The combined testing solution of two-node server and smart network card is semi-automated, resulting in low testing efficiency and poor reliability, and the inability to achieve automation and collaborative testing.
Through the master control program, multiple test nodes are controlled to execute the test stages in turn, match the test progress and conditions, adjust the test operations of abnormal nodes, and synchronize the test status after the node returns to normal to generate target test results.
The collaborative automated testing of multiple test nodes is realized, which reduces human errors, improves testing efficiency and reliability, and shortens the test cycle.
Smart Images

Figure CN120540916A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of server testing, and more specifically to a testing method, device, electronic device, and computer-readable storage medium. Background Art
[0002] Server performance and stability are crucial for data processing and network transmission. The combination of dual-node servers and SmartNICs has gained widespread adoption due to their advantages in resource integration and network optimization. However, because the testing solution for this combination is semi-automatic, test efficiency and reliability are low. Summary of the Invention
[0003] In view of the above problems, the present application provides a testing method, device, electronic device and computer-readable storage medium.
[0004] According to the first aspect of the present application, a testing method is provided, comprising: in a process of controlling multiple test nodes to sequentially execute test tasks of multiple test phases according to a preset test sequence, matching the test progress of each of the multiple test nodes with the test conditions of the multiple test phases to obtain a matching result, wherein the multiple test nodes include: multiple computing nodes and connection nodes; in a case where the matching result indicates that the test progress of at least one test node is abnormal, adjusting the test operations of the multiple test nodes according to the test status of the abnormal test node; in a case where it is determined that the abnormal test node has restored to a normal test status, synchronizing the current test status of the multiple test nodes so that the multiple test nodes continue to execute the test tasks according to the current test status to obtain test results of the multiple test phases; and generating target test results for the multiple test nodes according to the test results corresponding to the multiple test phases.
[0005] The second aspect of the present application provides a testing device, including: a matching module, which is used to match the test progress of multiple test nodes with the test conditions of multiple test stages in the process of controlling multiple test nodes to execute test tasks of multiple test stages in sequence according to a preset test order to obtain matching results, wherein the multiple test nodes include: multiple computing nodes and connection nodes; an adjustment module, which is used to adjust the test operations of multiple test nodes according to the test status of the abnormal test nodes when the matching result indicates that the test progress of at least one test node is abnormal; a synchronization module, which is used to synchronize the current test status of multiple test nodes when it is determined that the abnormal test node has restored to a normal test status, so that the multiple test nodes continue to execute the test tasks according to the current test status and obtain test results of multiple test stages; a generation module, which is used to generate target test results for multiple nodes based on the test results corresponding to the multiple test stages.
[0006] The third aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0007] The fourth aspect of the present application further provides a computer-readable storage medium storing a computer program or instructions, which implements the steps of the above method when the computer program or instructions are executed by a processor.
[0008] The fifth aspect of the present application further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when the above computer program or instructions are executed by a processor.
[0009] According to the embodiments of the present application, the main control program determines the test progress of the test node, and automatically determines the test task of the next test phase according to the preset test sequence, thereby realizing automated testing. The test nodes can be tested comprehensively and efficiently through multiple test phases. The main control program determines the abnormal test node based on the matching result, and adjusts the test operations of multiple test nodes according to the test status of the abnormal test node. It can automatically trigger abnormal operations, handle the test process abnormally in time, and ensure the smooth progress of the test process. After the abnormal test node returns to the normal test state, the current test status of multiple test nodes is synchronized to ensure that the test progress of multiple test nodes is consistent. Through the collaboration of the main control program, test progress judgment, and exception handling, collaborative automated testing of multiple test nodes is achieved, reducing errors. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above contents and other objects, features and advantages of the present application will become more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings, in which:
[0011] Figure 1 A schematic diagram schematically shows the structure of a dual-node server with a smart network card according to an embodiment of the present application;
[0012] Figure 2 A diagram schematically illustrates an application scenario of a testing method, apparatus, electronic device, and computer-readable storage medium according to an embodiment of the present application;
[0013] Figure 3 The following schematically shows a flow chart of a testing method according to an embodiment of the present application;
[0014] Figure 4 A flowchart schematically illustrates a test task for executing multiple test phases according to an embodiment of the present application;
[0015] Figure 5Schematically shows a flow chart of executing a test task in multiple test phases according to another embodiment of the present application;
[0016] Figure 6 Schematically shows a flowchart of the firmware testing phase according to an embodiment of the present application;
[0017] Figure 7 Schematically shows a flow chart of a sharing mode test phase according to an embodiment of the present application;
[0018] Figure 8 Schematically shows a flow chart of a through mode test phase according to an embodiment of the present application;
[0019] Figure 9 The following schematically shows a flow chart of the health check phase according to an embodiment of the present application;
[0020] Figure 10 The following schematically shows a structural block diagram of a testing device according to an embodiment of the present application;
[0021] Figure 11 A block diagram of an electronic device suitable for implementing a testing method according to an embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0022] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.
[0023] The terms used herein are only for describing specific embodiments and are not intended to limit this application. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0024] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0025] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0026] A dual-node server is a server architecture that integrates two independent computing nodes. The computing nodes contain processors, memory, and storage resources, and share power supplies, cooling modules, and smart network cards. The two computing nodes can operate independently or work together. A smart network card is a network adapter with independent computing resources. A smart network card supports functions such as data acceleration and protocol offloading, and can independently run an operating system. Smart network cards include, for example, multi-host smart network cards. A multi-host smart network card allows a physical network card to connect to multiple computing or storage hosts simultaneously, allowing multiple hosts to share the network functions and bandwidth resources of the network card. For example, a multi-host smart network card can virtualize a single network card into multiple logical channels, each channel directly connected to a different host, thereby achieving physical isolation and independent network configuration.
[0027] Figure 1 The schematic diagram schematically shows the structure of a dual-node server with a smart network card according to an embodiment of the present application. Figure 1 The dual-node server with SmartNIC architecture shown includes:
[0028] First computing node 110, which is provided with a USB interface 111 corresponding to first computing node 110; second computing node 120, which is provided with a USB interface 121 corresponding to second computing node 120. First computing node 110 and second computing node 120 both have components such as a CPU, memory, and a hard disk, and each of first computing node 110 and second computing node 120 has at least one USB interface. Midplane 130, which is used for communication and power supply, and multiple components can be connected through midplane 130 to form a whole. Smart NIC 140, which has an independent processor, memory, and hard disk and can be independently installed in the system. Smart NIC 140 itself can also be considered a node and has a corresponding USB interface 141 and two network ports 142. In addition, the structure of the dual-node server with the smart network card also includes: a power supply 150, a fan 160, and an input / output board (IO board for short) 170, wherein the input / output board 170 has a baseboard management controller (BMC management port) 171.
[0029] Most of the testing methods for the dual-node server with smart network card structure are semi-automatic testing. For example, basic scripts are written to realize the automatic execution of some test steps. However, during the test process, manual execution of test scripts, restarting servers and judging test results are still required. In addition, semi-automatic testing will lead to deficiencies in the integration of test processes. For example, different test links cannot effectively coordinate with each other, and a large amount of manual intervention is still required to switch test modes and judge test progress, such as manually switching network card modes and restarting the entire machine, resulting in low test efficiency. In addition, semi-automatic testing requires manual execution of different switching operations and manual judgments between multiple test nodes, which increases manual workload, consumes more energy, and makes the test process more prone to errors, which in turn leads to low reliability of test results. The above problems make it impossible to achieve automated and collaborative testing of multiple test nodes.
[0030] In view of this, an embodiment of the present application provides a testing method, including: in the process of controlling multiple test nodes to execute test tasks of multiple test stages in sequence according to a preset test order, matching the test progress of each of the multiple test nodes with the test conditions of the multiple test stages to obtain a matching result, wherein the multiple test nodes include: multiple computing nodes and connection nodes; when the matching result indicates that the test progress of at least one test node is abnormal, adjusting the test operations of the multiple test nodes according to the test status of the abnormal test node; when it is determined that the abnormal test node has restored to a normal test status, synchronizing the current test status of the multiple test nodes so that the multiple test nodes continue to execute the test tasks according to the current test status to obtain test results of multiple test stages; generating target test results for multiple nodes based on the test results corresponding to the multiple test stages.
[0031] Figure 2 The following schematically illustrates an application scenario diagram of a test method, device, electronic device, and computer-readable storage medium according to an embodiment of the present application. Figure 2 As shown, the application scenarios according to this embodiment may include:
[0032] The master server 210 can select a server as the master server and configure the Dynamic Host Configuration Protocol (DHCP) service to assign IP addresses to devices within the local area network. The master server can have functions such as running the master control program and / or testing components and storing logs. A switch 220 and several network cables 230 are used to form a local area network between the master server 210, the first computing node 110, the second computing node 120, and the smart network card 140. The first computing node 110 and the second computing node 120 can be connected to the local area network via a USB-to-RJ45 adapter card, and the BMC management port is also connected to the local area network.
[0033] The first computing node 110 and its corresponding USB interface 111, the second computing node 120 and its corresponding USB interface 121, the midplane 130, the SmartNIC 140, one USB interface 141 and two network ports 142 corresponding to the SmartNIC 140, the power supply 150, the fan 160, and the I / O board 170. The I / O board 170 includes a baseboard management controller 171. The two network ports 142 corresponding to the SmartNIC 140 are connected via a loop line 143. Specifically, the two ends of a cable are connected to a network port on the SmartNIC.
[0034] Taking multiple computing nodes including a first computing node 110 and a second computing node 120, and a connection node including an intelligent network card 140 as an example, the test method of the embodiment of the present application can be executed by the main control program, and in the process of sequentially executing the test tasks of multiple test phases, the test progress of the test node can be matched with the test conditions to obtain a matching result; when the matching result indicates that the test progress of at least one test node is abnormal, the test operations of multiple test nodes are adjusted according to the test status of the abnormal test node; when it is determined that the abnormal test node has restored to a normal test status, the current test status of the multiple test nodes is synchronized; and target test results for the first computing node 110, the second computing node 120 and the intelligent network card 140 are generated according to the test results of multiple test phases. Figure 2 The number of network cables and computing nodes shown in the figure is merely illustrative. Any number of network cables and computing nodes may be used depending on the implementation requirements.
[0035] The following will be based on Figure 2 The scene described by Figures 3 to 9 The test method of the embodiment of the present application is described in detail. Figure 3 The flowchart of the testing method according to the embodiment of the present application is schematically shown. Figure 3 The testing method of this embodiment includes operations S310 to S340.
[0036] In operation S310, while controlling the plurality of test nodes to sequentially execute test tasks of the plurality of test phases according to a preset test sequence, the test progress of each of the plurality of test nodes is matched with the test conditions of the plurality of test phases to obtain a matching result. The plurality of test nodes include computing nodes and connection nodes. The connection nodes may include SmartNICs (which themselves may also be considered as test nodes).
[0037] The execution order and test tasks of multiple test phases can be predetermined. Test conditions can include prerequisites for executing test tasks, such as the test condition of the firmware test phase including that a computing node must be in a shutdown state.
[0038] The test method of the embodiment of the present application can be executed by a preset main control program. The main control program can be deployed on a main server or a test node, and control multiple test nodes to perform operations such as powering on and off, calling test scripts, etc. through an in-band or out-of-band management interface (such as BMC), so that multiple test nodes can perform test tasks. For example: the main control program can control multiple test nodes to perform test tasks of multiple test phases in sequence according to a preset test sequence, wherein, in the current test phase, the main control program can obtain the test progress of each of the multiple test nodes, and match the test progress with the test conditions of the current test phase, so as to determine whether the test progress of the test node can meet the test conditions of the current test phase based on the matching result.
[0039] In operation S320 , if the matching result indicates that the test progress of at least one test node is abnormal, the test operations of the plurality of test nodes are adjusted according to the test status of the abnormal test node.
[0040] A test progress anomaly may include the test progress not meeting the test conditions. For example, during the firmware test phase, the second computing node needs to be shut down. However, after the main control program matches the test progress of the second computing node with the test conditions of the firmware test phase, the matching result indicates that the second computing node is not in the shutdown state. In other words, the test progress of the second computing node is abnormal, and the abnormal test node is the second computing node.
[0041] Abnormal test progress may prevent the abnormal test node from completing its test tasks and may also affect the test progress of other normal test nodes, causing them to be unable to complete subsequent test tasks. Therefore, it is necessary to adjust the test operations of multiple test nodes based on the test status of the abnormal test node so that the multiple test nodes can continue to execute test tasks normally. Adjusting the test operations of multiple test nodes may include, for example, performing fault recovery operations on the abnormal test node (for example, polling the abnormal test node's test status) and pausing the test progress of other test nodes except the abnormal test node to synchronize the test progress of other test nodes with the abnormal test node, thereby achieving coordinated testing of multiple test nodes.
[0042] In operation S330, when it is determined that the abnormal test node recovers to a normal test state, the current test states of the plurality of test nodes are synchronized so that the plurality of test nodes continue to execute the test tasks according to the current test states and obtain test results of the plurality of test phases.
[0043] The main control program can monitor the test status of the abnormal test node in real time to determine whether the abnormal test node has recovered to a normal test status. Synchronizing the current test status of multiple test nodes may include: keeping the current test progress of the multiple test nodes consistent.
[0044] Synchronizing test status ensures that all nodes operate under the same test conditions and environment, avoiding data inconsistencies caused by state asynchrony. This improves the accuracy and reliability of test results and helps coordinate the progress of multiple test nodes, enabling them to test at a consistent pace. This also prevents situations where some test nodes have completed certain test steps while others have not, leading to biased test data aggregation and a lack of accurate reflection of the overall system performance.
[0045] In operation S340, target test results for multiple nodes are generated based on the test results corresponding to the multiple test phases. For example, after obtaining multiple test results corresponding to the multiple test phases, the multiple test results can be comprehensively evaluated to generate the target test results.
[0046] According to an embodiment of the present application, the main control program automatically determines the test task for the next test phase based on the test progress of the test node and the preset test sequence, replacing manual judgment of the test progress. Compared with manual responsibility for the entire test process, this achieves automated testing and reduces the error rate of the test process. Through multiple test phases, comprehensive and efficient testing of test nodes can be achieved. Compared with only testing multiple test nodes in a single phase, the status of multiple test nodes can be comprehensively and accurately evaluated. By matching the test progress of each of the multiple test nodes with the test conditions of the multiple test phases by the main control program, determining whether there are abnormal test nodes based on the matching results, and adjusting the test operations of the multiple test nodes based on the test status of the abnormal test nodes, abnormal operations can be automatically triggered, and the test process can be handled in a timely manner, ensuring the smooth progress of the test process. When the main control program determines that the abnormal test node has returned to the normal test state, it synchronizes the current test status of the multiple test nodes, ensuring that the test progress of the multiple test nodes is consistent, and further achieving coordinated control of the multiple test nodes. Thus, through the coordination of the main control program, test progress judgment, and exception handling, coordinated and automated testing of multiple test nodes is achieved, avoiding human errors. The test cycle of the automated collaborative test of the dual-node server with smart network card structure achieved by the above method can be shortened by more than half compared with the test time of traditional serial testing.
[0047] According to an embodiment of the present application, the testing method also includes: when executing the test task using a main control program running on multiple test nodes respectively: determining the test progress of any test node by locally reading the test log of any test node among the multiple test nodes; and remotely querying the test logs of other test nodes among the multiple test nodes to determine the test progress of other test nodes.
[0048] The main control program can run in multiple computing nodes and connection nodes respectively. For the main control program on any test node, the test log in the test node can be read locally, and the test logs of other test nodes can be queried remotely to determine the test progress based on the test logs. For example, when multiple test nodes include a first computing node, a second computing node, and a connection node (such as a smart network card node), the main control program running on the first computing node can read the test log of the first computing node locally, and remotely obtain the test logs of the second computing node and the smart network card node. Furthermore, the main control program on any test node can also remotely obtain the power on / off status of other test nodes in the multiple test nodes.
[0049] For example, automated testing software including a main control program and test components can be developed. The main control program can be run on multiple test nodes respectively. The test component contains scripts for implementing the test of each functional module and the related test tools used. The scripts of each functional module, for example, include modules for implementing test tasks in multiple stages. The automated testing software can also include stress testing tools for network cards, CPUs, memories, hard disks, etc. The test components can be stored on the main server. After the test nodes are powered on, multiple test nodes enter the operating system respectively, automatically copy the test components from the main server, and automatically run the main control program. The main control program can upload the test logs to the main server in real time, and can access the test logs of other test nodes.
[0050] At each start-up, the master control system first determines the test node of the current operating environment. It then queries and analyzes the historical test logs of the test node to determine the test progress since the last startup of the test node and determines the subsequent test steps. The master control system executes different test steps and calls different test modules based on the operating environment. It also determines the test progress of the test node or other test nodes by querying and parsing the logs.
[0051] Related testing methods are usually only suitable for testing scenarios suitable for a single or a small number of servers and cannot meet the testing needs for large-scale servers. By running the master control program on multiple test nodes, this distributed master control mode does not occupy the resources of the main server and is suitable for scenarios such as factories or data centers where a large number of servers need to be tested simultaneously. In addition, the master control program can remotely query the test logs and power on / off status of other test nodes to achieve collaborative testing of multiple test nodes in large-scale server scenarios. However, when using master control programs running on multiple test nodes to perform test tasks, the control logic of the test process is relatively complex, and the workload of preliminary development is relatively large.
[0052] According to an embodiment of the present application, the testing method also includes: when executing the test task using a main control program running on a server: sending a test progress acquisition instruction to multiple test nodes; and determining the test progress of each of the multiple test nodes in response to the test status information returned by the multiple test nodes.
[0053] The master control program can also run on a server different from the servers corresponding to the multiple computing nodes, such as a master server. To determine test progress, the master control program can send test progress acquisition instructions to the multiple test nodes, enabling the multiple test nodes to respond and return test status information. The master control program then determines the test progress based on the test status information of each of the multiple test nodes. The master control program can remotely call test scripts and copy files via the Secure Shell (SSH) or Secure Copy (SCP) protocols, enabling the test nodes to execute test tasks according to the test scripts and copied files. The master control program can also remotely control node power on and off via the BMC interface. In addition, the master control program running on the server can communicate with the test nodes via IP to implement functions such as file transfer, remote test execution, log storage, parallel testing, and remote power on and off. The master control program can also remotely control the power on, power off, and restart of multiple test nodes via the BMC interface. For example, the master control program can use the Intelligent Platform Management Interface (IPMI) interface to remotely control power on, power off, and restart via preset commands.
[0054] For example, if multiple test nodes include a first compute node, a second compute node, and a connection node (such as a SmartNIC node), when the master control program begins running, it can transfer test components stored on the master server to designated directories on the operating systems of the first compute node, the second compute node, and the connection node. For example, a pre-set command can be used to transfer files, where the pre-set name includes the test component and the IP address of the destination test node, thereby implementing the file transfer function.
[0055] The main control program can also remotely execute commands and scripts on the first computing node, the second computing node, and the connection node through the secure shell protocol. For example, it can use preset commands to decompress the test components and execute the test scripts in the target directory, thereby realizing the remote execution of the test function.
[0056] The master control program can also copy the test log from the test node to the main server for storage through the secure shell protocol. For example, a preset command can be used to copy the test node or smart network card file to the main server directory, thereby realizing the log storage function. In addition, the master control program can also issue test commands to the first computing node, the second computing node, and the connection node at the same time, so that the first computing node, the second computing node, and the connection node can run tests in parallel, thereby realizing the parallel testing function. The master control program can query the power on and off status based on network protocol communication, log analysis, and out-of-band management interfaces (such as BMC) to synchronize the test progress across nodes.
[0057] When the master control program runs on a server, the control logic is simple, making parallel testing easy. However, when testing a large number of servers, each server under test requires a master control process to run on the master server, which places high demands on the master server's performance and increases testing costs.
[0058] The running location of the main control program can be selected according to the actual test scenario to meet diverse test needs and adapt to different test scenarios. For example, in a large-scale test scenario, the main control program can be run on multiple test nodes separately, so as not to occupy the main server resources. In the scenario of testing a single or fewer servers, the main control program can be run on the server first.
[0059] According to an embodiment of the present application, adjusting the test operations of the plurality of test nodes includes: polling the test status of the abnormal test node until the abnormal test node recovers to a normal test status or a maximum polling number is reached.
[0060] For example, for an abnormal test node, in order to obtain the status of the abnormal test node in a timely manner and to be informed in time when the abnormal test node returns to normal so as to advance the test progress, the test status of the abnormal test node can be polled at a predetermined time interval. When it is determined that the abnormal test node has returned to normal test status, polling is stopped and multiple test nodes are controlled to continue executing the test task. When the pre-set maximum number of polling times is reached, it means that the abnormal test node has not yet returned to normal test status. At this time, the abnormal test node can be restarted or an alarm message can be displayed to prompt the staff to handle the abnormal test node.
[0061] For other test nodes except the abnormal test node, adjusting the test operations of the plurality of test nodes may include: pausing the test tasks of the other test nodes to ensure that the test progress of the other test nodes can be synchronized with the abnormal test node.
[0062] When the test progress of at least one test node is determined to be abnormal, the test operations of multiple test nodes are adjusted to ensure that the test progress of multiple test nodes is consistent. The test status of the abnormal test node can be obtained in a timely manner through the polling operation, which facilitates the advancement of the test progress.
[0063] Figure 4 The flowchart of executing the test task of multiple test phases according to the embodiment of the present application is schematically shown. Figure 4 As shown, the test process of this embodiment may include operations S410 to S440.
[0064] In operation S410, a test task of a firmware test phase is executed;
[0065] In operation S420, a test task of the shared mode test phase is executed;
[0066] In operation S430, a test task of a pass-through mode test phase is performed;
[0067] In operation S440 , the test task of the health check phase is performed.
[0068] Depend on Figure 4 It can be seen that the multiple test phases include the firmware test phase, the shared mode test phase, and the direct mode test phase. In the shared mode test phase, multiple computing nodes share a network port of the network card node; in the direct mode test phase, multiple computing nodes correspond to different network ports of the network card node.
[0069] The multiple test phases may also include a health check phase, which is used to test whether the system log reports errors, whether the BMC issues alarms, etc. Switching between different test phases usually requires a full machine restart to take effect.
[0070] Multiple compute nodes can be connected to connection nodes in a variety of ways. For example, if the connection node is a Smart NIC, the Smart NIC has two modes: direct and shared. In direct mode, two compute nodes are directly connected to a network port on the Smart NIC, each occupying a dedicated network port. In shared mode, the Smart NIC's CPU manages and uses both network ports. It coordinates the two ports, determining which port to send data from and how to allocate network bandwidth.
[0071] Related testing methods typically don't divide test phases, or only divide them into firmware refresh phases and other phases, because they can't perform targeted testing on test nodes based on their characteristics. However, by rationally dividing the test phases into multiple phases, including firmware testing phases, shared mode testing phases, and pass-through mode testing phases, it is possible to fully consider the different modes of the Smart NIC based on the characteristics of the dual-node server and Smart NIC combination, thereby conducting comprehensive and efficient testing of the test nodes.
[0072] Figure 5 The flowchart of executing a test task of multiple test phases according to another embodiment of the present application is schematically shown. Figure 5 As shown, the test process of this embodiment may include operations S510 to S590.
[0073] like Figure 5 As shown, multiple test phases correspond to multiple boot operations. For example, the first boot enters phase 1, which is the firmware test phase; the second boot enters phase 2, which is the shared mode test phase; the third boot enters phase 3, which is the pass-through mode test phase; the fourth boot enters phase 4, which is the health check phase. Figure 5 As shown:
[0074] In the firmware testing phase, operations S510 to S520 are performed sequentially after the device is powered on for the first time.
[0075] In operation S510 , the smart network card firmware is refreshed.
[0076] In operation S520 , the entire device is restarted to make the firmware effective.
[0077] In the shared mode test phase, operation S530 is performed after the second boot, including: performing the shared mode stress test of the smart network card and the computing node hardware stress test in parallel.
[0078] In the pass-through mode test phase, operations S540 to S560 are performed after the device is powered on for the third time.
[0079] In operation S540 , the smart network card is set to a pass-through mode.
[0080] In operation S550 , the entire device is restarted to enable the pass-through mode.
[0081] In operation S560 , a bidirectional traffic test is performed between the multiple computing nodes via the loop line, and the smart network cards perform hardware stress testing in parallel.
[0082] During the health check phase, after the fourth power-on, operations S570 to S590 are executed.
[0083] In operation S570 , the smart network card is reset to a sharing mode.
[0084] In operation S580 , the entire device is restarted to enable the sharing mode.
[0085] In operation S590 , a health check is performed.
[0086] During various testing phases, test nodes may need to be powered on and off. The master control program can control the power on and off status of test nodes through in-band or out-of-band management interfaces, such as the BMC. Through these interfaces, the master control program can remotely control the power on and off of multiple test nodes, eliminating frequent manual power on and off operations, reducing human intervention, and improving testing efficiency.
[0087] According to an embodiment of the present application, specifically, the test tasks in the firmware testing phase include: executing a firmware refresh operation; when it is determined that the firmware refresh operation is completed and the second computing node among the multiple computing nodes is in a shutdown state, controlling the first computing node among the multiple computing nodes to restart, wherein the connection node restarts synchronously with the first computing node.
[0088] Before executing the firmware refresh operation, the main control program can first control the second computing node to shut down. During the process of executing the firmware refresh on the first computing node, if the second computing node is in the power-on state, it may interfere with the refresh process. For example, the second computing node may be using certain resources of the firmware refresh object, resulting in conflicts or errors in the refresh process, and the operation of the second computing node may occupy system resources, affecting the refresh operation of the first computing node. Therefore, the second computing node can be controlled to be in the power-off state first to ensure that the subsequent firmware refresh operation can be executed smoothly. The firmware refresh operation is performed after the second computing node is shut down. The firmware refresh operation may include refreshing or upgrading the BMC, complex programmable logic device (CPLD) and the like of the smart network card.
[0089] After determining that the first computing node has completed the firmware refresh operation, it is possible to determine whether the second computing node is in a shutdown state. If the second computing node is not shut down, wait for a predetermined period of time before determining whether the second computing node is in a shutdown state. For example, wait for 10 seconds and then query again whether the second computing node is in a shutdown state. If it is determined that the second computing node is shut down, the first computing node among the multiple computing nodes can be controlled to restart so that the firmware update takes effect.
[0090] The connected node and the first compute node can be restarted synchronously through a coordinated power-on and power-off mechanism. This coordinated power-on and power-off mechanism involves coordinating the power on and off of the connected node with multiple compute nodes. Specifically, the connected node is powered on when any compute node is powered on, and is powered off when all compute nodes are powered off. For example, if the connected node is a Smart NIC, and the multiple compute nodes include a first compute node and a second compute node, the Smart NIC is powered on when either the first compute node or the second compute node is powered on, and is powered off when both the first compute node and the second compute node are powered off.
[0091] Take the distributed master control mode, the computing nodes including the first computing node and the second computing node, and the connection nodes including the smart network card as an example: Figure 6 The flowchart of the firmware testing phase according to the embodiment of the present application is schematically shown. Figure 6 The test process of this embodiment may include operations S610 to S670 , and the execution subject is a main control program running on multiple test nodes respectively.
[0092] In operation S610 , the main control program automatically starts running after the computer is powered on for the first time.
[0093] In operation S620 , the main control program first determines whether it is currently running on the first computing node, the second computing node, or the smart network card, thereby determining its own operating environment.
[0094] In operation S630, the main control program running on the second computing node first controls the second computing node to shut down. For example, the second computing node can be controlled to shut down through a BMC interface, which can be achieved by issuing a preset shutdown command.
[0095] In operation S640, the main control program running on the first computing node executes a firmware refresh operation. Executing the firmware refresh operation includes refreshing the firmware of the SmartNIC. For example, the main control program running on the first computing node may invoke a module for performing a firmware refresh operation to refresh or upgrade the firmware of the SmartNIC, such as the BMC and CPLD.
[0096] In operation S650, after the smart network card firmware is refreshed, the main control program running on the first computing node determines whether the second computing node is shut down. For example, the main control program running on the first computing node can determine whether the second computing node is in a shutdown state through the BMC interface. For example, the power on / off status can be queried through a BMC out-of-band command, and the BMC out-of-band command includes the BMC IP address, user name, and password of the second computing node. If it is determined that the second computing node is not shut down, the main control program running on the first computing node can query whether the second computing node is shut down every predetermined period of time, for example, once every 10 seconds.
[0097] In operation S660, the main control program running on the first computing node may restart the first computing node to make the firmware effective if it is determined that the second computing node is shut down.
[0098] In operation S670, after restarting the first computing node, the second computing node can be controlled to start up in order to perform the test tasks of the subsequent test phase. Restarting the first computing node first and then controlling the second computing node to start up can ensure that the test process is executed in an orderly manner.
[0099] Depend on Figure 6 It can be seen that the main control program running on the smart network card does not perform any operation during the firmware testing phase, and the smart network card can be restarted along with the first computing node based on the coordinated power-on and power-off mechanism.
[0100] The main control program controls the execution of the firmware testing phase and the operations between different test nodes, avoiding manual switching between multiple test nodes, performing different operations, and making manual judgments. This avoids the situation where operators are more likely to make mistakes due to manual energy consumption, and improves the efficiency and accuracy of the firmware testing phase.
[0101] According to an embodiment of the present application, the test tasks of the shared mode test phase include: sending shared mode test instructions to enable multiple computing nodes and connection nodes to execute the test tasks of the shared mode test phase in parallel; when it is determined that the test tasks of the shared mode test phase are completed, adjusting the connection nodes from shared mode to pass-through mode.
[0102] Take the distributed master control mode, multiple computing nodes including the first computing node, the second computing node, and the connection node including the smart network card as an example: Figure 7 The flowchart of the sharing mode test phase according to the embodiment of the present application is schematically shown. Figure 7 As shown, the test process of this embodiment may include operations S710 to S790, and the execution subject is the main control program.
[0103] In operation S710, the main control program automatically starts running after the second power-on.
[0104] In operation S720 , the main control program first determines whether it is currently running on the first computing node, the second computing node, or the smart network card, thereby determining its own operating environment.
[0105] In operation S730, the main control program running on multiple test nodes respectively controls multiple test nodes in parallel to execute test tasks of the shared mode test phase. For example, the main control program can send shared mode test instructions to the running test nodes so that multiple test nodes can execute test tasks of the shared mode test phase in parallel, thereby improving test efficiency.
[0106] The test tasks of the connection node executing the shared mode test phase may include: the main control program running on the smart network card executes operation S731 on the smart network card, that is, the shared mode test. The shared mode test may include a network port connectivity test and a stress test. For example, the two network ports of the smart network card may be controlled to connect to each other through an external loop line. Specifically, one network port can be a client and the other network port can be a server. The two network ports send and receive data packets to each other through the loop line outside the network card, thereby performing a network port connectivity test and a stress test on the smart network card in shared mode. The test tasks of the first computing node and the second computing node executing the shared mode test phase may include: respectively executing operation S732 on the first computing node, that is, the first hardware stress test, and executing operation S733 on the second computing node, that is, the second stress test.
[0107] In operation S740 , after it is determined that the second computing node has completed the hardware stress test, the main control program running on the second computing node may control the second computing node to shut down.
[0108] In operation S750 , after determining that the smart network card has completed the shared mode test, the main control program running on the smart network card may set the smart network card from the shared mode to the pass-through mode to execute the next stage of the test task.
[0109] In operation S760, after determining that the first computing node has completed the hardware stress test, the main control program running on the first computing node can determine whether the smart network card is in pass-through mode. If the smart network card is not in pass-through mode, subsequent operations must wait until the smart network card is set to pass-through mode, thereby ensuring that the test process can be executed in order. The status of the smart network card can be polled to determine whether the smart network card is in pass-through mode.
[0110] In operation S770, if the SmartNIC is determined to be in pass-through mode, the main control program running on the first computing node determines whether the second computing node is shut down. For example, the main control program can determine whether the second computing node is shut down through the BMC interface. If the second computing node is determined to be not shut down, the main control program can query whether the second computing phase is shut down every predetermined period of time.
[0111] In operation S780, if the master program running on the first computing node determines that the second computing node is shut down, it can restart the first computing node to perform the next phase of testing. Based on the coordinated power-on and power-off mechanism, the smart network card restarts after the first computing node is restarted.
[0112] In operation S790, after the first computing node is restarted, the second computing node may be controlled to start up so as to execute the test tasks of the subsequent test phase.
[0113] By controlling multiple test nodes to simultaneously start and execute test tasks in the shared mode test phase in parallel, the idle waiting time of hardware resources is reduced, hardware resources can be fully utilized, and the test cycle is shortened. Parallel testing is closer to actual usage scenarios, which can keep the entire machine under greater pressure, making it easier to detect potential problems and improving test reliability.
[0114] According to an embodiment of the present application, the test tasks of the direct mode test phase include: sending a direct mode test instruction so that multiple computing nodes and connection nodes simultaneously execute the test tasks of the direct mode test phase; when it is determined that the connection node has completed the test tasks of the direct mode test phase, querying whether multiple computing nodes have completed the test tasks of the direct mode test phase according to a preset time interval; when it is determined that multiple computing nodes have completed the test tasks of the direct mode test phase, adjusting the connection node from the direct mode to the shared mode.
[0115] Take the distributed master control mode, multiple computing nodes including the first computing node, the second computing node, and the connection node including the smart network card as an example: Figure 8 Schematically shows a flow chart of the through mode test phase according to an embodiment of the present application. Figure 8 As shown, the test process of this embodiment may include operations S810 to S890, and the execution subject is the main control program.
[0116] In operation S810, the main control program automatically starts running after the third power-on.
[0117] In operation S820 , the main control program first determines whether it is currently running on the first computing node, the second computing node, or the smart network card, thereby determining its own operating environment.
[0118] In operation S830, the main control program running on multiple test nodes respectively controls multiple test nodes to simultaneously execute test tasks of the direct mode test phase. For example, the main control program can send direct mode test instructions to multiple test nodes so that multiple test nodes can execute test tasks of the direct mode test phase in parallel, thereby improving test efficiency.
[0119] For example, the test tasks for executing the passthrough mode test phase on the SmartNIC may include: performing a stress test on the SmartNIC's CPU, memory, and hard disk. The test tasks for executing the passthrough mode test phase on the first compute node may include: configuring a first IP address for the network port to assign an IP address to the network card, and then enabling the server of the stress testing tool so that the stress testing tool runs as a server on the first compute node.
[0120] The test task of performing the pass-through mode test phase on the second computing node may include: configuring a second IP address for the network port, thereby allocating an IP address to the network card, and ensuring that the network port IP addresses of the first computing node and the second computing node are in the same network segment through the first IP address setting and the second IP address setting;
[0121] The stress testing tool is then controlled to run as a client on the second computing node. In addition, the server IP is specified as the IP of the first computing node. In this way, in the smart network card pass-through mode, the first computing node and the second computing node control the smart network card to bidirectionally connect the external loop line and send and receive data packets in both directions, thereby achieving the purpose of network port connectivity and stress testing of the test node.
[0122] In operation S840, the main control program running on the smart network card determines whether the computing node has completed the test. When determining the first computing node and / or the second computing node, it can wait and determine whether the computing node has completed the test at a predetermined time interval, for example, querying whether the computing node has completed the test every 10 seconds.
[0123] In operation S850, when it is determined that the computing node has completed the test, the main control program running on the smart network card sets the smart network card from the direct mode to the shared mode, so that the smart network card can perform the next stage of testing tasks in the shared mode.
[0124] In operation S860, the main control program running on the first computing node determines whether the smart network card has been set to shared mode. If it is determined that the smart network card has not been set to shared mode, subsequent operations are performed after the smart network card setting is completed. During the waiting period, whether the smart network card setting is completed can be queried at a predetermined time interval, for example, once every 10 seconds.
[0125] In operation S870, after determining that the smart network card has been set to shared mode, the main control program running on the first computing node determines whether the second computing node has been shut down. If it is determined that the second computing node has not been shut down, subsequent operations are performed after the second computing node is shut down. During the waiting period, the smart network card can be queried at a predetermined time interval to determine whether the second computing node has been shut down.
[0126] In operation S880, when it is determined that the second computing node has been shut down, the main control program running on the first computing node restarts the first computing node.
[0127] In operation S890, after restarting the first computing node, the second computing node is controlled to start up, so that the first computing node and the second computing node can perform the next phase of the test task. Based on the coordinated power-on and power-off mechanism, the smart network card is restarted after the first computing node is restarted.
[0128] By performing parallel testing on multiple test nodes in stages such as direct mode testing, multiple test nodes do not interfere with each other, which can shorten the overall test cycle and improve test efficiency. It avoids the situation where only one node can be operated at a time during manual operation and the test process is executed serially. When one node is being tested, other nodes are idle most of the time.
[0129] The multiple test phases may also include a health check phase. The test tasks in the health check phase may include, for example, checking whether the device is in place, whether there are errors in the system log, and whether there are BMC alarms. Multiple test nodes may be controlled to execute the test tasks in the health check phase in parallel.
[0130] Take the distributed master control mode, multiple computing nodes including the first computing node, the second computing node, and the connection node including the smart network card as an example: Figure 9 The flowchart of the health check phase according to the embodiment of the present application is schematically shown. Figure 9 As shown, the test process of this embodiment may include operations S910 to S930, and the execution subject is the main control program.
[0131] In operation S910, the main control program automatically starts running after the fourth power-on.
[0132] In operation S920 , it is determined whether the system is running on a computing node or a smart network card.
[0133] In operation S930 , the main control programs respectively running on the multiple test nodes control the multiple test nodes in parallel to execute the test tasks of the health check phase.
[0134] By executing health check tests on multiple test nodes in parallel, the test cycle can be shortened and test efficiency improved. Health checks can comprehensively detect device presence, log errors, and BMC alarms, improving fault detection rates and test result reliability.
[0135] In multiple test phases, the main control program controls the switching operations between multiple test nodes. Compared with the related methods in which manual operations frequently switch between multiple nodes, perform different operations, and make judgments, this reduces the manual workload and the probability of human error, and improves test reliability.
[0136] Based on the above test method, this application also provides a test device. Figure 10 The device is described in detail. Figure 10 The following schematically shows a structural block diagram of a test device according to an embodiment of the present application. Figure 10 As shown, the testing device 1000 of this embodiment includes a matching module 1010 , an adjustment module 1020 , a synchronization module 1030 and a generation module 1040 .
[0137] Matching module 1010 is configured to match the test progress of each of the multiple test nodes with the test conditions of the multiple test phases during the process of controlling the multiple test nodes to sequentially execute test tasks of the multiple test phases according to a preset test sequence, thereby obtaining a matching result. The multiple test nodes include multiple computing nodes and connection nodes. In one embodiment, matching module 1010 can be configured to perform operation S310 described above, which is not further described here.
[0138] The adjustment module 1020 is configured to adjust the test operations of the plurality of test nodes according to the test status of the abnormal test node when the matching result indicates that the test progress of at least one test node is abnormal. In one embodiment, the adjustment module 1020 may be configured to perform the operation S320 described above, which will not be described in detail here.
[0139] Synchronization module 1030 is configured to synchronize the current test states of multiple test nodes when it is determined that the abnormal test node has returned to a normal test state, so that the multiple test nodes can continue to execute the test task based on the current test state and obtain test results for multiple test phases. In one embodiment, synchronization module 1030 can be used to perform operation S330 described above and will not be further described.
[0140] The generating module 1040 is used to generate target test results for multiple test nodes according to the test results corresponding to the multiple test phases. In one embodiment, the synchronizing module 1030 can be used to perform the operation S340 described above, which will not be described in detail here.
[0141] The testing device includes a first determination module, which is used to determine the test progress of any test node by locally reading the test log of any test node among the multiple test nodes when using the main control program running on multiple test nodes respectively to perform the test task; the second determination module is used to remotely query the test logs of other test nodes among the multiple test nodes to determine the test progress of other test nodes.
[0142] The testing device includes a sending module for sending a test progress acquisition instruction to multiple test nodes when executing a test task using a main control program running on a server; and a third determining module for determining the test progress of each of the multiple test nodes in response to the test status information returned by the multiple test nodes.
[0143] The adjustment module includes a polling submodule for polling the test status of the abnormal test node until the abnormal test node recovers to a normal test status or reaches a maximum polling number.
[0144] The testing device also includes a firmware refresh module for executing test tasks during the firmware testing phase. The firmware refresh module includes an execution submodule for executing the firmware refresh operation; and a control submodule for controlling a first computing node among the multiple computing nodes to restart when the firmware refresh operation is completed and a second computing node among the multiple computing nodes is in a shutdown state, wherein the connection node restarts synchronously with the first computing node.
[0145] The testing apparatus also includes a shared mode test module configured to execute test tasks during the shared mode test phase. The shared mode test module includes a first sending submodule configured to send a shared mode test instruction to cause a first computing node, a second computing node, and a connection node among the plurality of computing nodes to execute the test tasks during the shared mode test phase in parallel; and a first adjustment submodule configured to adjust the connection node from a shared mode to a pass-through mode upon determining that the test tasks during the shared mode test phase have been completed.
[0146] The testing device also includes a direct mode test module for executing test tasks in the direct mode test phase. The direct mode test module includes: a second sending submodule for sending a direct mode test instruction so that multiple computing nodes and the connection node simultaneously execute the test tasks in the direct mode test phase; a query submodule for, upon determining that the network card node has completed the test tasks in the direct mode test phase, querying, at preset time intervals, whether the multiple computing nodes have completed the test tasks in the direct mode test phase; and a second adjustment submodule for, upon determining that the multiple computing nodes have completed the test tasks in the direct mode test phase, adjusting the operating mode of the connection node from the direct mode to the shared mode.
[0147] According to embodiments of the present application, any multiple modules among the matching module 1010, adjustment module 1020, synchronization module 1030, and generation module 1040 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. At least one of the matching module 1010, adjustment module 1020, synchronization module 1030, and generation module 1040 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the matching module 1010 , the adjustment module 1020 , the synchronization module 1030 and the generation module 1040 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.
[0148] Figure 11 Schematically shows a block diagram of an electronic device suitable for implementing the test method according to an embodiment of the present application. Figure 11As shown, electronic device 1100 according to an embodiment of the present application includes a processor 1101, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1102 or programs loaded from storage 1108 into a random access memory (RAM) 1103. Processor 1101 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or related chipsets, and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)). Processor 1101 may also include onboard memory for cache purposes. Processor 1101 may include a single processing unit or multiple processing units for performing the various actions of the method flow according to the embodiment of the present application. RAM 1103 stores various programs and data required for the operation of electronic device 1100. Processor 1101, ROM 1102, and RAM 1103 are interconnected via a bus 1104. Processor 1101 performs the various actions of the method flow according to the embodiment of the present application by executing the programs in ROM 1102 and / or RAM 1103. It should be noted that the program may also be stored in one or more memories other than the ROM 1102 and the RAM 1103. The processor 1101 may also execute the various operations of the method flow according to the embodiment of the present application by executing the program stored in the one or more memories.
[0149] According to an embodiment of the present application, electronic device 1100 may further include an input / output (I / O) interface 1105, which is also connected to bus 1104. Electronic device 1100 may also include one or more of the following components connected to I / O interface 1105: an input section 1106 including a keyboard, mouse, etc.; an output section 1107 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 1108 including a hard disk; and a network interface card 1109 including a LAN card, modem, etc. 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to I / O interface 1105 as needed. Removable media 1111, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 1110 as needed, so that computer programs read from the removable media can be installed into storage section 1108 as needed.
[0150] The present application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present application. The computer-readable storage medium may be a non-volatile computer-readable storage medium, such as, but not limited to, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of the present application, the computer-readable storage medium may include ROM 1102 and / or RAM 1103 described above, and / or one or more memories other than ROM 1102 and RAM 1103.
[0151] Embodiments of the present application also include a computer program product comprising a computer program containing program code for executing the method illustrated in the flowchart. When the computer program product is executed in a computer system, the program code is used to cause the computer system to implement the method provided in the embodiments of the present application. When the computer program is executed by processor 1101, the aforementioned functions defined in the system / device of the embodiments of the present application are performed. According to embodiments of the present application, the systems, devices, modules, units, etc. described above may be implemented using computer program modules. In one embodiment, the computer program may be implemented on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal over a network medium and downloaded and installed via 1109 and / or installed from removable media 1111. The program code contained in the computer program may be transmitted using any suitable network medium, including but not limited to wireless, wired, etc., or any suitable combination thereof. In such an embodiment, the computer program may be downloaded and installed from a network via 1109 and / or installed from removable media 1111. When the computer program is executed by the processor 1101, the above functions defined in the system of the embodiment of the present application are executed. According to the embodiment of the present application, the system, device, apparatus, module, unit, etc. described above can be implemented by a computer program module.
[0152] According to an embodiment of the present application, the program code for executing the computer program provided by the embodiment of the present application can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0153] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should be noted that the functions marked in the boxes in some alternative implementations can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart and the combination of the boxes in the block diagram or flow chart can be implemented with a special hardware-based system that performs the specified function or operation, or can be implemented with a combination of special hardware and computer instructions.
[0154] Those skilled in the art will appreciate that the features described in the various embodiments of the present application may be combined and / or coupled in a variety of ways, even if such combinations or couplings are not explicitly described in the present application. In particular, without departing from the spirit and teachings of the present application, the features described in the various embodiments of the present application may be combined and / or coupled in a variety of ways. All of these combinations and / or couplings fall within the scope of the present application. The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although the various embodiments have been described above separately, this does not mean that the measures in the various embodiments cannot be advantageously used in combination. Without departing from the scope of the present application, those skilled in the art may make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present application.
Claims
1. A testing method, characterized in that: The method comprises: In a process of controlling a plurality of test nodes to sequentially execute test tasks of a plurality of test phases according to a preset test sequence, matching the test progress of each of the plurality of test nodes with the test conditions of the plurality of test phases to obtain a matching result, wherein the plurality of test nodes include: computing nodes and connection nodes; If the matching result indicates that the test progress of at least one test node is abnormal, adjusting the test operations of the plurality of test nodes according to the test status of the abnormal test node; When it is determined that the abnormal test node has recovered to a normal test state, synchronizing the current test states of the multiple test nodes so that the multiple test nodes continue to execute the test task according to the current test states and obtain test results of multiple test phases; Target test results for the plurality of test nodes are generated according to the test results corresponding to the plurality of test phases.
2. The method according to claim 1, characterized in that The method further includes, when executing the test task using main control programs respectively running on the plurality of test nodes: Determining the test progress of any test node by locally reading the test log of any test node among the multiple test nodes; Remotely query the test logs of other test nodes among the multiple test nodes to determine the test progress of the other test nodes.
3. The method according to claim 1, characterized in that The method further includes, when executing the test task using a main control program running on a server: Sending a test progress acquisition instruction to the multiple test nodes; In response to the test status information returned by the multiple test nodes, the test progress of each of the multiple test nodes is determined.
4. The method according to claim 1, wherein The test operation of adjusting the plurality of test nodes includes: The test status of the abnormal test node is polled until the abnormal test node returns to a normal test status or a maximum polling number is reached.
5. The method according to claim 1, wherein The computing nodes include multiple ones; the multiple test phases include: a firmware test phase, a shared mode test phase, and a pass-through mode test phase, wherein, in the shared mode test phase, multiple computing nodes share one network port of the network card node; in the pass-through mode test phase, the multiple computing nodes respectively correspond to different network ports of the network card node.
6. The method according to claim 5, characterized in that The test tasks of the firmware testing phase include: Execute the firmware flash operation; When it is determined that the firmware refresh operation is completed and the second computing node among the multiple computing nodes is in a shutdown state, the first computing node among the multiple computing nodes is controlled to restart, wherein the connecting node is restarted synchronously with the first computing node.
7. The method according to claim 5, characterized in that The test tasks of the shared mode test phase include: Sending a shared mode test instruction to enable the plurality of computing nodes and the connection node to execute test tasks of the shared mode test phase in parallel; When it is determined that the test task of the shared mode test phase is completed, the connection node is adjusted from the shared mode to the direct mode.
8. The method according to claim 5, characterized in that The test tasks of the direct mode test phase include: Sending a pass-through mode test instruction to enable the plurality of computing nodes and the connection node to simultaneously perform the test tasks of the pass-through mode test phase; In the case of determining that the connection node completes the test task of the direct mode test phase, querying the plurality of computing nodes at preset time intervals whether they complete the test tasks of the direct mode test phase; When it is determined that the plurality of computing nodes complete the test tasks of the pass-through mode test phase, the operation mode of the connection node is adjusted from the pass-through mode to the sharing mode.
9. A testing device, characterized in that: The device comprises: a matching module, configured to match the test progress of each of the plurality of test nodes with the test conditions of the plurality of test phases in a process of controlling the plurality of test nodes to sequentially execute the test tasks of the plurality of test phases according to a preset test sequence, and obtain a matching result, wherein the plurality of test nodes include: computing nodes and connection nodes; an adjusting module, configured to adjust the test operations of the plurality of test nodes according to the test status of the abnormal test node when the matching result indicates that the test progress of at least one test node is abnormal; a synchronization module, configured to synchronize the current test states of the plurality of test nodes when it is determined that the abnormal test node has recovered to a normal test state, so that the plurality of test nodes continue to execute the test task according to the current test state and obtain test results of a plurality of test phases; A generating module is configured to generate target test results for the plurality of test nodes according to the test results corresponding to the plurality of test phases.
10. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.