Network fault repairing method, device and equipment and storage medium

By obtaining network information of the server operating system and management system, calculating the overall accessibility index, and accurately locating and repairing the target server with network failure, the problem of test failure caused by network failure in server automated batch testing is solved, and test reliability is improved.

CN119484250BActive Publication Date: 2025-10-21INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411620584.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-13
Publication Date
2025-10-21
Estimated Expiration
2044-11-13

AI Technical Summary

Technical Problem

During the automated batch testing of servers, the complexity of the test scenarios and the mutual interference between test cases may cause network failures, affecting the normal progress of the test and reducing the reliability of the automated test.

Method used

By obtaining network information of the server operating system and management system, the overall accessibility index is calculated, and the target server of the network failure is accurately located and repaired, including the use of smart devices for hardware interface docking, system reinstallation, and backup server replacement.

Benefits of technology

Improved the reliability of server automated batch testing to ensure the normal execution and successful completion of test tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119484250B_ABST
    Figure CN119484250B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of servers, and discloses a network fault repairing method, device, equipment and storage medium, which are applied to a test system for testing at least one server. The method comprises: acquiring first network information between each server operating system and second network information between each server management system; determining an overall accessibility index of one or more servers currently tested based on the first network information and the second network information; in the case that the overall accessibility index meets a preset condition, determining a target server with a network fault according to a test task performed for each server, and repairing the network fault of the target server. The reliability of server automated batch testing can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of server technology, and in particular to a network fault repair method, apparatus, device, and storage medium. Background Art

[0002] Before servers leave the factory, various functions are typically tested to ensure product quality. To reduce the workload, some technologies use test systems to perform automated batch testing of servers. However, when executing automated testing tasks, the complexity of test scenarios and interference between test cases can lead to network failures on servers, hindering the smooth progress of automated testing. The reliability of automated batch server testing still needs to be improved. Summary of the Invention

[0003] In view of this, the present disclosure provides a network fault repair method, a network fault repair device, an electronic device, and a computer-readable storage medium, which can improve the reliability of automated batch testing of servers.

[0004] In a first aspect, the present disclosure provides a network fault repair method, which is applied to a test system, wherein the test system is used to test at least one server; the method comprises:

[0005] Acquire first network information between each server operating system and second network information between each server management system;

[0006] determining an overall accessibility index of one or more servers currently being tested based on the first network information and the second network information;

[0007] When the overall accessibility index meets a preset condition, a target server having a network failure is determined according to the test tasks performed on each of the servers, and the network failure of the target server is repaired.

[0008] In a second aspect, the present disclosure provides a network fault repair device, which is applied to a test system for testing at least one server; the device includes:

[0009] A network information acquisition module, configured to acquire first network information between each server operating system and second network information between each server management system;

[0010] an accessibility index determining module, configured to determine an overall accessibility index of one or more servers currently being tested based on the first network information and the second network information;

[0011] The fault repair module is used to determine the target server with network failure according to the test tasks performed on each of the servers when the overall accessibility index meets the preset conditions, and repair the network failure of the target server.

[0012] In a third aspect, the present disclosure provides an electronic device comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, computer instructions are stored in the memory, and the processor executes the above method by executing the computer instructions.

[0013] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the above method.

[0014] In some embodiments of the present disclosure, the overall accessibility index determined based on the above-mentioned first network information and second network information can reflect the overall probability of a network failure occurring in one or more servers currently being tested. Furthermore, when the overall accessibility index meets the preset conditions, the target server with a network failure can be accurately located and the network failure of the target server can be repaired based on the test tasks performed on each server, thereby preventing the problem of server test failure caused by network failure, thereby greatly improving the reliability of automated batch testing of servers. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the related technologies, the following briefly introduces the drawings required for use in the specific embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0016] Figure 1 This is a schematic diagram of a scenario of automated batch testing of servers provided by some embodiments of the present disclosure;

[0017] Figure 2 is a flowchart of a network fault repair method provided by some embodiments of the present disclosure;

[0018] Figure 3 This is a module diagram of a network fault repair device provided by an embodiment of the present disclosure;

[0019] Figure 4 It is a structural diagram of an electronic device provided by some embodiments of the present disclosure. DETAILED DESCRIPTION

[0020] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present disclosure.

[0021] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0022] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "some embodiments" or "the embodiment" should be understood as "at least some embodiments". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below.

[0023] Herein, unless explicitly stated otherwise, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.

[0024] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0025] It is understandable that before using the technical solutions disclosed in the various embodiments of the present disclosure, the type, scope of use, usage scenarios, etc. of the information involved in the present disclosure should be informed to relevant users and authorization should be obtained from relevant users in an appropriate manner in accordance with relevant laws and regulations. The relevant users may include any type of right holders, such as individuals, enterprises, and groups.

[0026] For example, in response to receiving an active request from a user, a prompt message is sent to the relevant user to clearly prompt the relevant user that the operation requested to be performed will require obtaining and using the information of the relevant user, so that the relevant user can independently choose whether to provide information to the software or hardware such as the electronic device, application, server or storage medium that executes the operation of the technical solution of the present disclosure based on the prompt message.

[0027] As an optional but non-limiting implementation, in response to receiving an active request from a relevant user, a prompt message may be sent to the relevant user in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide information to the electronic device.

[0028] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0029] See also Figure 1 , which is a schematic diagram of a scenario for automated batch testing of servers provided in some embodiments of the present disclosure. Figure 1 In the example, servers A, B, and C are the servers to be tested. The test system executes test cases and scripts on each server to complete the server test. Server testing can be divided into two aspects: hardware testing and software testing.

[0030] Hardware testing may include, but is not limited to, central processing unit (CPU) testing, memory testing, storage testing, network testing, power supply testing, heat dissipation testing, motherboard testing, and hardware compatibility testing. For example, testing the effectiveness of server motherboard interfaces. Another example is testing the maximum read and write speeds and capacity supported by storage devices.

[0031] Software testing can include, but is not limited to, functional testing, performance testing, security and reliability testing, and software compatibility testing. For example, testing the server's Basic Input / Output System (BIOS) functionality. Another example is testing the server's power-off recovery function.

[0032] In some embodiments, servers A, B, and C may be servers of the same type. The so-called same type means that the software and hardware of servers A, B, and C are the same. In the test system, different test case scripts can be run for servers A, B, and C to execute different test tasks of the same type of servers in parallel. For example, run the test case script a1 to test the BIOS function and operating system function in server A; run the test case script b1 to test the motherboard heat dissipation and interface validity in server B. And so on. Finally, the test results of servers A, B, and C are summarized to obtain the test results of the corresponding type of servers. In these embodiments, different functions of servers of the same type can be tested in parallel, and the test efficiency is high.

[0033] In some embodiments, servers A, B, and C may be different types of servers. The so-called different types refer to differences in the software or hardware of servers A, B, and C. In the test system, each type of server may have one or more corresponding test case scripts. By running the test case scripts corresponding to each type of server, the test task for the corresponding type of server can be executed, and the test results for the corresponding type of server can be obtained. For example, assuming that server A is an M1 type server, corresponding to test case scripts a21 and a22; server B is an M2 type server, corresponding to test case scripts b21, b22, and b23. Then, after running test scripts a21 and a22 to execute the test task on server A, the test results for the M1 type server can be obtained; after running test scripts b21, b22, and b23 to execute the test task on server B, the test results for the M2 type server can be obtained. In these embodiments, the test system can run the test case scripts for each type of server in parallel, thereby performing parallel testing on each type of server, which has higher testing efficiency.

[0034] Based on the above description, this solution of running test case scripts through the test system to test multiple servers in parallel is the automated batch testing of servers.

[0035] It should be noted that, first of all, Figure 1 This is just an example of 3 servers, which does not constitute a limitation to the present disclosure. In the actual server automation batch testing scenario, the number of servers tested in parallel can be set according to actual needs, such as 5 servers or 10 servers. Secondly, the server type of batch testing, the correspondence between the server and the test case script, can also be set according to actual needs. The two test scenarios listed above do not constitute a limitation to the present disclosure. For example, among the multiple servers tested in parallel, some servers can be of the same type, and some servers can be of different types.

[0036] Continue reading Figure 1. In some embodiments of the present disclosure, intelligent devices (such as intelligent robots) may also be included in the area where the server is located. The test system can be communicatively connected with the intelligent device. During the batch testing of the server, for testing needs, the test system can also issue control instructions to the intelligent device to control the intelligent device to perform specified operations on one or more of the servers. For example, when the baseboard management control system (Baseboard Management Controller, BMC) of server A fails, resulting in the inability to remotely power on and off server A through the baseboard management control system, the test system can control the intelligent device to move to the location of server A and perform power-on or power-off operations on server A.

[0037] based on Figure 1 For the server automated batch test scenario shown, although it can greatly improve the server test efficiency and reduce the workload of server testing, due to the complexity of the test scenario and the mutual interference between test cases, it may cause network failures in the server, making it impossible for the test system to access the server's operating system or baseboard management and control system, thereby affecting the normal progress of the automated test. For example, when executing the test case script a21, the baseboard management and control system of server A is shut down for testing purposes, and after the test is completed, the baseboard management and control system is not restarted. In this way, after the test case script a21 is executed, if the baseboard management and control system of server A needs to be accessed when continuing to execute the test case script a22, a failure to access the baseboard management and control system will occur, making it impossible for the automated test to continue. That is, the reliability of the server automated batch test is not high enough.

[0038] In view of this, the present disclosure provides a network fault repair method, which can repair the network fault of the server in the scenario of server automated batch testing, thereby improving the reliability of server automated batch testing. The network fault repair method can be applied to Figure 1 The test system in the test system, or the electronic equipment used to run the test system. The electronic equipment may include but is not limited to tablet computers, laptop computers, desktop computers, servers, controllers, etc. Figure 2 , which is a flow chart of a network fault repair method provided in some embodiments of the present disclosure. Figure 2 In the network fault repair method, the following steps may be included:

[0039] Step S201: Acquire first network information between each server operating system and second network information between each server management system.

[0040] In this embodiment, the server operating system refers to the server's OS (Operating System), which is primarily used for the server's device management, file management system, memory management, multitasking, and user interface. The server management system refers to the server's baseboard management and control system, which is primarily used for device information management, status monitoring, and remote control. Generally speaking, the difference between a server operating system and a server management system is that a server operating system primarily provides a stable operating environment for applications and services and manages the server's hardware and software resources, while a server management system primarily monitors and manages the server's hardware and provides remote management capabilities.

[0041] The first network information may represent the communication quality between the test system and each server operating system, including but not limited to the disconnection status (i.e., connected or disconnected) between the test system and each server operating system, the first response time of the server operating system when communicating with each server operating system, the packet loss rate during the communication process, etc. Similarly, the second network information may represent the communication quality between the test system and each server management system, including but not limited to the disconnection status between the test system and each server management system, the first response time of the server management system when communicating with each server management system, the packet loss rate during the communication process, etc.

[0042] Specifically, in a server, the server operating system and the server management system may each have a corresponding IP address. The first network information and the second network information may be obtained based on a network diagnostic command and the IP addresses of the server operating system and the server management system. For example, by using the ping command, a data packet may be sent to the IP address of the server operating system, and the first network information may be obtained based on the response of the server operating system.

[0043] Furthermore, in the scenario of automated batch server testing, when executing certain test tasks on a server, the test system and the server operating system or server management system are inherently disconnected. If, during the execution of these specific test tasks, a disconnection between the test system and the server operating system or server management system is detected using techniques such as the ping command, this does not necessarily indicate a true network failure. This is because, under normal circumstances, the test case scripts used to execute these specific test tasks will include network recovery instructions. After the test tasks are completed, the network recovery instructions can be used to restore the connection between the test system and the server operating system or server management system. For example, suppose that when running test case script a3 to test the server's BIOS functions, after the server boots up, it will enter the BIOS system but not the server operating system. That is, during the execution of test case script a3, the test system and the server operating system are disconnected. After the BIOS function test is completed, the network recovery instructions in test case script a3 can be executed to restore the network connection between the test system and the server operating system. Therefore, during the execution of the test case script a3, the disconnection between the test system and the server operating system detected by technical means such as the ping command is not accurate.

[0044] In view of this, in this embodiment, the first network information can be obtained based on the following logic: starting from a specified time point, if the test system and the server operating system are detected to be disconnected for a first preset duration, then it can be determined that the test system and the server operating system are disconnected; conversely, it can be considered that the test system and the server operating system are connected. The first preset duration can be the maximum execution time of a single test task. In this way, the situation where the network between the test system and the server operating system is normally disconnected can be effectively avoided, thereby ensuring the accuracy of the obtained first network information.

[0045] Furthermore, if it is determined that the test system is connected to the server operating system, the duration from the specified time point to the first receipt of a response from the server operating system indicating a successful connection can be used as the first response duration of the server operating system. For example, assuming that 10:30:20 is used as the specified time point, and a ping command is executed every 1 second, if the ping commands sent between the 20th and 39th seconds all receive responses indicating a connection failure, and the first response indicating a successful connection is received at the 40th second, then the duration between the 20th and 40th seconds (i.e., 20 seconds) can be used as the first response duration of the server operating system.

[0046] Using a similar principle to obtaining the first network information, the second network information can be obtained based on the following logic: starting from a specified time point, if the test system and the server management system are detected to be disconnected for a first preset period of time, then the test system and the server management system can be determined to be disconnected; conversely, the test system and the server management system can be considered to be connected. Furthermore, if the test system and the server management system are determined to be connected, the duration from the specified time point to the first receipt of a response from the server management system indicating a successful connection can be used as the second response duration of the server management system.

[0047] The first network information and the second network information can be used to reflect whether a network failure has occurred between the test system and the server. Specifically, if, based on the first network information, it is determined that the first response time of each server operating system is relatively long, or that relatively few server operating systems are connected to the test system, this can indicate that there is a high probability of a network failure between the test system and the server operating system. Similarly, if, based on the second network information, it is determined that the first response time of each server management system is relatively long, or that relatively few server management systems are connected to the test system, this can indicate that there is a high probability of a network failure between the test system and the server management system.

[0048] Step S202: determining the overall accessibility index of one or more servers currently being tested based on the first network information and the second network information.

[0049] The overall accessibility index is used to indicate the probability of network failure for one or more servers being tested. This means that all servers being tested are considered a server set, and the overall accessibility index is used to indicate the probability of network failure for that server set as a whole.

[0050] In this embodiment, the first network information and the second network information may be fused and calculated to determine the overall accessibility index. For example, the overall accessibility index may be calculated as follows:

[0051] Calculate the average of the first response time of each server operating system and the second response time of each server management system to obtain the overall response time of the one or more servers currently being tested;

[0052] Determine the total number of server operating systems and server management systems connected to the test system;

[0053] The product of the overall response time and the total number of systems is taken as the overall accessibility index.

[0054] In subsequent embodiments of the present disclosure, a better method for calculating the overall accessibility index is provided, which will not be described here in detail.

[0055] Step S203 : When the overall accessibility index meets the preset conditions, the target server having the network failure is determined according to the test tasks executed on each server, and the network failure of the target server is repaired.

[0056] Specifically, depending on the calculation method of the overall accessibility index, the overall accessibility index may be inversely proportional or directly proportional to the probability of a network failure occurring in the server set. Specifically, if the overall accessibility index is inversely proportional to the probability of a network failure occurring in the server set, if the overall accessibility index is less than a first threshold, it may indicate that the overall accessibility index meets the preset condition. If the overall accessibility index is directly proportional to the probability of a network failure occurring in the server set, if the overall accessibility index is greater than a second threshold, it may indicate that the overall accessibility index meets the preset condition.

[0057] Based on the above description, when the overall accessibility index meets the preset conditions, it means that the probability of a network failure in the server set is relatively high. In this case, one or more servers currently being tested can be traversed, and the network disconnection between the test system and each server operating system or server management system can be detected. Among them, for the server operating system A1 and server management system A2 in any server A, if it is detected that there is a network disconnection between the test system and the server operating system A1 or the server management system A2, then according to the test task performed on server A, it is determined whether the network disconnection between the test system and the server operating system A1 or the server management system A2 is normal. If it is normal, it is determined that there is no network failure in server A. If it is abnormal, it is determined that server A is the target server with a network failure. For example, suppose that during the traversal process, it is detected that there is a network disconnection between the test system and the server operating system A1. In this case, if it is found that the test task performed on server A is to detect whether the BIOS function is normal (without entering the server operating system A1), it can be determined that the network disconnection between the test system and the server operating system A1 is normal, that is, there is no network failure in server A. On the contrary, if it is found that the test task performed on server A is to test whether the server management system A2 can accurately control the fan in server A (without affecting the server operating system A1), it can be determined that the network disconnection between the test system and the server operating system A1 is abnormal. Therefore, it can be determined that server A is the target server with a network failure.

[0058] After identifying the target server experiencing a network failure, the network failure can be automatically repaired using pre-set instructions, ensuring the normal execution of subsequent test tasks. For example, if a network disconnection is detected between the test system and server operating system A1, and analysis reveals that the cause of the disconnection is that server A is not started, the smart device in the area where server A is located can be controlled to start server A. This resolves the network disconnection between the test system and server operating system A1, ensuring the normal execution of subsequent test tasks.

[0059] To sum up, in some embodiments of the present disclosure, the overall accessibility index determined based on the above-mentioned first network information and second network information can reflect the overall probability of a network failure occurring in one or more servers currently being tested. Furthermore, when the overall accessibility index meets the preset conditions, the target server with a network failure can be accurately located and the network failure of the target server can be repaired based on the test tasks performed on each server, thereby preventing the problem of server test failure caused by network failure, thereby greatly improving the reliability of automated batch testing of servers.

[0060] The solution of this application is further explained below.

[0061] In some embodiments, determining the overall accessibility index of the currently tested one or more servers based on the first network information and the second network information in step S201 may be performed by following steps 1) and 2):

[0062] 1) Based on the first network information, a first accessibility index of the server operating system is determined, and based on the second network information, a second accessibility index of the server management system is determined.

[0063] Specifically, based on the disconnection status between each server operating system, the proportion of server operating systems in a connected state can be counted, and based on a preset operating system response time threshold, the proportion of server operating systems in a connected state and the first response time of the server operating system in a connected state, the first accessibility index of the server operating system can be determined.

[0064] The operating system response time threshold can be the maximum execution time of a single test task. The first accessibility index is used to characterize the probability of network failure of all server operating systems currently being tested. Its calculation formula can be shown as expression (1):

[0065]

[0066] Among them, AI osIndicates the first accessibility index of the server operating system;

[0067] N os Indicates the percentage of server operating systems in connected state;

[0068] T os_resp Indicates the first response time of the server operating system in the connected state. If multiple server operating systems have their own corresponding first response time, T os_resp It may be a value obtained by averaging the first response times of the multiple server operating systems;

[0069] T os_threah Indicates the operating system response time threshold.

[0070] Similarly, based on the disconnection status between each server management system, the proportion of server management systems in a connected state can be counted, and based on the preset management system response time threshold, the proportion of server management systems in a connected state and the second response time of the server management system in a connected state, the second accessibility index of the server management system can be determined.

[0071] The calculation formula of the accessibility index of the server management system can be shown as expression (2):

[0072]

[0073] Among them, AI bmc A second accessibility index representing the server management system;

[0074] N bmc Indicates the percentage of server management systems that are in a connected state;

[0075] T bmc_resp Indicates the first response time of the server management system in the connected state. If multiple server management systems have their own corresponding first response time, T bmc_resp It may be a value obtained by averaging the first response times of the multiple server management systems;

[0076] T bmc_threah Indicates the operating system response time threshold.

[0077] 2) The first accessibility index and the second accessibility index are fused and calculated to obtain an overall accessibility index.

[0078] Specifically, the first accessibility index and the second accessibility index can be added together to obtain the overall accessibility index. That is, the overall accessibility index AI can be expressed as follows:

[0079]

[0080] The overall accessibility index AI obtained based on expression (3) can be inversely proportional to the probability of network failure occurring in the server set.

[0081] In the above embodiment, when determining the overall accessibility index AI, the number of server operating systems and server management systems in a connected state, as well as the response time of the server operating systems and server management systems are comprehensively considered, and the obtained overall accessibility index is more objective and accurate.

[0082] In some embodiments, repairing the network failure of the target server in step S203 includes:

[0083] According to the specified repair operation steps, control the intelligent devices in the area where the target server is located to repair the network failure of the target server;

[0084] If the network failure of the target server fails to be repaired by the smart device, if there is a system in the server management system and the server operating system that is in a connected state, the network failure of the target server is repaired based on the system in the connected state;

[0085] If the network failure of the target server cannot be repaired by the connected system, if there is a backup server that is the same as the target server, the backup server is used to replace the target server.

[0086] Specifically, when repairing a network fault, you can try to repair the target server's network fault in the order of smart device > server management system or server operating system > backup server. This way, you can maximize the guarantee that the target server's network fault can be repaired.

[0087] In some embodiments, controlling the smart devices in the area where the target server is located to repair the network failure of the target server according to the specified repair operation steps may include:

[0088] When the target server is in a shutdown state, control the intelligent device to power on the target server;

[0089] When the target server is powered on, the intelligent device is controlled to connect to the hardware interface of the target server, and the software configuration of the target server is updated through the hardware interface to repair the network failure of the target server.

[0090] Specifically, updating the software configuration of the target server through the hardware interface may include:

[0091] 1) Update the target server's network configuration. For example, if the target server is currently in the BIOS system, set the target server to boot into the server operating system the next time it boots. Once this is set, reboot the target server to boot into the server operating system. This will restore the network connection between the test system and the server operating system.

[0092] 2) Reinstall the server operating system or server management system.

[0093] Specifically, by reinstalling the system, a failed server operating system or server management system can be repaired, and the network connection between the test system and the server operating system or server management system can be restored.

[0094] In some embodiments, repairing a network failure of a target server based on a system in a connected state may include:

[0095] When the server operating system is in a connected state and a network failure occurs in the server management system, repair the network failure of the server management system through the server operating system;

[0096] When the server management system is in a connected state and a network fault occurs in the server operating system, the network fault of the server operating system is repaired through the server management system.

[0097] For example, while the server management system is connected, you can power on, reboot, or start the target server in the device management interface of the server management system, thereby starting the server operating system and restoring the network connection between the test system and the server operating system. For another example, while the server operating system is connected, executing a specified command in the server operating system can start the server management system and restore the network connection between the test system and the server management system.

[0098] In some embodiments, after replacing the target server with the backup server, the method of the present disclosure may further include:

[0099] Set the media access control address of the backup server to the media access control address of the target server;

[0100] Setting the Internet Protocol address of the server operating system of the backup server to the Internet Protocol address of the server operating system of the target server;

[0101] The Internet Protocol address of the server management system of the backup server is set to the Internet Protocol address of the server management system of the target server.

[0102] Specifically, a media access control address is also known as a MAC address, and an internet protocol address is also known as an IP address. By setting the backup server's media access control address and internet protocol address to be the same as the target server's, the test server can establish a connection with the backup server based on the target server's address and use the backup server as the target server to complete the test task for the target server.

[0103] This completes the entire description of the method disclosed herein.

[0104] See also Figure 3 , which is a module diagram of a network fault repair device provided by an embodiment of the present disclosure. Figure 3 In the embodiment, the network fault recovery device includes:

[0105] The network information acquisition module 301 is used to acquire first network information between each server operating system and second network information between each server management system;

[0106] an accessibility index determination module 302 for determining an overall accessibility index of one or more servers currently being tested based on the first network information and the second network information;

[0107] The fault repair module 303 is used to determine the target server with network failure according to the test tasks performed on each server when the overall accessibility index meets the preset conditions, and repair the network failure of the target server.

[0108] In some embodiments, the accessibility index determination module 302 is specifically configured to:

[0109] Determining a first accessibility index of the server operating system based on the first network information, and determining a second accessibility index of the server management system based on the second network information;

[0110] The first accessibility index and the second accessibility index are fused and calculated to obtain an overall accessibility index.

[0111] In some embodiments, the first network information includes the disconnection status between each server operating system and the first response time of the server operating system when communicating with each server operating system; the second network information includes the disconnection status between each server management system and the second response time of the server management system when communicating with each server management system; the accessibility index determination module 302 is specifically used to:

[0112] Based on the disconnection status of each server operating system, the percentage of server operating systems that are connected is counted;

[0113] Determining a first accessibility index of the server operating system based on a preset operating system response time threshold, a proportion of server operating systems in a connected state, and a first response time of the server operating systems in a connected state;

[0114] Based on the disconnection status with each server management system, the percentage of server management systems in the connected state is counted;

[0115] A second accessibility index of the server management system is determined based on a preset management system response time threshold, a proportion of server management systems in a connected state, and a second response time of the server management system in a connected state.

[0116] In some embodiments, the fault recovery module 303 is specifically configured to:

[0117] According to the specified repair operation steps, control the intelligent devices in the area where the target server is located to repair the network failure of the target server;

[0118] If the network failure of the target server fails to be repaired by the smart device, if there is a system in the server management system and the server operating system that is in a connected state, the network failure of the target server is repaired based on the system in the connected state;

[0119] If the network failure of the target server cannot be repaired by the connected system, if there is a backup server that is the same as the target server, the backup server is used to replace the target server.

[0120] In some embodiments, the fault recovery module 303 is specifically configured to:

[0121] When the target server is in a shutdown state, control the intelligent device to power on the target server;

[0122] When the target server is powered on, the intelligent device is controlled to connect to the hardware interface of the target server, and the software configuration of the target server is updated through the hardware interface to repair the network failure of the target server.

[0123] In some embodiments, after the target server is replaced with the standby server, the fault recovery module 303 is further configured to:

[0124] Set the media access control address of the backup server to the media access control address of the target server;

[0125] Setting the Internet Protocol address of the server operating system of the backup server to the Internet Protocol address of the server operating system of the target server;

[0126] The Internet Protocol address of the server management system of the backup server is set to the Internet Protocol address of the server management system of the target server.

[0127] In some embodiments, the fault recovery module 303 is specifically configured to:

[0128] When the server operating system is in a connected state and a network failure occurs in the server management system, repair the network failure of the server management system through the server operating system;

[0129] When the server management system is in a connected state and a network fault occurs in the server operating system, the network fault of the server operating system is repaired through the server management system.

[0130] The network fault repair device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0131] The network fault repair device disclosed in the present invention has the same beneficial effects as the above-mentioned network fault repair method, which will not be described in detail here.

[0132] The present disclosure also provides an electronic device having the above Figure 3 The network fault repair device shown.

[0133] See also Figure 4 , is a schematic diagram of the structure of an electronic device provided by some embodiments of the present disclosure. Figure 4As shown, the electronic device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 4 A processor 10 is taken as an example.

[0134] The processor 10 may be a first PCIe device, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CPLD), a field programmable gate array (FPGA), a general purpose array logic (GAL), or any combination thereof.

[0135] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0136] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0137] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0138] The electronic device further includes a communication interface 30 for the electronic device to communicate with other devices or a communication network.

[0139] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0140] A portion of the present disclosure may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present disclosure through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes but is not limited to a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0141] Although the embodiments of the present disclosure have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A network fault repair method, characterized in that: Applied to a test system, the test system is used to test at least one server; the method includes: Acquire first network information between each server operating system and second network information between each server management system; determining an overall accessibility index of one or more servers currently being tested based on the first network information and the second network information; When the overall accessibility index satisfies a preset condition, determining a target server having a network failure according to the test tasks performed on each of the servers, and repairing the network failure of the target server; The first network information includes the disconnection status between each of the server operating systems and the first response time of the server operating system when communicating with each of the server operating systems; the second network information includes the disconnection status between each of the server management systems and the second response time of the server management system when communicating with each of the server management systems; and determining the overall accessibility index of the currently tested one or more servers based on the first network information and the second network information includes: Based on the disconnection status between each of the server operating systems, counting the proportion of server operating systems in a connected state; Determining a first accessibility index of the server operating system based on a preset operating system response time threshold, a proportion of server operating systems in a connected state, and a first response time of the server operating systems in a connected state; Based on the disconnection status between each of the server management systems, counting the proportion of server management systems in a connected state; Determining a second accessibility index of the server management system based on a preset management system response time threshold, a proportion of server management systems in a connected state, and a second response time of the server management system in a connected state; The first accessibility index and the second accessibility index are fused and calculated to obtain the overall accessibility index.

2. The method according to claim 1, characterized in that The repairing of the network failure of the target server includes: According to the specified repair operation steps, control the smart devices in the area where the target server is located to repair the network failure of the target server; In the case where the network fault of the target server fails to be repaired by the smart device, if there is a system in a connected state between the server management system and the server operating system, the network fault of the target server is repaired based on the system in a connected state; In the case that the network failure of the target server cannot be repaired by the system in the connected state, if there is a backup server identical to the target server, the target server is replaced by the backup server.

3. The method according to claim 2, characterized in that The step of controlling the smart devices in the area where the target server is located to repair the network failure of the target server according to the specified repair operation steps includes: When the target server is in a shutdown state, controlling the smart device to power on the target server; When the target server is powered on, the smart device is controlled to connect to the hardware interface of the target server, and the software configuration of the target server is updated through the hardware interface to repair the network failure of the target server.

4. The method according to claim 2, characterized in that After replacing the target server with the standby server, the method further includes: Setting the media access control address of the backup server to the media access control address of the target server; Setting the Internet Protocol address of the server operating system of the standby server to the Internet Protocol address of the server operating system of the target server; The Internet Protocol address of the server management system of the standby server is set as the Internet Protocol address of the server management system of the target server.

5. The method according to claim 2, characterized in that The repairing of the network failure of the target server based on the system in the connected state includes: When the server operating system is in a connected state and a network failure occurs in the server management system, repairing the network failure of the server management system by using the server operating system; When the server management system is in a connected state and a network failure occurs in the server operating system, the network failure of the server operating system is repaired by the server management system.

6. A network fault repair device, characterized in that: Applicable to a test system, the test system is used to test at least one server; the device includes: A network information acquisition module, configured to acquire first network information between each server operating system and second network information between each server management system; an accessibility index determination module, configured to determine an overall accessibility index of one or more servers currently being tested based on the first network information and the second network information, wherein the first network information includes a disconnection status with each of the server operating systems and a first response time of the server operating system when communicating with each of the server operating systems; and the second network information includes a disconnection status with each of the server management systems and a second response time of the server management system when communicating with each of the server management systems. When determining the overall accessibility index, the module calculates the proportion of server operating systems in a connected state based on the disconnection status with each of the server operating systems; determines the first accessibility index of the server operating system based on a preset operating system response time threshold, the proportion of server operating systems in a connected state, and the first response time of the server operating system in a connected state; calculates the proportion of server management systems in a connected state based on the disconnection status with each of the server management systems; and determines the second accessibility index of the server management system based on a preset management system response time threshold, the proportion of server management systems in a connected state, and the second response time of the server management system in a connected state; and fuses the first accessibility index and the second accessibility index to obtain the overall accessibility index. The fault repair module is used to determine the target server with network failure according to the test tasks performed on each of the servers when the overall accessibility index meets the preset conditions, and repair the network failure of the target server.

7. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the network fault repair method according to any one of claims 1 to 5 by executing the computer instructions.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the network fault repair method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Server capable of achieving network abnormity repair and network abnormity repair method thereof

    CN103490946A

  • Artificial intelligence early warning system

    CN109447048A