Server troubleshooting method, device, storage medium and program product

By combining a preset fault information set with a preset positioning device, efficient and accurate troubleshooting of server faults is achieved, solving the time-consuming and error-prone problems of existing technologies and improving troubleshooting efficiency and automation levels.

CN120469844BActive Publication Date: 2025-10-10INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510954764.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-10
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

Existing technologies require a lot of time and are prone to errors in server fault troubleshooting. They also require high technical skills from troubleshooters and the manual troubleshooting process is complicated, resulting in low efficiency.

Method used

By determining the troubleshooting steps based on a preset fault information set, the theoretical location information of the target node to be troubleshooted is displayed to the user, and the actual location information is displayed using a preset positioning device. Fault troubleshooting is performed by combining the theoretical and actual location information, and the troubleshooting conclusion is output.

Benefits of technology

It improves the accuracy and efficiency of troubleshooting, lowers the professional threshold, simplifies the troubleshooting process, reduces hardware loss and manual workload, and improves the automation and intelligence level of fault diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469844B_ABST
    Figure CN120469844B_ABST
Patent Text Reader

Abstract

The application provides a server troubleshooting method and device, a storage medium and a program product, which can be applied to the technical field of computers. The server troubleshooting method comprises the following steps: determining a troubleshooting step corresponding to a target server fault phenomenon input by a user based on a preset fault information set; determining a target node to be troubleshooted corresponding to the troubleshooting step, and showing the user theoretical position information of the node to be troubleshooted in a target server; according to an execution sequence of the troubleshooting step, sending the theoretical position information to a preset positioning device, so that the preset positioning device shows the user actual position information of the target node to be troubleshooted in the target server based on the received theoretical position information; and in response to obtaining troubleshooting feedback information, outputting a troubleshooting conclusion for the target server fault phenomenon based on the troubleshooting feedback information, wherein the troubleshooting feedback information is obtained by the user based on the theoretical position information and the actual position information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computers, and more particularly to a server troubleshooting method, device, storage medium and program product. BACKGROUND

[0002] When a server fails, maintenance personnel need to determine the overall link situation of the server in combination with the board schematic diagram, server link diagram, etc., and then determine the fault location. For example, if a hard disk cannot be detected in a server during the test phase, the overall hard disk link needs to be checked. However, the process of checking the overall link situation manually needs to refer to a large number of schematic diagrams, link diagrams, etc., and also needs to find the entity location of the content to be checked on the server physical object. The reference process and the finding process need to spend a lot of time, and the checking process is prone to errors. SUMMARY

[0003] In view of the above problems, the present application provides a server troubleshooting method, device, storage medium and program product.

[0004] According to a first aspect of the present application, a server troubleshooting method is provided, comprising: determining a troubleshooting step corresponding to a target server fault phenomenon input by a user based on a preset fault information set; determining a target to-be-troubleshooted node corresponding to the troubleshooting step, and showing the user theoretical position information of the target to-be-troubleshooted node in a target server; according to an execution order of the troubleshooting step, sending the theoretical position information to a preset positioning device, so that the preset positioning device shows the user actual position information of the target to-be-troubleshooted node in the target server based on the received theoretical position information; in response to obtaining troubleshooting feedback information, outputting a fault troubleshooting conclusion for the target server fault phenomenon based on the troubleshooting feedback information, wherein the troubleshooting feedback information is obtained by the user based on the theoretical position information and the actual position information executing the troubleshooting step.

[0005] A second aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0006] A third aspect of the present application further provides a computer-readable storage medium having a computer program or instructions stored thereon, wherein the computer program or instructions are executed by a processor to implement the steps of the above method.

[0007] A fourth aspect of the present application further provides a computer program product comprising a computer program or instructions, wherein the computer program or instructions are executed by a processor to implement the steps of the above method. BRIEF DESCRIPTION OF DRAWINGS

[0008] The above content of the present application and other purposes, features and advantages will be more apparent through the following description of embodiments of the present application with reference to the accompanying drawings, in which:

[0009] Figure 1 An application scenario diagram of a server troubleshooting method, device, storage medium and program product according to an embodiment of the present application is shown.

[0010] Figure 2 A flowchart of a server troubleshooting method according to an embodiment of the present application is shown.

[0011] Figure 3 A flowchart of a server troubleshooting method according to another embodiment of the present application is shown.

[0012] Figure 4 A schematic diagram of a user display interface according to an embodiment of the present application is shown.

[0013] Figure 5 A structural block diagram of a server troubleshooting device according to an embodiment of the present application is shown.

[0014] Figure 6 A block diagram of an electronic device suitable for implementing a server troubleshooting method according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0015] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. It should be understood, however, that the description is merely exemplary of the present application, and is not intended to limit the scope of the present application. In the following detailed description of embodiments of the present application, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that one or more embodiments of the present application can be practiced without these specific details. In other instances, well-known structures and functions have not been described in detail in order to avoid obscuring aspects of the present application.

[0016] The terms used herein are merely used to describe specific embodiments, and are not intended to limit the present application. The terms "include", "comprise" and the like used herein indicate the presence of the features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0017] All terms used herein (including technical and scientific terms) have meanings commonly understood by one of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having meanings consistent with the context of the specification, and should not be interpreted in an idealized or overly formal manner.

[0018] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0019] In the server industry, during small-batch trial production, early shipments, or normal shipments of products, if problems such as poor internal circuit contact, damaged components, or bandwidth reduction in server links occur, maintenance personnel are usually required to use a multimeter to perform continuity tests on the cables and board ends of the server using server-related schematics, server link diagrams, and other information to locate the failure point or link.

[0020] Among them, the schematic diagram is also called the electrical schematic diagram, which includes various circuit component symbols and the connection methods between them. By identifying the various circuit component symbols and connection methods in the schematic diagram, you can understand the actual working principle of the circuit. The server link diagram can also be called the server structure topology diagram, which can show the hardware composition inside the server, such as the central processing unit (CPU), memory, hard disk, network interface, etc. The server link diagram can be used to understand the server architecture and hardware configuration, and because the server link diagram shows the connection status and network layout of network devices, it helps to analyze and optimize network performance. You can use the topology diagram template to customize the creation of the topology diagram and modify and adjust the topology diagram as needed.

[0021] The related method generally determines the overall link condition of the server by cross-checking the principle diagram and the server link diagram. For example, if a hard disk of a server cannot be detected in the test stage, the overall hard disk link needs to be checked, which includes, for example, hard disk-backplane-backplane cable-mainboard-CPU-power supply. Among them, "power supply-CPU-mainboard" in the overall hard disk link needs to be checked based on the principle diagram of the mainboard, "mainboard-backplane cable-backplane" needs to be checked based on the server link diagram, and "backplane-hard disk" needs to be checked based on the principle diagram of the backplane. The checking process needs to be checked based on a large number of principle diagrams and link diagrams, and the checking steps need to be determined manually, and it also needs to be converted to the actual object to find the specific point information of each component to check the fault of the component, so the overall checking process needs to spend a lot of time, which leads to low checking efficiency, and the checking process is prone to errors, and in the case of error in some point, the whole checking link may be wrong, so the technical level of the checking personnel is required, and the operation accuracy of the checking personnel is required, and the manual checking process is prone to error diagnosis due to errors.

[0022] Further, the related server whole machine on-off fault checking method is generally a single device checking, for example, in the case that a hard disk of a server cannot be detected in the test stage, the method of checking the whole hard disk link is generally to replace and cross-verify the hard disk, backplane, backplane cable, CPU, memory, power supply, and mainboard one by one to locate the faulty device. However, the replacement and cross-verification needs to spend a lot of time, and the operation of disassembling and assembling the whole server also needs to consume a lot of time, in addition, the replacement and cross-verification is prone to cause damage to the components due to the disassembly and assembly operation, which leads to the destruction of the on-site server environment, if the fault reason is not determined in the checking process, and the fault phenomenon has disappeared, because the on-site server environment has been destroyed, it is difficult for the checking personnel to further analyze the fault phenomenon.

[0023] Therefore, the embodiments of the present application provide a server fault checking method, which comprises: determining a checking step corresponding to a target server fault phenomenon input by a user based on a preset fault information set; determining a target to-be-checked node corresponding to the checking step, and showing the user theoretical position information of the target to-be-checked node in a target server; sending the theoretical position information to a preset positioning device according to an execution order of the checking step, so that the preset positioning device shows the user actual position information of the target to-be-checked node in the target server based on the received theoretical position information; and outputting a fault checking conclusion for the target server fault phenomenon based on checking feedback information in response to obtaining the checking feedback information, wherein the checking feedback information is obtained by the user based on the theoretical position information and the actual position information.

[0024] Figure 1 An application scenario diagram of the server troubleshooting method, device, storage medium and program product according to the embodiments of the present application is shown.

[0025] As shown in Figure 1 , the application scenario according to the embodiments can include a digital process system 101 and a preset positioning device 102. The digital process system 101 can show the theoretical position information of the target to-be-troubleshoot node in the target server, for example, in the case where the target to-be-troubleshoot node includes node 1, node 2…node n, the digital process system 101 can show the theoretical position information of node 1, node 2…node n respectively, for example, show theoretical position information 1, theoretical position information 2…theoretical position information n. The digital process system 101 can also send the theoretical position information to the preset positioning device 102, and the preset positioning device 102 can determine and show the actual position information of the target to-be-troubleshoot node in the target server based on the received theoretical position information, for example, the actual position information can include actual position information 1, actual position information 2…actual position information n.

[0026] For example, in the case where the target server fails, the user can input the target server failure phenomenon into the digital process system 101, determine the troubleshooting steps corresponding to the target server failure phenomenon input by the user based on the preset failure information set; determine the target to-be-troubleshoot node corresponding to the troubleshooting steps, and show the user the theoretical position information of the target to-be-troubleshoot node in the target server; according to the execution order of the troubleshooting steps, send the theoretical position information to the preset positioning device 102, so that the preset positioning device 102 shows the actual position information of the target to-be-troubleshoot node in the target server based on the received theoretical position information; in response to obtaining the troubleshooting feedback information, output the troubleshooting conclusion for the target server failure phenomenon based on the troubleshooting feedback information, wherein the troubleshooting feedback information is obtained by the user based on the theoretical position information and the actual position information executing the troubleshooting steps.

[0027] The server troubleshooting method according to the embodiments of the present application will be described in detail below. Figures 2-4

[0028] Figure 2 A flowchart of the server troubleshooting method according to the embodiments of the present application is shown. As Figure 2 shown, the server troubleshooting method of the embodiments includes operations S210-S240.

[0029] In operation S210, the troubleshooting steps corresponding to the target server failure phenomenon input by the user are determined based on the preset failure information set.

[0030] ​Since different fault phenomena can correspond to different troubleshooting steps, a plurality of historical fault phenomena of the server can be acquired in advance, and historical troubleshooting steps for the plurality of historical fault phenomena can be acquired, for example, a plurality of historical fault phenomena of different types of servers respectively, and historical troubleshooting steps for the plurality of historical fault phenomena respectively. Based on the preset fault information set, historical fault phenomena and historical troubleshooting steps corresponding to different types of servers can be determined. For example, when the user inputs the target server fault phenomenon of the target server, the troubleshooting step corresponding to the target server fault phenomenon of the target server can be determined from the preset fault information set based on the type of the target server and the target server fault phenomenon.

[0031] In operation S220, the target to-be-troubleshoot node corresponding to the troubleshooting step is determined, and the user is shown the theoretical position information of the target to-be-troubleshoot node in the target server.

[0032] The preset fault information set can also include to-be-troubleshoot nodes corresponding to the troubleshooting step, for example, including components of the server. Different troubleshooting steps can correspond to corresponding to-be-troubleshoot nodes, for example, troubleshooting step 1 can correspond to to-be-troubleshoot nodes 1 and 2, and troubleshooting step 2 can correspond to to-be-troubleshoot node 3. The to-be-troubleshoot nodes need to be troubleshooted to determine whether the fault phenomenon is caused by the fault of the to-be-troubleshoot node.

[0033] For example, the preset fault information set can be imported into the digital process system in advance, and after the user inputs the target server fault phenomenon into the digital process system, the troubleshooting step corresponding to the target server fault phenomenon can be automatically matched based on the type of the target server and the target server fault phenomenon, and the target to-be-troubleshoot node can be automatically determined through the preset fault information set.

[0034] A three-dimensional model corresponding to the target server can be shown, for example, a three-dimensional model of a component, which can be generated by a three-dimensional rendering software built in the digital process system. The theoretical position information of the target to-be-troubleshoot node in the target server can include: presenting the three-dimensional model of the target to-be-troubleshoot node in a flashing animation. Further, the theoretical position information of the target to-be-troubleshoot node in the target server can also include: labeling the target to-be-troubleshoot node related information, for example, labeling the pin number, etc.

[0035] For example, in the case where the target to-be-troubleshoot node corresponding to the troubleshooting step 1 includes the component 1 and the component 2, the user can be shown the three-dimensional model of the component 1 and the component 2 in the server whole machine respectively, and the above three-dimensional model can be highlighted and flashed, and the pin number of the above component can also be labeled, for example, "test point: Pin1 and Pin2" can be labeled.

[0036] For example, the user can be sequentially shown the theoretical position information of the target to-be-troubleshoot node in the target server according to the execution order of the troubleshooting steps, for example, in the troubleshooting step 1 stage, the three-dimensional model of the target to-be-troubleshoot node corresponding to the troubleshooting step 1 is highlighted and flashed, and in the troubleshooting step 2 stage, the three-dimensional model of the target to-be-troubleshoot node corresponding to the troubleshooting step 2 is highlighted and flashed. Specifically, the highlighting and flashing of the three-dimensional model of the target to-be-troubleshoot node in the troubleshooting step 1 can be stopped in the troubleshooting step 2 stage, or the highlighting and flashing of the three-dimensional model of the target to-be-troubleshoot node in the troubleshooting step 1 can be continued.

[0037] In operation S230, the theoretical position information is sent to the preset positioning device according to the execution order of the troubleshooting steps, so that the preset positioning device shows the actual position information of the target to-be-troubleshoot node in the target server to the user based on the received theoretical position information.

[0038] For example, the theoretical position information can be sequentially sent to the preset positioning device according to the execution order of the troubleshooting steps, so that the preset positioning device sequentially shows the actual position information of the target to-be-troubleshoot node in the target server to the user based on the sequentially received theoretical position information.

[0039] For example, in the troubleshooting step 1 stage, after the theoretical position information 1 corresponding to the troubleshooting step 1 stage is determined, the theoretical position information 1 is sent to the preset positioning device, and the actual position information 1 corresponding to the target to-be-troubleshoot node in the troubleshooting step 1 is determined by the preset positioning device based on the received theoretical position information 1, and the actual position information 1 is shown to the user. Then, in the troubleshooting step 2 stage, the theoretical position information 2 corresponding to the troubleshooting step 2 is sent to the preset positioning device, and the actual position information 2 is determined and shown to the user by the preset positioning device based on the received theoretical position information 2.

[0040] The actual position information of the target to-be-troubleshoot node in the target server can be shown to the user by shining a laser light on the components of the server physical object corresponding to the target to-be-troubleshoot node.

[0041] For example, the mechanical arm system can be used to accurately position the overall structure of the server based on the edge of the server case as the coordinate position, and then the laser light is shone on the components of the server physical object, and the user can know the actual position of the target to-be-troubleshoot node in the target server through the shining position.

[0042] After the user is shown the theoretical location information and the actual location information of the target to-be-troubleshoot node in the target server, the user can quickly and accurately locate the position of the target to-be-troubleshoot node in the server entity based on the theoretical location information and the actual location information, and perform troubleshooting. For example, if the target to-be-troubleshoot node corresponding to the troubleshooting step 1 includes the component 1, the user can directly determine the shape, position, and other information of the component 1 based on the 3D model of the component 1 and the lighting of the component 1 of the server physical object, and perform troubleshooting.

[0043] In operation S240, in response to obtaining the troubleshooting feedback information, a troubleshooting conclusion for the target server fault phenomenon is output based on the troubleshooting feedback information, wherein the troubleshooting feedback information is obtained by the user based on the theoretical location information and the actual location information.

[0044] After the user performs the troubleshooting step based on the theoretical location information and the actual location information on the target to-be-troubleshoot node, the troubleshooting feedback information corresponding to the target to-be-troubleshoot node is obtained. For example, after the user uses the multimeter to troubleshoot the target to-be-troubleshoot node, the current value, voltage value, and other troubleshooting feedback information corresponding to the target to-be-troubleshoot node are obtained. After obtaining the troubleshooting feedback information, the user can input the troubleshooting feedback information into the digital process system, so that the system can generate a troubleshooting conclusion based on the troubleshooting feedback information. The troubleshooting conclusion can include the name of the fault node, the component number, the cause of the fault, and the like.

[0045] According to the embodiments of the present application, the theoretical location information of the target to-be-troubleshoot node in the target server is shown to the user by highlighting and flickering the 3D model corresponding to the target to-be-troubleshoot node, and the actual location information of the target to-be-troubleshoot node in the target server is shown to the user by lighting the component of the server physical object. On the one hand, this can achieve relatively accurate fault location, avoid blind troubleshooting and component cross-replacement verification, effectively reduce hardware loss and troubleshooting time; on the other hand, it can convert abstract circuit principles and fault locations into intuitive visual information, so that non-professionals can easily understand and determine the location of the to-be-troubleshoot node without having to manually interpret schematic diagrams, circuit diagrams, and other professional drawings to find the actual location of the to-be-troubleshoot node. This simplifies the complex troubleshooting process, reduces the professional threshold and manual workload, and improves troubleshooting efficiency and accuracy. In addition, by integrating the above functions in the same system by combining the determination of the troubleshooting step corresponding to the target server fault phenomenon based on the preset fault information set and the showing of the theoretical location information and the actual location information of the target to-be-troubleshoot node in the target server to the user, a complete server troubleshooting system can be formed, thereby improving the automation and intelligence level of fault diagnosis.

[0046] According to an embodiment of the present application, the preset fault information set includes a correspondence between a server fault phenomenon and a troubleshooting path, and the troubleshooting path includes at least one troubleshooting step node and a connection sequence between at least one troubleshooting step node.

[0047] For example, based on the acquired multiple historical fault phenomena and corresponding historical troubleshooting steps, the corresponding relationship between the server fault phenomenon and the troubleshooting steps can be determined, and the execution order of the troubleshooting steps can be determined. For a server fault phenomenon, at least one troubleshooting step node corresponding to the server fault phenomenon can be generated, and each troubleshooting step node can be associated with a troubleshooting step. For example, the troubleshooting step associated with troubleshooting step node 1 is "confirming the on / off status of the hard disk to the motherboard", and the troubleshooting step associated with troubleshooting step node 2 connected to troubleshooting step node 1 is "confirming the on / off status of the motherboard to the power supply". According to the connection order, it can be determined that after executing troubleshooting step 1, troubleshooting step 2 can be executed.

[0048] The troubleshooting steps corresponding to the target server failure phenomenon input by the user may include an initial troubleshooting step and subsequent troubleshooting steps. Determining the troubleshooting steps corresponding to the target server failure phenomenon input by the user based on a preset fault information set may include: determining a target troubleshooting path corresponding to the target server failure phenomenon from the preset fault information set; determining the initial troubleshooting step corresponding to the initial troubleshooting step node in the target troubleshooting path; and determining subsequent troubleshooting steps based on the initial feedback data corresponding to the initial troubleshooting step.

[0049] Figure 3 FIG. 1 shows a flow chart of a server fault troubleshooting method according to another embodiment of the present application. Figure 3 As shown, the server fault troubleshooting method of this embodiment includes operations S310 to S360.

[0050] In operation S310 , a starting troubleshooting step is determined.

[0051] For example, because the preset fault information set includes a correspondence between server fault symptoms and troubleshooting paths, the target server fault symptom can be determined from the preset fault information set to determine the troubleshooting path corresponding to the target server fault symptom, thereby obtaining a target troubleshooting path. The target troubleshooting path can include multiple troubleshooting step nodes, and the troubleshooting step node that is ranked first in the connection order can be used as the starting troubleshooting step node, and the troubleshooting step corresponding to the starting troubleshooting step node can be used as the starting troubleshooting step.

[0052] In operation S320 , the theoretical location information and the actual location information of the target node to be checked corresponding to the initial checking step are displayed.

[0053] For example, the three-dimensional model of the target to-be-troubleshoot node corresponding to the starting troubleshooting step can be highlighted and displayed to show the theoretical position information, and the laser light is shone on the component of the server physical object corresponding to the target to-be-troubleshoot node to show the actual position information.

[0054] In operation S330, the starting feedback data input by the user is received.

[0055] The starting feedback data can include troubleshooting data obtained by the user performing the starting troubleshooting step. For example, the user can perform the starting troubleshooting step on the target to-be-troubleshoot node according to the theoretical position information and the actual position information, such as performing an operation of checking the on-off state of the hard disk to the mainboard, to obtain corresponding troubleshooting data, which includes, for example, the resistance value, the voltage value, and the like of the component corresponding to the starting troubleshooting step. Further, the digital process system can receive the above starting feedback data input by the user.

[0056] In operation S340, the subsequent troubleshooting step is determined.

[0057] According to an embodiment of the present application, the troubleshooting path further includes a feedback data value condition associated with at least one troubleshooting step node. According to the starting feedback data corresponding to the starting troubleshooting step, the subsequent troubleshooting step is determined by matching the starting feedback data with the feedback data value condition associated with the starting troubleshooting step node, and determining the subsequent troubleshooting step according to the matching result.

[0058] For example, the feedback data value condition can represent a standard range that should be met by the feedback data obtained by the user performing the troubleshooting step corresponding to the troubleshooting step node in the case where the to-be-troubleshoot node corresponding to the troubleshooting step node does not have a fault. For example, the troubleshooting step corresponding to the troubleshooting step node 1 is "checking the on-off state of the hard disk to the mainboard", and the feedback data value condition associated with the troubleshooting step node 1 is 0.477Ω to 1.3Ω, indicating that in the case where the to-be-troubleshoot node corresponding to the step does not have a fault, the feedback data obtained by troubleshooting with a multimeter should be between 0.477Ω and 1.3Ω.

[0059] Matching the starting feedback data with the feedback data value condition associated with the starting troubleshooting step node can determine whether the starting feedback data falls within the feedback data value condition. In the case where it is determined that the starting feedback data falls within the feedback data value condition, it indicates that the to-be-troubleshoot node corresponding to the starting troubleshooting step does not have a fault, and therefore the troubleshooting step corresponding to the troubleshooting step node directly connected to the starting troubleshooting step node can be performed next, that is, the subsequent troubleshooting step includes the troubleshooting step corresponding to the troubleshooting step node directly connected to the starting troubleshooting step node.

[0060] If it is determined based on the matching result that the initial feedback data does not fall within the feedback data numerical condition, it indicates that the to-be-investigated node corresponding to the initial investigation step has a fault. Further, the subsequent investigation step can further investigate the to-be-investigated node corresponding to the initial investigation step to determine the fault cause of the to-be-investigated node.

[0061] By determining the target investigation path corresponding to the target server fault phenomenon from the preset fault information set, determining the initial investigation step corresponding to the initial investigation step node in the target investigation path, and matching the initial feedback data with the feedback data numerical condition associated with the initial investigation step node, the subsequent investigation step is determined according to the matching result, which can automatically and accurately determine the investigation step. Compared with artificial device monomer investigation and component-by-component replacement cross verification, the probability of investigation failure is reduced, and the investigation accuracy and efficiency are improved.

[0062] In operation S350, the theoretical position information and the actual position information of the target to-be-investigated node corresponding to the subsequent investigation step are displayed. For example, the three-dimensional model of the target to-be-investigated node corresponding to the subsequent investigation step can be displayed in high light and flicker to display the theoretical position information, and the laser light is projected onto the server physical device corresponding to the target to-be-investigated node to display the actual position information.

[0063] In operation S360, the fault investigation conclusion is output.

[0064] Based on the feedback information of the user for the initial investigation step and the subsequent investigation step, the investigation feedback information corresponding to the target server fault phenomenon can be obtained, the fault cause can be determined according to the investigation feedback information, and the solution to the fault cause can also be determined to obtain the fault investigation conclusion. For example, the target server fault phenomenon input by the user includes that the hard disk at position NO2 cannot be detected, and based on the feedback information of the user for the initial investigation step and the subsequent investigation step, it can be determined that the fault cause is the hard disk backboard monomer fault, and the solution is to replace the hard disk backboard.

[0065] According to the embodiments of the present application, the server fault investigation method can further include the following operations.

[0066] The target identification information of the target to-be-investigated node is determined, wherein the target identification information may, for example, include the component number of the target to-be-investigated node.

[0067] The target coordinate information and the target rendering rule corresponding to the target to-be-investigated node are read from the pre-constructed node mapping relationship set according to the target identification information.

[0068] For example, the pre-constructed node mapping relationship set can include coordinate information corresponding to at least one node to be checked, wherein the at least one node to be checked includes a target node to be checked, and the coordinate information can represent the position of the node to be checked in the server. The node mapping relationship set can also include rendering rules corresponding to at least one node to be checked, wherein the rendering rules can include a target rendering rule, and the target rendering rule can represent a visual marking method corresponding to the node to be checked, such as highlighting the node to be checked, or flashing, or using the above methods together. The rendering rules of the at least one node to be checked can be the same or different, and can be set according to actual conditions.

[0069] Based on the target coordinate information, a target model device corresponding to the target node to be checked in the pre-constructed server model is determined.

[0070] The server model can include models of a plurality of components corresponding to the server, for example, can include three-dimensional models of a plurality of components. According to the target coordinate information, the target model device corresponding to the target node to be checked can be accurately located from the server model, for example, the target coordinate information includes (x1, y1, z1), and the model device at the position (x1, y1, z1) can be located from the server model, thereby determining the target model device corresponding to the target node to be checked.

[0071] Based on the target rendering rule, the target model device is visually marked.

[0072] For example, the visual marking includes highlighting and flashing, and through the visual marking, the user can be shown the theoretical position information of the target node to be checked in the target server.

[0073] By determining the target model device corresponding to the target node to be checked in the pre-constructed server model based on the target coordinate information, and visually marking the target model device based on the target rendering rule, the user can be intuitively shown the theoretical position of the target node to be checked, thereby facilitating the user to accurately locate the target node to be checked.

[0074] According to an embodiment of the present application, the server troubleshooting method further includes determining and showing the user troubleshooting step execution information corresponding to the troubleshooting step, and the troubleshooting step execution information includes at least one of the following: troubleshooting device usage information, node to be checked material information.

[0075] The relevant schematic diagram (such as a board schematic diagram) and the relevant link diagram (such as a link diagram between devices) corresponding to the server can be acquired in advance. Based on the relevant schematic diagram and the relevant link diagram, the component information of the server, such as the component information of a main card, a backboard, a power board, a fan board, a power supply, a memory, a CPU, and a cable, is determined. The working information of the server components is also determined. The working information of the server components includes, for example, pin power supply information, serial general-purpose input / output information, inter-integrated circuit bus signal information, and data transmission link information. For example, the working information of the server components and the component information can be determined by a staff member based on the relevant schematic diagram and the relevant link diagram in advance. By analyzing the component information and the working information of the server components, the overall situation of the server can be determined more accurately, such as the component number, the model of a component, and the connection relationship of the node to be checked. For example, it can be determined that the power supply link of a hard disk at a certain position includes: the hard disk, a hard disk backboard numbered NO1, a hard disk backboard component numbered NO1, a backboard power supply line, a mainboard connector, a mainboard bus, a mainboard power supply interface, and a power supply.

[0076] For example, the device usage information to be checked can include the usage requirement information of a multimeter, such as the usage mode of the multimeter. The usage mode includes, for example, a continuity file, a resistance file, and other files. The device usage information to be checked can also include an automatically generated multimeter picture and the theoretical reference value of the multimeter corresponding to the troubleshooting step. The material information of the node to be checked can include, for example, the component number (Part Number, PN) and the like.

[0077] Figure 4 A schematic diagram of a user display interface according to an embodiment of the present application is shown.

[0078] The user can input a fault phenomenon, for example, input a target server fault phenomenon, which can include "hard disk with location number NO2 cannot be detected, link power supply is abnormal". After determining the troubleshooting steps, the user can be shown the troubleshooting steps and the troubleshooting step execution information corresponding to the troubleshooting steps, for example, the user can be shown that troubleshooting step 1 specifically includes "confirming the on-off state of the hard disk to the mainboard", troubleshooting step 2 specifically includes "confirming the on-off state of the mainboard to the power supply", troubleshooting step 3 specifically includes "first round of confirming whether the hard disk backplane is abnormal", and troubleshooting step 4 specifically includes "second round of confirming whether the hard disk backplane is abnormal". For troubleshooting steps 1-4, the theoretical reference value of the multimeter can be 0.477Ω-1.3Ω. The theoretical reference value of the multimeter, for example, can be used to represent the interval that the resistance measured by the multimeter should meet when the node to be checked is not faulty, and the multimeter settings corresponding to the above four troubleshooting steps can be on-off, red meter pen corresponding to voltage resistance measurement hole, and black meter pen corresponding to common end jack. In addition, the component numbers of the nodes to be checked corresponding to troubleshooting steps 1-4 can also be shown, for example, including component numbers XXX1-XXX6. Further, the component information of the nodes to be checked corresponding to troubleshooting steps 1-4 can also be shown, for example, component model number, etc. Further, the troubleshooting conclusions corresponding to troubleshooting steps 1-4 can also be shown. The troubleshooting conclusion can be obtained according to the user's execution of the troubleshooting steps based on the theoretical position information and the actual position information. For troubleshooting steps 1-3, the troubleshooting conclusion can be "empty", for example, indicating that no abnormalities have been found at this troubleshooting step. At troubleshooting step 4, the troubleshooting conclusion can be determined and shown as "hard disk backplane single fault, replace hard disk backplane".

[0079] According to an embodiment of the present application, the server troubleshooting method further comprises: in response to obtaining the conclusion feedback information of the user on the troubleshooting conclusion, performing an update operation on the preset fault information set based on the conclusion feedback information.

[0080] The conclusion feedback information can include the user's feedback on the troubleshooting conclusion, for example, in the case where the user feedbacks that the troubleshooting conclusion does not correspond to the target server fault phenomenon, such as the case where the target server fault phenomenon still exists after maintenance based on the troubleshooting conclusion. The update operation on the troubleshooting information set can be performed, for example, which can include adjusting or updating the information in the preset fault information set. The troubleshooting conclusion does not correspond to the target server fault phenomenon, which may be because the troubleshooting step node, feedback data value condition, etc. in the preset fault information set may not be accurate enough, so it can be adjusted. Or, it may be because the preset fault information set is missing information corresponding to the target server fault phenomenon, so information corresponding to the target server fault phenomenon can be added.

[0081] By performing the updating operation on the preset fault information set based on the conclusion feedback information, the preset fault information set can be dynamically adjusted based on user feedback, so that the troubleshooting conclusion can be more accurately output subsequently.

[0082] Based on the above server troubleshooting method, the application further provides a server troubleshooting device. The following will be combined with Figure 5 The device will be described in detail.

[0083] Figure 5 The structure block diagram of the server troubleshooting device according to the embodiment of the application is shown.

[0084] As Figure 5 shown, the server troubleshooting device 500 of the embodiment includes a first determination module 510, a second determination module 520, a sending module 530 and an output module 540.

[0085] The first determination module 510 is configured to determine the troubleshooting step corresponding to the target server fault phenomenon input by the user based on the preset fault information set. In an embodiment, the first determination module 510 can be configured to perform the operation S210 described above, and details are not repeated here.

[0086] The second determination module 520 is configured to determine the target to-be-troubleshoot node corresponding to the troubleshooting step, and show the user the theoretical position information of the target to-be-troubleshoot node in the target server. In an embodiment, the second determination module 520 can be configured to perform the operation S220 described above, and details are not repeated here.

[0087] The sending module 530 is configured to send the theoretical position information to the preset positioning device according to the execution order of the troubleshooting step, so that the preset positioning device shows the user the actual position information of the target to-be-troubleshoot node in the target server based on the received theoretical position information. In an embodiment, the sending module 530 can be configured to perform the operation S230 described above, and details are not repeated here.

[0088] The output module 540 is configured to output the troubleshooting conclusion for the target server fault phenomenon based on the troubleshooting feedback information in response to obtaining the troubleshooting feedback information, wherein the troubleshooting feedback information is obtained by the user based on the theoretical position information and the actual position information. In an embodiment, the output module 540 can be configured to perform the operation S240 described above, and details are not repeated here.

[0089] According to the embodiment of the application, the first determination module 510 includes a first determination sub-module, a second determination sub-module and a third determination sub-module.

[0090] The first determining sub-module is configured to determine a target troubleshooting path corresponding to the target server failure phenomenon from a preset failure information set; the second determining sub-module is configured to determine a starting troubleshooting step corresponding to a starting troubleshooting step node in the target troubleshooting path; and the third determining sub-module is configured to determine a subsequent troubleshooting step according to starting feedback data corresponding to the starting troubleshooting step.

[0091] According to an embodiment of the present application, the third determining sub-module comprises a matching unit. The matching unit is configured to match the starting feedback data with a feedback data value condition associated with the starting troubleshooting step node, and determine the subsequent troubleshooting step according to a matching result.

[0092] According to an embodiment of the present application, the server troubleshooting device 500 further comprises a third determining module, a reading module, a fourth determining module, and a visual marking module.

[0093] The third determining module is configured to determine target identification information of a target node to be troubleshooted; the reading module is configured to read target coordinate information and a target rendering rule corresponding to the target node to be troubleshooted from a pre-constructed node mapping relationship set according to the target identification information; the fourth determining module is configured to determine a target model device corresponding to the target node to be troubleshooted in a pre-constructed server model based on the target coordinate information; and the visual marking module is configured to visually mark the target model device based on the target rendering rule.

[0094] According to an embodiment of the present application, the server troubleshooting device 500 further comprises a fifth determining module.

[0095] The fifth determining module is configured to determine and display troubleshooting step execution information corresponding to the troubleshooting step to the user, the troubleshooting step execution information comprising at least one of the following: troubleshooting equipment usage information, and node-to-be-troubleshooted material information.

[0096] According to an embodiment of the present application, the server troubleshooting device 500 further comprises an execution module.

[0097] The execution module is configured to, in response to obtaining conclusion feedback information of the user on the troubleshooting conclusion, perform an update operation on the preset failure information set based on the conclusion feedback information.

[0098] According to embodiments of the present application, any multiple modules among the first determination module 510, the second determination module 520, the sending module 530, and the output module 540 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present application, at least one of the first determination module 510, the second determination module 520, the sending module 530, and the output module 540 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the first determination module 510 , the second determination module 520 , the sending module 530 and the output module 540 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.

[0099] Figure 6 A block diagram of an electronic device suitable for implementing a server fault troubleshooting method according to an embodiment of the present application is shown.

[0100] like Figure 6 As shown, an electronic device 600 according to an embodiment of the present application includes a processor 601, which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 602 or a program loaded from a storage unit 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present application.

[0101] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via the bus 604. The processor 601 performs various operations of the method flow according to the embodiments of the present application by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the programs can also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 can also perform various operations of the method flow according to the embodiments of the present application by executing the programs stored in the one or more memories.

[0102] According to the embodiments of the present application, the electronic device 600 can further include an input / output (I / O) interface 605, which is also connected to the bus 604. The electronic device 600 can further include one or more of the following components connected to the input / output (I / O) interface 605: an input part 606 including a keyboard, a mouse, and the like; an output part 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), and the like, and a speaker, and the like; a storage part 608 including a hard disk, and the like; and a communication part 609 including a network interface card such as a LAN card, a modem, and the like. The communication part 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as necessary. A removable medium 611 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is mounted on the drive 610 as necessary, so that a computer program read therefrom is installed in the storage part 608 as necessary.

[0103] The present application also provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments; or can exist separately without being assembled into the device / apparatus / system. The above computer readable storage medium carries one or more programs, when the one or more programs are executed, the method according to the embodiments of the present application is implemented.

[0104] According to an embodiment of the present application, the computer readable storage medium can be a non-transitory computer readable storage medium, for example, can include but not limited to: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In this application, a computer readable storage medium can be any tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, the computer readable storage medium can include one or more of the above-described ROM 602 and / or RAM 603 and / or memories other than the ROM 602 and the RAM 603.

[0105] Embodiments of the present application also include a computer program product, which includes a computer program containing program codes for executing the methods shown in the flowcharts. When the computer program product is run in a computer system, the program codes are used to make the computer system implement the methods provided by the embodiments of the present application.

[0106] The above-described functions defined in the system / device / apparatus of the embodiments of the present application are performed when the computer program is executed by the processor 601. According to an embodiment of the present application, the above-described system, device, module, unit, etc. can be implemented by computer program modules.

[0107] In one embodiment, the computer program can rely on tangible storage media such as optical storage media, magnetic storage media, etc. In another embodiment, the computer program can also be transmitted, distributed, and downloaded in the form of signals on network media. The computer program is downloaded and installed through the communication part 609 and / or installed from the detachable medium 611. The program codes contained in the computer program can be transmitted by any appropriate network media, including but not limited to wireless, wired, etc., or any suitable combination of the foregoing.

[0108] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609 and / or installed from the detachable medium 611. When the computer program is executed by the processor 601, the above-described functions defined in the system of the embodiments of the present application are performed. According to an embodiment of the present application, the above-described system, device, apparatus, module, unit, etc. can be implemented by computer program modules.

[0109] According to an embodiment of the present application, the program code for executing the computer program provided by the embodiment of the present application can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0110] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0111] It will be understood by those skilled in the art that the features described in the various embodiments of this application may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in this application. In particular, the features described in the various embodiments of this application may be combined and / or coupled in various ways without departing from the spirit and teachings of this application. All such combinations and / or couplings fall within the scope of this application.

[0112] The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present application, those skilled in the art may make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present application.

Claims

1. A server fault troubleshooting method, characterized in that: The method comprises: Determining, based on a preset fault information set, a troubleshooting step corresponding to a target server fault phenomenon input by a user, wherein the preset fault information set includes nodes to be troubleshooted corresponding to the troubleshooting step, and the nodes to be troubleshooted include components of the server; Determine the target node to be checked corresponding to the troubleshooting step; The 3D model of the target node to be checked is presented with a flashing animation; According to the execution order of the troubleshooting steps, the theoretical position information is sent to the preset positioning device, so that the preset positioning device, based on the received theoretical position information, uses a laser light to illuminate the physical components of the server corresponding to the target node to be checked; In response to obtaining troubleshooting feedback information, a troubleshooting conclusion for the target server fault phenomenon is output based on the troubleshooting feedback information, wherein the troubleshooting feedback information is obtained by the user performing the troubleshooting steps based on the theoretical location information and the actual location information.

2. The method according to claim 1, characterized in that The preset fault information set includes a correspondence between a server fault phenomenon and a troubleshooting path, and the troubleshooting path includes at least one troubleshooting step node and a connection sequence between at least one of the troubleshooting step nodes.

3. The method according to claim 2, characterized in that The step of determining, based on the preset fault information set, the fault phenomenon of the target server input by the user and corresponding to the fault includes: Determining a target troubleshooting path corresponding to the target server fault phenomenon from the preset fault information set; Determining a starting troubleshooting step corresponding to a starting troubleshooting step node in the target troubleshooting path; Determine subsequent troubleshooting steps based on initial feedback data corresponding to the initial troubleshooting step.

4. The method according to claim 3, characterized in that The troubleshooting path further includes a feedback data value condition associated with the at least one troubleshooting step node; The determining of subsequent troubleshooting steps based on the initial feedback data corresponding to the initial troubleshooting step includes: The initial feedback data is matched with the feedback data value condition associated with the initial troubleshooting step node, and the subsequent troubleshooting step is determined according to the matching result.

5. The method according to claim 1, wherein The method further comprises: Determine the target identification information of the target node to be checked; Reading target coordinate information and target rendering rules corresponding to the target node to be checked from a pre-built node mapping relationship set according to the target identification information; Determining a target model component corresponding to the target node to be checked in a pre-built server model based on the target coordinate information; The target model device is visually marked based on the target rendering rule.

6. The method according to claim 1, characterized in that The method further comprises: Determine and display to the user the troubleshooting step execution information corresponding to the troubleshooting step, wherein the troubleshooting step execution information includes at least one of the following: troubleshooting equipment usage information and node material information to be troubleshooted.

7. The method according to claim 1, characterized in that The method further comprises: In response to obtaining conclusion feedback information of the user on the fault troubleshooting conclusion, an update operation of the preset fault information set is performed based on the conclusion feedback information.

8. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

9. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Troubleshooting method and device, storage medium and electronic equipment

    CN117389792A