Fault monitoring method and device, electronic equipment and storage medium
By adjusting and monitoring the timeout exit time range of PCIe devices in the server, the problem of fault monitoring of false alarms in the current technology is solved, a more stable and reliable server operation is achieved, and the performance of PCIe devices is optimized.
Patent Information
- Application Number
- CN202411388499.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-09-30
AI Technical Summary
There are false alarm problems in the fault monitoring of the timeout exit mechanism in the existing technology, which affects the normal operation of the system.
By obtaining the first completion timeout exit time range supported by each PCIe device in the server, global modification is made based on the input preset timeout exit time adjustment instruction, the modified second completion timeout exit time range is obtained. After restarting the server, set the timeout exit time for all PCIe devices, and determine the fault monitoring results by comparing the set timeout time range with the original range supported by the device.
It significantly reduces false positives caused by improper configuration of the timeout exit mechanism, ensures the stability and reliability of the server, and provides a flexible mechanism to adapt to changing system requirements and environmental conditions, optimizes the performance of PCIe devices and the overall server operation efficiency.
Smart Images

Figure CN119996252A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a fault monitoring method, device, electronic equipment and storage medium. Background Art
[0002] In the communication process of PCIe devices, when a device called Requester sends a request, it often needs to wait for the response of another device Completer before continuing to perform subsequent operations. However, in some cases, such as configuration errors or system failures, the Requester may not receive a response from the Completer in time. In order to ensure that the system can recover from this waiting state and continue to operate normally, a mechanism is needed to intervene. This is where the Completion Timeout mechanism comes into play: it sets a timeout limit for the Requester to wait for a response. Once this time limit is exceeded, the Requester will no longer wait, but trigger the timeout exit mechanism, allowing the system to continue to perform subsequent tasks and avoid stagnation caused by waiting.
[0003] However, in the related technologies, the timeout exit mechanism may still have problems such as false alarms, which affect the normal operation of the system. Therefore, how to better monitor the faults of the timeout exit mechanism has become an urgent problem to be solved in the industry. Summary of the invention
[0004] The present invention provides a fault monitoring method, device, electronic equipment and storage medium, which are used to solve the defect of how to better perform fault monitoring of a timeout exit mechanism in the prior art.
[0005] The present invention provides a fault monitoring method, comprising the following steps: Get the first completion timeout exit time range supported by each PCIE device in the server; According to the input preset timeout exit time adjustment instruction, each of the first completion timeout exit time ranges is globally modified to obtain a modified second completion timeout exit time range; wherein the preset timeout exit time adjustment instruction includes an adjustment method for multiple preset completion timeout exit times; Restarting the server, and setting a completion timeout exit time for all PCIE devices in the server according to the second completion timeout exit time range; The fault monitoring result is determined according to a comparison result of the second completion timeout exit time range corresponding to the PCIE device after the restart and the first completion timeout exit time range supported by each PCIE device.
[0006] According to a fault monitoring method provided by the present invention, the step of obtaining the first completion timeout exit time range supported by each PCIE device in the server includes: Obtain device information of all PCIE devices in the server; According to the preset register command, obtain a first completion timeout exit time range supported by the PCIE device; Among them, the first completion timeout exit time range includes at least one of the following: a first preset time range [a1, b1], a second preset time range [c1, d1], a third preset time range [e1, f1], a fourth preset time range [g1, h1], a1≤b1≤c1≤d1≤f1≤g1≤h1.
[0007] According to a fault monitoring method provided by the present invention, the globally modifying each of the first completion timeout exit time ranges according to the input preset timeout exit time adjustment instruction to obtain the modified second completion timeout exit time range includes: Exporting basic input and output system options of the server through a server configuration tool; Through the basic input and output system option search and the PCIE device global timeout setting, the first preset time range [a1, b1] in each of the first completion timeout exit time ranges is modified to the fifth preset time range [a2, b1], the second preset time range [c1, d1] is modified to the sixth preset time range [c2, d1], the third preset time range [e1, f1] is modified to the seventh preset time range [e2, f1], and the fourth preset time range [g1, h1] is modified to the eighth preset time range [g2, h1]; each modified second completion timeout exit time range is obtained; wherein the second completion timeout exit time range includes at least one of the following: the fifth preset time range [a2, b1], the sixth preset time range [c2, d1], the seventh preset time range [e2, f1], and the eighth preset time range [g2, h1]; b1>a2>a1, d1>c2>c1, f1>e2>e1, h1>g2>g1; Executing the import instruction of the second completion timeout exit time range through the server configuration tool to obtain a modified basic input and output system option; The modified basic input / output system option is used to set a global completion timeout exit time according to the second completion timeout exit time range after the server is restarted.
[0008] According to a fault monitoring method provided by the present invention, the server is restarted, and a completion timeout exit time is set for all PCIE devices in the server according to the second completion timeout exit time range, including: The server is restarted, and after the basic input and output system of the system performs system parameter initialization, the corresponding second completion timeout exit time range is loaded for the PCIE device to implement the timeout exit time setting for all PCIE devices in the server.
[0009] According to a fault monitoring method provided by the present invention, before the step of determining the fault monitoring result according to the comparison result of the second completion timeout exit time range corresponding to the PCIE device after the restart and the first completion timeout exit time range supported by each PCIE device, the method further includes: After restarting the server, querying the target value of the second completion timeout exit time range of the register representing the register in the PCIE through a device query instruction; Wherein, the target value includes at least one of the following: a first target value of the fifth preset time range [a2, b1], a second target value of the sixth preset time range [c2, d1], a third target value of the seventh preset time range [e2, f1], and a fourth target value of the eighth preset time range [g2, h1].
[0010] According to a fault monitoring method provided by the present invention, the fault monitoring result is determined according to the comparison result between the second completion timeout exit time range corresponding to the PCIE device after the restart and the first completion timeout exit time range supported by each PCIE device, including: In the case where the first completion timeout exit time range includes the second completion timeout exit time range, the register check passes, and a completion timeout exit time fault monitoring log of the PCIE device where the register is located is generated.
[0011] According to a fault monitoring method provided by the present invention, the fault monitoring result is determined according to the comparison result between the second completion timeout exit time range corresponding to the PCIE device after the restart and the first completion timeout exit time range supported by each PCIE device, including: In a case where the second completion timeout exit time range is not included in the first completion timeout exit time range, the register check fails, and basic input / output system configuration fault information of the PCIE device where the register is located is generated.
[0012] The present invention also provides a fault monitoring device, comprising the following modules: An acquisition module is used to acquire the first completion timeout exit time range supported by each PCIE device in the server; A modification module, used to globally modify each of the first completion timeout exit time ranges according to the input preset timeout exit time adjustment instruction, so as to obtain a modified second completion timeout exit time range; wherein the preset timeout exit time adjustment instruction includes an adjustment method for multiple preset completion timeout exit times; A restart module, used to restart the server, and set a completion timeout exit time for all PCIE devices in the server according to the second completion timeout exit time range; The monitoring module is used to determine the fault monitoring result according to the comparison result between the second completion timeout exit time range corresponding to the PCIE device after restart and the first completion timeout exit time range supported by each PCIE device.
[0013] According to the fault monitoring device provided by the present invention, the device is also used for: Obtain device information of all PCIE devices in the server; According to the preset register command, obtain a first completion timeout exit time range supported by the PCIE device; Among them, the first completion timeout exit time range includes at least one of the following: a first preset time range [a1, b1], a second preset time range [c1, d1], a third preset time range [e1, f1], a fourth preset time range [g1, h1], a1≤b1≤c1≤d1≤f1≤g1≤h1.
[0014] According to the fault monitoring device provided by the present invention, the device is also used for: Exporting basic input and output system options of the server through a server configuration tool; Through the basic input and output system option search and the PCIE device global timeout setting, the first preset time range [a1, b1] in each of the first completion timeout exit time ranges is modified to the fifth preset time range [a2, b1], the second preset time range [c1, d1] is modified to the sixth preset time range [c2, d1], the third preset time range [e1, f1] is modified to the seventh preset time range [e2, f1], and the fourth preset time range [g1, h1] is modified to the eighth preset time range [g2, h1]; each modified second completion timeout exit time range is obtained; wherein the second completion timeout exit time range includes at least one of the following: the fifth preset time range [a2, b1], the sixth preset time range [c2, d1], the seventh preset time range [e2, f1], and the eighth preset time range [g2, h1]; b1>a2>a1, d1>c2>c1, f1>e2>e1, h1>g2>g1; Executing the import instruction of the second completion timeout exit time range through the server configuration tool to obtain a modified basic input and output system option; The modified basic input / output system option is used to set a global completion timeout exit time according to the second completion timeout exit time range after the server is restarted.
[0015] According to the fault monitoring device provided by the present invention, the device is also used for: The server is restarted, and after the basic input and output system of the system performs system parameter initialization, the corresponding second completion timeout exit time range is loaded for the PCIE device to implement the timeout exit time setting for all PCIE devices in the server.
[0016] According to the fault monitoring device provided by the present invention, the device is also used for: After restarting the server, querying the target value of the second completion timeout exit time range of the register representing the register in the PCIE through a device query instruction; Wherein, the target value includes at least one of the following: a first target value of the fifth preset time range [a2, b1], a second target value of the sixth preset time range [c2, d1], a third target value of the seventh preset time range [e2, f1], and a fourth target value of the eighth preset time range [g2, h1].
[0017] According to the fault monitoring device provided by the present invention, the device is also used for: In the case where the first completion timeout exit time range includes the second completion timeout exit time range, the register check passes, and a completion timeout exit time fault monitoring log of the PCIE device where the register is located is generated.
[0018] According to the fault monitoring device provided by the present invention, the device is also used for: In a case where the second completion timeout exit time range is not included in the first completion timeout exit time range, the register check fails, and basic input / output system configuration fault information of the PCIE device where the register is located is generated.
[0019] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-mentioned fault monitoring methods is implemented.
[0020] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the fault monitoring method described in any one of the above methods is implemented.
[0021] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the fault monitoring method described above is implemented.
[0022] The fault monitoring method, device, electronic device and storage medium provided by the present invention, the system will obtain the first completion timeout exit time range supported by each PCIE device in the server. According to the input preset timeout exit time adjustment instruction, the timeout range supported by each device is globally modified to obtain the modified second completion timeout exit time range. After restarting the server, according to the modified timeout range, the completion timeout exit time is set for all PCIE devices in the server. Then the system will determine the fault monitoring result according to the comparison result of the second completion timeout exit time range corresponding to each register in the PCIE device and the first completion timeout exit time range supported by the device. If the completion timeout time set by the device is not within the supported range, there may be a fault and further investigation is required. False alarms caused by improper configuration of the timeout exit mechanism can be significantly reduced to ensure the stability and reliability of the server. At the same time, it also provides a flexible mechanism to adapt to changing system requirements and environmental conditions, thereby optimizing the performance of PCIe devices and the operating efficiency of the overall server. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0024] Figure 1 It is a flow chart of the fault monitoring method provided by the present invention.
[0025] Figure 2 A query flow chart provided for an embodiment of the present application.
[0026] Figure 3 A schematic diagram of the verification process provided for an embodiment of the present application.
[0027] Figure 4 A schematic diagram of the structure of a fault monitoring device provided in an embodiment of the present application; Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0029] Figure 1 It is a flow chart of the fault monitoring method provided by the present invention, such as Figure 1 As shown, the method includes the following: Step 110, obtaining the first completion timeout exit time range supported by each PCIE device in the server; In an embodiment of the present application, each PCIe device has a series of configuration registers that contain the device's functional and performance parameters, including the supported completion timeout exit time range.
[0030] Use system management tools or access directly through the PCIe configuration space to read the values of these registers. For example, use the lspci command to list all PCIe devices and their related information, including the supported timeout ranges.
[0031] Parse the supported timeout ranges from the device configuration registers. This information is usually in the form of bit fields and needs to be parsed according to the PCIe specification.
[0032] Assume that there are two PCIE devices in the server. Device A supports a timeout range of 50us-10ms, and device B supports a timeout range of 10ms-250ms.
[0033] Step 120, according to the input preset timeout exit time adjustment instruction, globally modify each of the first completion timeout exit time ranges to obtain a modified second completion timeout exit time range; wherein the preset timeout exit time adjustment instruction includes an adjustment method for multiple preset completion timeout exit times; In the embodiment of the present application, the timeout range and adjustment method to be adjusted are understood. The instruction may include increasing or decreasing the timeout time, setting a specific timeout threshold, etc.
[0034] Modify the first completion timeout exit time range for all PCIe devices in the server. This usually needs to be done through the BIOS / firmware configuration interface or using system management tools.
[0035] Use command-line tools (such as lspci with scripts) or graphical tools to batch update device timeout settings. Apply the preset timeout adjustment to the configuration of each device. This may involve writing new values to the configuration registers of the PCIe devices.
[0036] Record the modified second completion timeout exit time range for each device. These records can be used for future reference and further adjustments.
[0037] Step 130, restarting the server, and setting a completion timeout exit time for all PCIE devices in the server according to the second completion timeout exit time range; In an embodiment of the present application, a restart operation of the server is performed. This can be done through a physical restart button, a system command (such as reboot in Linux or remote restart through a BMC tool).
[0038] During system startup, the BIOS will load the updated configuration, including the settings for the second completion timeout exit time range. These settings take effect when the BIOS initializes the PCIe device.
[0039] Observe the server boot process to ensure that the system boots smoothly and that no new hardware errors or configuration issues have occurred.
[0040] Once the server boots up, use system management tools or command-line tools such as lspci to check the configuration of each PCIe device to verify that the new timeout settings have been correctly applied.
[0041] Step 140, determining a fault monitoring result according to a comparison result between the second completion timeout exit time range corresponding to the PCIE device after restart and the first completion timeout exit time range supported by each PCIE device.
[0042] After the server restarts and loads the new timeout settings, use system management tools or command-line tools such as lspci to collect configuration register data for each PCIe device.
[0043] From the collected register data, the current timeout exit time range of each device is extracted, that is, the second completion timeout exit time range, and the second completion timeout exit time range of each device is compared with the original first completion timeout exit time range.
[0044] Confirm whether the modified timeout setting is within the preset support range of the device. Compare the results and record them. If the modified timeout setting is within the support range, record it as normal; if it is not within the support range, record it as abnormal. Generate a fault monitoring report based on the comparison results. The report should list the original timeout range, the modified timeout range, and the comparison results for each device in detail.
[0045] In an embodiment of the present application, the system obtains the first completion timeout exit time range supported by each PCIE device in the server. According to the input preset timeout exit time adjustment instruction, the timeout range supported by each device is globally modified to obtain the modified second completion timeout exit time range. After restarting the server, the completion timeout exit time is set for all PCIE devices in the server according to the modified timeout range. Then the system determines the fault monitoring result based on the comparison result of the second completion timeout exit time range corresponding to each register in the PCIE device and the first completion timeout exit time range supported by the device. If the completion timeout time set by the device is not within the supported range, there may be a fault and further investigation is required. False alarms caused by improper configuration of the timeout exit mechanism can be significantly reduced to ensure the stability and reliability of the server. At the same time, it also provides a flexible mechanism to adapt to changing system requirements and environmental conditions, thereby optimizing the performance of PCIe devices and the operating efficiency of the overall server.
[0046] Optionally, the acquiring of the first completion timeout exit time range supported by each PCIE device in the server includes: Obtain device information of all PCIE devices in the server; According to the preset register command, obtain a first completion timeout exit time range supported by the PCIE device; The first completion timeout exit time range includes at least one of the following: a first preset time range [a1, b1], a second preset time range [c1, d1], a third preset time range [e1, f1], and a fourth preset time range [g1, h1].
[0047] In the embodiment of the present application, the bus / device / function of all currently inserted PCIE devices can be obtained through lspci -tv in the system, and the CompletionTimeout Range supported by the current PCIE device itself can be viewed through the register command. The execution command is as follows: lspci -s bus:device.function –xxx (obtained from bus:device.function), check the value of bit[0-3] of offset 64 is X, and check the supported range of PCI-E Completion Timeout (CTO) based on the value of X
[0048] 0001b:RangA; 0010b:RangB0011b:RangA and RangB
[0049] 0110b:RangB and RangC0111b:RangA,RankB and RangC
[0050] 1110b:RangB,RangCand RangD1111b:RangA,RangB,RangC and RangD
[0051] In an optional embodiment, the first preset time range [a1, b1] is: 50us-10ms, the second preset time range [c1, d1] is 10ms-250ms, the third preset time range [e1, f1] is 250ms-4s, and the fourth preset time range is 4s-64s.
[0052] Taking the value of X as 0111 as an example, this device itself supports the first preset time range [a1, b1], the second preset time range [c1, d1], the third preset time range [e1, f1], and the fourth preset time range [g1, h1]. You can view that the Completion Timeout value supported by the current device is 50us-4s.
[0053] Figure 2 The query flow chart provided in the embodiment of the present application is as follows: Figure 2 As shown, including: First, PCIE device enumeration is performed, then IO initialization is performed, and then PCIEBAR register polling is performed to check the PCIEDevice Control 2 Register enable status; When the fourth preset time range [g1, h1] is supported, CTO is set to 4s-64s; when the third preset time range [e1, f1] is supported, CTO is set to 250ms-4s; when the second preset time range [c1, d1] is supported, CTO is set to 10ms-250ms; when the first preset time range [a1, b1] is supported, CTO is set to 50usms-10ms.
[0054] In the embodiment of the present application, the configuration register of the PCIE device contains the timeout range information supported by the device. By reading this register, the completion timeout range supported by the device can be obtained. By obtaining the timeout range supported by the device, the characteristics and performance of the device can be understood, providing a basis for subsequent configuration and optimization.
[0055] Optionally, the globally modifying each of the first completion timeout exit time ranges according to the input preset timeout exit time adjustment instruction to obtain the modified second completion timeout exit time range includes: Exporting basic input and output system options of the server through a server configuration tool; Through the basic input and output system option search and the PCIE device global timeout setting, the first preset time range [a1, b1] in each of the first completion timeout exit time ranges is modified to the fifth preset time range [a2, b1], the second preset time range [c1, d1] is modified to the sixth preset time range [c2, d1], the third preset time range [e1, f1] is modified to the seventh preset time range [e2, f1], and the fourth preset time range [g1, h1] is modified to the eighth preset time range [g2, h1]; each modified second completion timeout exit time range is obtained; wherein the second completion timeout exit time range includes at least one of the following: the fifth preset time range [a2, b1], the sixth preset time range [c2, d1], the seventh preset time range [e2, f1], and the eighth preset time range [g2, h1]; b1>a2>a1, d1>c2>c1, f1>e2>e1, h1>g2>g1; Executing the import instruction of the second completion timeout exit time range through the server configuration tool to obtain a modified basic input and output system option; The modified basic input / output system option is used to set a global completion timeout exit time according to the second completion timeout exit time range after the server is restarted.
[0056] In the embodiment of the present application, a server configuration tool is used to export the basic input and output system (BIOS) options of the current server. This step usually involves using a BIOS interface or using a specific export command.
[0057] For example, export the bios options via the SCE tool . / SCELnx64 / o / s bios.txt / b.
[0058] Find and set Enable PCI-E Completion Timeout (Global) to Global (Global is the master switch) in bios.txt, set PCI-E Global Timeout Value to 50us to 100us / 1msto 10ms / 16ms to 55ms / 65ms to 210ms / 260ms to 900ms / 1s to 3.5s / 4s to 13s, and save the changes.
[0059] Specifically, through the basic input and output system option search and the PCIE device global timeout setting, the first preset time range [a1, b1] in each of the first completion timeout exit time ranges is modified to the fifth preset time range [a2, b1], the second preset time range [c1, d1] is modified to the sixth preset time range [c2, d1], the third preset time range [e1, f1] is modified to the seventh preset time range [e2, f1], and the fourth preset time range [g1, h1] is modified to the eighth preset time range [g2, h1]; each modified second completion timeout exit time range is obtained; wherein the second completion timeout exit time range includes at least one of the following: the fifth preset time range [a2, b1], the sixth preset time range [c2, d1], the seventh preset time range [e2, f1], and the eighth preset time range [g2, h1]; b1>a2>a1, d1>c2>c1, f1>e2>e1, h1>g2>g1; For example: Modify the data of 50 microseconds (us) to 100 microseconds. Modify the data of 1 millisecond (ms) to 10 milliseconds. Modify the data of 6 milliseconds to 55 milliseconds. Modify the data of 65 milliseconds to 210 milliseconds. Modify the data of 260 milliseconds to 900 milliseconds. Modify the data of 1 second (s) to 3.5 seconds. Modify the data of 4 seconds to 13 seconds. These modifications will generate a new second completion timeout exit time range.
[0060] Then, import this modification through the SCE tool and execute the following command: . / SCELnx64 / i / s bios.txt / b to get the modified basic input and output system options.
[0061] In the embodiment of the present application, by adjusting the timeout period, false alarms caused by improper timeout settings can be reduced. Appropriate timeout settings can improve the overall stability of the system and avoid system failures caused by waiting timeouts.
[0062] Optionally, restarting the server, and setting a completion timeout exit time for all PCIE devices in the server according to the second completion timeout exit time range, includes: The server is restarted, and after the basic input and output system of the system performs system parameter initialization, the corresponding second completion timeout exit time range is loaded for the PCIE device to implement the timeout exit time setting for all PCIE devices in the server.
[0063] In the embodiment of the present application, the server restart is triggered by a system command, a physical restart button, or a remote management tool (such as IPMI). During the server startup process, the BIOS (Basic Input Output System) will perform a self-test and initialize system parameters.
[0064] The BIOS loads a new timeout exit time configuration for each PCIe device based on the imported second completion timeout exit time range setting. As the PCIe device is initialized, the new timeout setting is applied, overwriting the previous default or original setting.
[0065] Once the server boots up, use system management tools or command-line tools such as lspci to check the configuration of each PCIe device to verify that the new timeout settings have been correctly applied.
[0066] Optionally, before the step of determining the fault monitoring result according to the comparison result of the second completion timeout exit time range corresponding to the PCIE device after the restart and the first completion timeout exit time range supported by each PCIE device, the step further includes: After restarting the server, querying the target value of the second completion timeout exit time range of the register representing the register in the PCIE through a device query instruction; Among them, the first target value of the fifth preset time range [a2, b1], the second target value of the sixth preset time range [c2, d1], the third target value of the seventh preset time range [e2, f1], and the fourth target value of the eighth preset time range [g2, h1].
[0067] In an embodiment of the present application, after restarting the server, the target value of bit [0-3] of the PCIE device register offset 68 is queried to determine the second completion timeout exit time range of each register.
[0068] Restart the server to make the modified BIOS option take effect. Then, query the target value of bit[0-3] of PCIE device register offset 68 through the device query command. This value represents the current completion timeout exit time range of the device.
[0069] According to the queried target value, determine the corresponding completion timeout exit time range. The corresponding relationship between the target value and the timeout range is as follows: 0001: 50us to 100us
[0070] 0010: 1ms to 10ms
[0071] 0101: 16ms to 55ms
[0072] 0110: 65ms to 210ms
[0073] 1001: 260ms to 900ms
[0074] 1010: 1s to 3.5s
[0075] 0000: 4s to 13s
[0076] In an embodiment of the present application, querying the register value can verify whether the modified BIOS option is set correctly and determine the current completion timeout exit time range of the device.
[0077] Optionally, determining the fault monitoring result according to a comparison result between the second completion timeout exit time range corresponding to the PCIE device after the restart and the first completion timeout exit time range supported by each PCIE device includes: In the case where the first completion timeout exit time range includes the second completion timeout exit time range, the register check passes, and a completion timeout exit time fault monitoring log of the PCIE device where the register is located is generated.
[0078] In a case where the second completion timeout exit time range is not included in the first completion timeout exit time range, the register check fails, and basic input / output system configuration fault information of the PCIE device where the register is located is generated.
[0079] Figure 3 The verification process diagram provided in the embodiment of the present application is as follows: Figure 3As shown, the system compares the second completion timeout exit time range corresponding to the PCIE device register queried after restart with the first completion timeout exit time range supported by the device previously obtained.
[0080] If the second completion timeout exit time range is included in the first completion timeout exit time range, it means that the register setting is correct and the device can work normally within the timeout range it supports. Therefore, the register check passes, and the system will generate a completion timeout exit time fault monitoring log for the PCIE device to record the normal status.
[0081] If the second completion timeout exit time range is not included in the first completion timeout exit time range, this indicates that the register setting exceeds the range supported by the device, which may be a configuration error or device failure. In this case, the register check fails and the system will generate a basic input and output system (BIOS) configuration fault message, indicating that there is a problem with the configuration of the PCIE device.
[0082] The system will record detailed information about the PCIE device where the fault occurred, including device ID, register settings, supported timeout ranges, and actual set timeout ranges. The system may issue an alarm to notify maintenance personnel that a configuration fault has occurred and needs to be checked and repaired. Fault information will be recorded in the system's log file for subsequent analysis and troubleshooting.
[0083] In the embodiment of the present application, the system can promptly discover and report configuration problems of PCIE devices, helping operation and maintenance personnel to quickly locate and solve problems, thereby ensuring the stable operation of the system.
[0084] In an optional embodiment, a PCIe device is installed in the server, and necessary BIOS configuration is performed to ensure compatibility and performance optimization of the hardware and system settings.
[0085] Install the PCIe Completion Timeout Adaptive Device, which is a hardware or software solution designed to dynamically adjust the PCIe device timeout settings. After completing the installation, power on the server and boot the operating system.
[0086] After the system is started, the PCIe Completion Timeout adaptive program is automatically triggered to run. This program is designed to monitor all PCIe devices in the server and can track their status and performance indicators in real time.
[0087] If the monitoring program detects any PCIe device failure or timeout anomaly, it will automatically generate a detailed fault matching log.
[0088] These logs record the time, device identification, fault type, and possible cause of the fault event in detail and save them in the result.log file.
[0089] The device automatically sends the logs in the result.log file back to the central monitoring system or the analysis platform of the operation and maintenance team.
[0090] After receiving these logs, the operation and maintenance personnel use them to accurately locate the faulty device and perform in-depth analysis to determine the root cause of the failure.
[0091] Based on the results of log analysis, the operations team develops and executes appropriate troubleshooting measures, which may include adjusting timeout settings, hardware repair or replacement, driver updates, etc.
[0092] After troubleshooting is complete, continue to monitor device status to verify that the issue is resolved and to ensure that the server and PCIe devices return to normal operation.
[0093] In an embodiment of the present application, through this refined monitoring and response process, the data center can achieve rapid response to PCIe device failures, minimize service interruptions, and improve overall system stability and reliability.
[0094] The fault monitoring device provided by the present invention is described below. The fault monitoring device described below and the fault monitoring method described above can be referenced to each other.
[0095] Figure 4 A schematic diagram of the structure of a fault monitoring device provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, including: The acquisition module 410 is used to acquire the first completion timeout exit time range supported by each PCIE device in the server; The modification module 420 is used to globally modify each of the first completion timeout exit time ranges according to the input preset timeout exit time adjustment instruction to obtain a modified second completion timeout exit time range; wherein the preset timeout exit time adjustment instruction includes an adjustment method for multiple preset completion timeout exit times; The restart module 430 is used to restart the server and set the completion timeout exit time for all PCIE devices in the server according to the second completion timeout exit time range; The monitoring module 440 is used to determine the fault monitoring result according to the comparison result between the second completion timeout exit time range corresponding to the PCIE device after the restart and the first completion timeout exit time range supported by each PCIE device.
[0096] According to the fault monitoring device provided by the present invention, the device is also used for: Obtain device information of all PCIE devices in the server; According to the preset register command, obtain a first completion timeout exit time range supported by the PCIE device; Among them, the first completion timeout exit time range includes at least one of the following: a first preset time range [a1, b1], a second preset time range [c1, d1], a third preset time range [e1, f1], and a fourth preset time range [g1, h1], a1≤b1≤c1≤d1≤f1≤g1≤h1.
[0097] According to the fault monitoring device provided by the present invention, the device is also used for: Exporting basic input and output system options of the server through a server configuration tool; Through the basic input and output system option search and the PCIE device global timeout setting, the first preset time range [a1, b1] in each of the first completion timeout exit time ranges is modified to the fifth preset time range [a2, b1], the second preset time range [c1, d1] is modified to the sixth preset time range [c2, d1], the third preset time range [e1, f1] is modified to the seventh preset time range [e2, f1], and the fourth preset time range [g1, h1] is modified to the eighth preset time range [g2, h1]; each modified second completion timeout exit time range is obtained; wherein the second completion timeout exit time range includes at least one of the following: the fifth preset time range [a2, b1], the sixth preset time range [c2, d1], the seventh preset time range [e2, f1], and the eighth preset time range [g2, h1]; b1>a2>a1, d1>c2>c1, f1>e2>e1, h1>g2>g1; executing the import instruction of the second completion timeout exit time range through the server configuration tool to obtain the modified basic input and output system options; The modified basic input / output system option is used to set a global completion timeout exit time according to the second completion timeout exit time range after the server is restarted.
[0098] According to the fault monitoring device provided by the present invention, the device is also used for: The server is restarted, and after the basic input and output system of the system performs system parameter initialization, the corresponding second completion timeout exit time range is loaded for the PCIE device to implement the timeout exit time setting for all PCIE devices in the server.
[0099] According to the fault monitoring device provided by the present invention, the device is also used for: After restarting the server, querying the second completion timeout of the register in the PCIE through the device query instruction to obtain a target value of the second completion timeout exit time range of the register representing the register; Among them, the first target value of the fifth preset time range [a2, b1], the second target value of the sixth preset time range [c2, d1], the third target value of the seventh preset time range [e2, f1], and the fourth target value of the eighth preset time range [g2, h1].
[0100] According to the fault monitoring device provided by the present invention, the device is also used for: In the case where the first completion timeout exit time range includes the second completion timeout exit time range, the register check passes, and a completion timeout exit time fault monitoring log of the PCIE device where the register is located is generated.
[0101] According to the fault monitoring device provided by the present invention, the device is also used for: In a case where the second completion timeout exit time range is not included in the first completion timeout exit time range, the register check fails, and basic input / output system configuration fault information of the PCIE device where the register is located is generated.
[0102] In an embodiment of the present application, the system obtains the first completion timeout exit time range supported by each PCIE device in the server. According to the input preset timeout exit time adjustment instruction, the timeout range supported by each device is globally modified to obtain the modified second completion timeout exit time range. After restarting the server, the completion timeout exit time is set for all PCIE devices in the server according to the modified timeout range. Then the system determines the fault monitoring result based on the comparison result of the second completion timeout exit time range corresponding to each register in the PCIE device and the first completion timeout exit time range supported by the device. If the completion timeout time set by the device is not within the supported range, there may be a fault and further investigation is required. False alarms caused by improper configuration of the timeout exit mechanism can be significantly reduced to ensure the stability and reliability of the server. At the same time, it also provides a flexible mechanism to adapt to changing system requirements and environmental conditions, thereby optimizing the performance of PCIe devices and the operating efficiency of the overall server.
[0103] Figure 5 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 5As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute the fault monitoring method, which includes: obtaining the first completion timeout exit time range supported by each PCIE device in the server; According to the input preset timeout exit time adjustment instruction, each of the first completion timeout exit time ranges is globally modified to obtain a modified second completion timeout exit time range; wherein the preset timeout exit time adjustment instruction includes an adjustment method for multiple preset completion timeout exit times; Restarting the server, and setting a completion timeout exit time for all PCIE devices in the server according to the second completion timeout exit time range; The fault monitoring result is determined according to a comparison result of the second completion timeout exit time range corresponding to the PCIE device after the restart and the first completion timeout exit time range supported by each PCIE device.
[0104] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0105] On the other hand, the present invention further provides a computer program product, the computer program product includes a computer program, the computer program can be stored in a non-transitory computer-readable storage medium, when the computer program is executed by a processor, the computer can execute the fault monitoring method provided by the above methods, the method comprising: obtaining a first completion timeout exit time range supported by each PCIE device in the server; According to the input preset timeout exit time adjustment instruction, each of the first completion timeout exit time ranges is globally modified to obtain a modified second completion timeout exit time range; wherein the preset timeout exit time adjustment instruction includes an adjustment method for multiple preset completion timeout exit times; Restarting the server, and setting a completion timeout exit time for all PCIE devices in the server according to the second completion timeout exit time range; The fault monitoring result is determined according to a comparison result of the second completion timeout exit time range corresponding to the PCIE device after the restart and the first completion timeout exit time range supported by each PCIE device.
[0106] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to execute the fault monitoring method provided by the above methods, the method comprising: obtaining a first completion timeout exit time range supported by each PCIE device in the server; According to the input preset timeout exit time adjustment instruction, each of the first completion timeout exit time ranges is globally modified to obtain a modified second completion timeout exit time range; wherein the preset timeout exit time adjustment instruction includes an adjustment method for multiple preset completion timeout exit times; Restarting the server, and setting a completion timeout exit time for all PCIE devices in the server according to the second completion timeout exit time range; The fault monitoring result is determined according to a comparison result of the second completion timeout exit time range corresponding to the PCIE device after the restart and the first completion timeout exit time range supported by each PCIE device.
[0107] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0108] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A fault monitoring method, characterized in that: include: Get the first completion timeout exit time range supported by each PCIE device in the server; According to the input preset timeout exit time adjustment instruction, each of the first completion timeout exit time ranges is globally modified to obtain a modified second completion timeout exit time range; wherein the preset timeout exit time adjustment instruction includes an adjustment method for multiple preset completion timeout exit times; Restarting the server, and setting a completion timeout exit time for all PCIE devices in the server according to the second completion timeout exit time range; The fault monitoring result is determined according to a comparison result of the second completion timeout exit time range corresponding to the PCIE device after the restart and the first completion timeout exit time range supported by each PCIE device.
2. The fault monitoring method according to claim 1, characterized in that: The obtaining of the first completion timeout exit time range supported by each PCIE device in the server includes: Obtain device information of all PCIE devices in the server; According to the preset register command, obtain a first completion timeout exit time range supported by the PCIE device; Among them, the first completion timeout exit time range includes at least one of the following: a first preset time range [a1, b1], a second preset time range [c1, d1], a third preset time range [e1, f1], a fourth preset time range [g1, h1], a1≤b1≤c1≤d1≤f1≤g1≤h1.
3. The fault monitoring method according to claim 2, characterized in that: The globally modifying each of the first completion timeout exit time ranges according to the input preset timeout exit time adjustment instruction to obtain the modified second completion timeout exit time range includes: Exporting basic input and output system options of the server through a server configuration tool; Through the basic input and output system option search and the PCIE device global timeout setting, the first preset time range [a1, b1] in each of the first completion timeout exit time ranges is modified to the fifth preset time range [a2, b1], the second preset time range [c1, d1] is modified to the sixth preset time range [c2, d1], the third preset time range [e1, f1] is modified to the seventh preset time range [e2, f1], and the fourth preset time range [g1, h1] is modified to the eighth preset time range [g2, h1]; each modified second completion timeout exit time range is obtained; wherein the second completion timeout exit time range includes at least one of the following: the fifth preset time range [a2, b1], the sixth preset time range [c2, d1], the seventh preset time range [e2, f1], and the eighth preset time range [g2, h1]; b1>a2>a1, d1>c2>c1, f1>e2>e1, h1>g2>g1; Executing the import instruction of the second completion timeout exit time range through the server configuration tool to obtain a modified basic input and output system option; The modified basic input / output system option is used to set a global completion timeout exit time according to the second completion timeout exit time range after the server is restarted.
4. The fault monitoring method according to claim 1, characterized in that: Restarting the server, and setting a completion timeout exit time for all PCIE devices in the server according to the second completion timeout exit time range, including: The server is restarted, and after the basic input and output system of the system performs system parameter initialization, the corresponding second completion timeout exit time range is loaded for the PCIE device to implement the timeout exit time setting for all PCIE devices in the server.
5. The fault monitoring method according to claim 1, characterized in that: Before the step of determining the fault monitoring result according to the comparison result of the second completion timeout exit time range corresponding to the PCIE device after the restart and the first completion timeout exit time range supported by each PCIE device, the step further includes: After restarting the server, querying the target value of the second completion timeout exit time range of the register representing the register in the PCIE through a device query instruction; Wherein, the target value includes at least one of the following: a first target value of the fifth preset time range [a2, b1], a second target value of the sixth preset time range [c2, d1], a third target value of the seventh preset time range [e2, f1], and a fourth target value of the eighth preset time range [g2, h1].
6. The fault monitoring method according to claim 5, characterized in that: The determining of the fault monitoring result according to the comparison result between the second completion timeout exit time range corresponding to the PCIE device after the restart and the first completion timeout exit time range supported by each PCIE device includes: In the case where the first completion timeout exit time range includes the second completion timeout exit time range, the register check passes, and a completion timeout exit time fault monitoring log of the PCIE device where the register is located is generated.
7. The fault monitoring method according to claim 5, characterized in that: The determining of the fault monitoring result according to the comparison result between the second completion timeout exit time range corresponding to the PCIE device after the restart and the first completion timeout exit time range supported by each PCIE device includes: In a case where the second completion timeout exit time range is not included in the first completion timeout exit time range, the register check fails, and basic input / output system configuration fault information of the PCIE device where the register is located is generated.
8. A fault monitoring device, characterized in that: include: An acquisition module is used to acquire the first completion timeout exit time range supported by each PCIE device in the server; A modification module, used to globally modify each of the first completion timeout exit time ranges according to the input preset timeout exit time adjustment instruction, so as to obtain a modified second completion timeout exit time range; wherein the preset timeout exit time adjustment instruction includes an adjustment method for multiple preset completion timeout exit times; A restart module, used to restart the server and set a completion timeout exit time for all PCIE devices in the server according to the second completion timeout exit time range; The monitoring module is used to determine the fault monitoring result according to the comparison result between the second completion timeout exit time range corresponding to the PCIE device after restart and the first completion timeout exit time range supported by each PCIE device.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the fault monitoring method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the fault monitoring method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
PCIe link state detection method of server and server
CN118245295A
Pcie fault self-repairing method, apparatus and device, and readable storage medium
WO2022228499A1