Fault detection system and method, storage medium, electronic equipment and program product

By embedding an exception positioning program in the system management interrupt program, detecting the processor's enable registers and count registers, and analyzing software information, the problem of difficult positioning of system management interrupt failure is solved, and fast and accurate fault detection is achieved.

CN120276906AActive Publication Date: 2025-07-08INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510494928.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-08
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

In the process of triggering system management interrupts, software functions cannot be successfully executed, resulting in the inability to accurately detect the fault location, which increases the difficulty and time cost of R&D personnel for locating faults.

Method used

It provides a fault detection system, by embedding an exception positioning program in the program for system management interrupts, detecting the hardware results of the enable register and count register of the processor in turn, and detecting the software information of the basic input and output system chip when there is no hardware exception, and generating fault location information of the system management interrupts.

Benefits of technology

The ability to quickly and accurately locate the fault location of system management interrupts, reduces the manpower and time investment of R&D personnel, and improves the efficiency and accuracy of fault detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276906A_ABST
    Figure CN120276906A_ABST
Patent Text Reader

Abstract

The invention discloses a fault detection system and method, a storage medium, electronic equipment and a program product, and relates to the technical field of computers.The system comprises a system disk, a processor and a basic input and output system chip, and the processor is connected with the system disk and the basic input and output system chip; in response to the exception of the processor in the process of triggering the system management interruption, calling an exception positioning program in a system disk by the processor; the processor performs fault detection on an enabling register and a counting register in the processor in sequence based on the exception positioning program to obtain a hardware detection result of the processor; under the condition that the processor judges that the hardware detection result is not abnormal, the processor performs fault detection on software information of system management interruption in the basic input and output system chip to obtain a software detection result; and the processor generates fault position information of the system management interruption according to the hardware detection result and the software detection result. According to the invention, abnormal position information is determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a fault detection system, method, storage medium, electronic device, and program product. Background Art

[0002] A System Management Interrupt (SMI) is the only way for a processor to enter the System Management Mode (SMM). The system management interrupt is triggered by a system management interrupt signal received through a system management interrupt pin on the server processor or through an Advanced Programmable Interrupt Controller Bus (APIC Bus), causing the CPU to enter the SMM state. Among them, the system management interrupt is a non-maskable external interrupt.

[0003] Currently, in the process of triggering a system management interrupt in related technologies, there may be a situation where a software function cannot be executed successfully. However, since it involves two parts, namely the triggering of the system management interrupt and the installation of the system management interrupt, it is impossible to accurately detect the fault location of the system management interrupt. Summary of the Invention

[0004] The present disclosure provides a fault detection system, method, storage medium, electronic device, and program product. Its main purpose is to solve the problem that in the process of triggering a system management interrupt in related technologies, there may be a situation where a software function cannot be executed successfully. However, since it involves two parts, namely the triggering of the system management interrupt and the installation of the system management interrupt, it is impossible to accurately detect the fault location of the system management interrupt.

[0005] In a first aspect, the present application provides a fault detection system, including: a system disk, a processor, and a Basic Input / Output System (BIOS) chip. The processor is respectively connected to the system disk and the BIOS chip; In response to an exception occurring during the process of the processor triggering a system management interrupt, the processor calls an exception location program in the system disk, and the exception location program is embedded in the program of the system management interrupt; The processor sequentially performs fault detection on an enable register and a count register in the processor based on the exception location program to obtain a hardware detection result of the processor; When the processor determines that there is no exception in the hardware detection result, the processor performs fault detection on software information of the system management interrupt in the BIOS chip to obtain a software detection result; The processor generates fault location information of the system management interrupt based on the hardware detection result and the software detection result.

[0006] In a second aspect, the present application provides a fault detection method, including: In response to an exception occurring during the processor triggering a system management interrupt, an exception location program is called, and the exception location program is embedded in the program of the system management interrupt; Based on the exception location program, fault detection is sequentially performed on the enable register and the count register in the processor to obtain the hardware detection result of the processor; In the case where it is determined that there is no exception in the hardware detection result, fault detection is performed on the software information of the system management interrupt in the basic input / output system chip to obtain the software detection result; According to the hardware detection result and the software detection result, fault location information of the system management interrupt is generated.

[0007] In a third aspect, the present application provides a fault detection device, including: A calling module, configured to call an exception location program in response to an exception occurring during the processor triggering a system management interrupt, and the exception location program is embedded in the program of the system management interrupt; A detection module, configured to perform fault detection on the enable register and the count register in the processor based on the exception location program to obtain the hardware detection result of the processor; The detection module is further configured to perform fault detection on the software information of the system management interrupt in the basic input / output system chip in the case where it is determined that there is no exception in the hardware detection result to obtain the software detection result; A generation module, configured to generate fault location information of the system management interrupt according to the hardware detection result and the software detection result.

[0008] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0009] In a fifth aspect, the present application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the computer program, the method of the first aspect is implemented.

[0010] In a sixth aspect, the present application provides a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0011] The fault detection system, method, storage medium, electronic device and program product provided by the present disclosure, wherein the system includes: a system disk, a processor and a basic input / output system chip, and the processor is respectively connected to the system disk and the basic input / output system chip; in response to an exception occurring during the processor triggering a system management interrupt, the processor calls an exception location program in the system disk, and the exception location program is embedded in the program of the system management interrupt; the processor sequentially performs fault detection on the enable register and the count register in the processor based on the exception location program to obtain the hardware detection result of the processor; when the processor determines that there is no exception in the hardware detection result, it performs fault detection on the software information of the system management interrupt in the basic input / output system chip to obtain the software detection result; the processor generates the fault location information of the system management interrupt according to the hardware detection result and the software detection result. Compared with the related art, the present application can, in the case of an exception in the system management interrupt, call the exception location program pre-embedded in the program of the system management interrupt, and the exception location program can detect the enable register and the count register in the processor to determine whether there is an exception in the hardware part of the processor, which in turn causes the exception in the system management interrupt; it can also detect the software information of the system management interrupt to analyze whether there is an exception in the program code of the system management interrupt, which in turn causes the exception in the system management interrupt, and determine the location information that causes the exception in the system management interrupt according to the hardware detection result and the software detection result, that is, the enable register, the count register or the software program information. In addition, it can also reduce the manpower and time investment required for R & D personnel to determine the fault location after the system management interrupt exception.

[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0014] Figure 1 FIG. 1 shows a schematic structural diagram of a fault detection system provided by an embodiment of the present application; Figure 2 FIG. 2 shows a schematic flow diagram of a fault detection method provided by an embodiment of the present application; Figure 3 FIG. 3 shows a schematic flow diagram of another fault detection method provided by an embodiment of the present application; Figure 4Shows a schematic diagram of an example provided by an embodiment of the present application; Figure 1 Among them: 1 - System disk; 2 - Processor; 3 - Basic Input Output System chip. Specific implementation manners

[0015] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0016] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0017] The Basic Input Output System (BIOS), as the manager of the most basic and direct hardware settings and controls on the server motherboard, can provide more simple and user-friendly functions for the server. BIOS is a bridge connecting hardware devices and software programs. During the boot process, BIOS first performs self-check and initialization on the hardware devices. The self-check process includes self-checking the central processing unit (CPU), memory, motherboard, serial and parallel ports, and hard disk devices, and loading the driver programs for Peripheral Component Interconnect Express (PCIe) devices, input / output (I / O) devices, etc. After the detection of the hardware devices is completed, the startup item will be loaded and started to the operating system for the user to use the operating system to complete the operations required by the user.

[0018] The SMI is the only way for the processor to enter the System Management Mode (SMM). The SMI is an SMI signal that can be received through the SMI# pin on the server processor or through the Advanced Programmable Interrupt Controller Bus (APIC Bus). The SMI is a non-maskable external interrupt. After the SMI is issued, the CPU enters the SMM state. The SMM is an execution mode of a CPU in the x86 architecture by Intel and can only be entered through the SMI. The SMM is a special-purpose operation mode that provides functions for power management, system hardware control, or the execution of code by the Original Equipment Manufacturer Handler (OEM handler). The SMM can provide an independent and easily isolated processor environment that is transparent to the operating system or other software. When the SMI is generated, the processor will wait for all instructions to be executed and all memory operations to be completed, and then will switch to a special operating environment defined by a new address space. The SMI Handler will execute in this special environment. The key code and data of the SMI execution program are located in a physical memory area of this space, the System Management RAM (SMRAM). The processor will also save the current state to the SMRAM and then start executing the SMI Handler.

[0019] Therefore, after the machine boots into the operating system, some system events can be processed by triggering the SMI, or hardware control can be achieved, such as by implementing some special functions or virtual hardware functions, or used to avoid some hardware bugs (BUGs). The SMI is also often used on server systems to implement some special functions. After a SMI function is executed, the program that was being executed before the SMI trigger saved by the CPU will continue.

[0020] SMI is divided into software-triggered system management interrupt (Software SMI) and hardware-triggered system management interrupt (Hardware SMI) according to different triggering methods. For Software SMI, by writing the corresponding function number into the input / output port (IO Port) 0xB2, the function corresponding to this number can be triggered. As the name implies, Software SMI realizes the SMI function by combining software code, and the combined software code is installed and executed in the BIOS. The overall implementation steps of SoftWare SMI are as follows: set the function function that the user needs to implement after triggering SMI at the BIOS end, install a function at the BIOS end, and use this function to implement the function that the user needs to implement. However, this function will be installed during the process of the server starting up and executing the BIOS, but it will only be installed and will not be executed at this time. Only when the corresponding SMI is triggered will this function be executed to realize the function. Multiple Software SMI functions can be installed in the BIOS code, and these SoftWare SMI functions can be distinguished by a defined number. And multiple Software SMI functions installed under BISO will be connected to a software system management interrupt processor list (SW SMI Handler List) through functions. When a Software SMI is triggered under the OS, according to the corresponding function number written in the IO port 0xB2 by the triggered SoftWare SMI under the OS, it will enter the general entry function of the SW SMI at the BIOS end. In the entry function, find the function corresponding to the function number under the SofteWare SMI handler List according to certain rules, and then execute this function to realize the user-defined function.

[0021] Implement a user - specific function through Software SMI, which mainly includes the following processes: 1. First, install a Software SMI function with a specific function number on the BIOS side; 2. Under the OS, trigger the Software SMI by writing the corresponding function number to the IO port 0xB2 through an execution program or tool. Therefore, the implementation of the entire Software SMI function is divided into two parts: function installation on the BIOS side and SMI triggering under the OS. And regarding the characteristics of the CPU's SMM state, all interrupts and SMIs are masked in the SMM state, the SMI handler is non - reentrant, and only when the current SMI execution is completed will the next SMI program be executed. If another SMI comes during the execution of the current SMI, only one SMI will be latched. After the current SMI is executed, the latched SMI will be executed, and other SMIs will be ignored.

[0022] However, during the normal SMI function triggering process, there are factors that can cause the Software function to fail to execute successfully. Since it involves two parts, namely SMI triggering and SMI installation, and these two parts are on different carriers, it hinders the debugging process of R & D personnel. Usually, it is necessary for BIOS R & D and OS R & D to jointly complete the location of which party the problem lies in, and then the party with the problem solves the problem. Therefore, the initial location of the problem will consume a large amount of manpower and time, resulting in serious waste of resources.

[0023] In order to improve the technical problem that in the triggering process of system management interrupts in related technologies, there will be a situation where the software function cannot be executed successfully, but due to involving two parts, namely the triggering of system management interrupts and the installation of system management interrupts, it is impossible to accurately detect the fault location of system management interrupts. This embodiment provides a fault detection system, as Figure 1 shown. The system includes: a system disk 1, a processor 2, and a basic input / output system chip 3. The processor 2 is respectively connected to the system disk 1 and the basic input / output system chip 3; in response to an exception occurring during the processor 2 triggering a system management interrupt, the processor 2 calls an exception location program in the system disk 1, and the exception location program is embedded in the system management interrupt program; the processor 2 sequentially performs fault detection on the enable register and the count register in the processor 2 based on the exception location program to obtain the hardware detection result of the processor 2; when the processor 2 determines that there is no exception in the hardware detection result, it performs fault detection on the software information of the system management interrupt in the basic input / output system chip 3 to obtain the software detection result; the processor 2 generates the fault location information of the system management interrupt based on the hardware detection result and the software detection result.

[0024] In an embodiment of the present application, the system disk refers to a storage device that installs an operating system and its related files. It not only carries the core components of the operating system, such as the kernel, drivers, and services, but may also store application programs, user data, and various configuration files. Exemplarily, an exception location program is stored in the system disk 1 in the embodiment of the present application. The fault location information can be determined through the exception location program when the processor 2 triggers a system management interrupt.

[0025] Optionally, this embodiment provides a fault detection method, as Figure 2 shown, which is applied to the processor in the above-mentioned fault detection system. The method includes the following steps: Step 101: In response to an exception occurring during the processor triggering a system management interrupt, call the exception location program.

[0026] Among them, the exception location program is embedded in the program of the system management interrupt.

[0027] In an embodiment of the present application, the processor triggering a system management interrupt (SMI) is a special hardware interrupt mechanism, usually triggered by the motherboard chipset or the basic input / output system (BIOS) to perform low-level system management tasks. The priority of SMI is higher than that of all other interrupts (including non-maskable interrupt (NMI) and ordinary interrupts), and it runs in the system management mode (SMM), independent of the operating system.

[0028] In some examples, the ways to trigger a system management interrupt may include but are not limited to: 1. Chipset trigger: The motherboard chipset notifies the processor to trigger SMI through a specific signal (such as writing to I / O port 0xB2). 2. Timer trigger: The hardware timer can periodically trigger SMI to regularly check the system status. 3. External event trigger: Events such as pressing the power button, heat sensor alarm, etc. 4. BIOS / firmware: The BIOS or firmware can trigger SMI through specific instructions or operations. 5. Operating system: The underlying driver or tool triggers SMI by writing to a specific I / O port (such as 0xB2).

[0029] For this embodiment, an exception occurring during the process of triggering the system management interrupt may cause system instability, crashes, or other unpredictable behaviors. Specifically, the reasons for the exception may include, but are not limited to: Hardware level: 1. Motherboard or chipset failure: If there is a failure in the motherboard chipset or related hardware, it may cause the SMI to not be triggered or processed correctly. 2. Power supply problem: Unstable power supply interferes with the execution of the SMI. 3. Overheating or voltage anomaly: High temperature or voltage fluctuations trigger hardware errors. Firmware / BIOS level: 1. BIOS configuration error: Settings in the BIOS (such as power management policies) cause the SMI to be abnormal. 2. BIOS firmware error: If there are errors in the BIOS or firmware, it will affect the execution of the SMI handler. 3. Incompatible firmware version: Using a BIOS version that does not match the hardware. Software level: 1. Incorrect function number: Writing an invalid function number to an I / O port (such as 0xB2) causes the chipset to not be able to process it correctly. 2. Driver conflict: The underlying driver interferes with the triggering or execution of the SMI. 3. Operating system problem: Modules in the operating system conflict with the SMI handler. System Management RAM (SMRAM) related problems: 1. SMRAM memory corruption: If the data in the SMRAM is accidentally modified or corrupted, it causes the SMI handler to not run properly. 2. SMRAM address conflict: If other software or hardware occupies the address space of the SMRAM, it causes the SMI to fail.

[0030] It should be noted that the exception location program implemented in this application is embedded in the program of the system management interrupt. Specifically, a program tool that runs under an operating system (OS), that is, the exception location program in this embodiment of the application, can be set, and this program tool is embedded into the program process triggered by the system management interrupt (SMI). When there is a problem that the software system management interrupt (SW SMI) cannot be executed normally, this program tool, that is, the exception location program in this embodiment of the application, is called.

[0031] Step 102: Based on the exception location program, perform fault detection on the enable register and the count register in the processor in sequence to obtain the hardware detection result of the processor.

[0032] In an embodiment of the present application, the enable register can be the Software System Management Interrupt Enable Register (SMI_EN) in the processor, which is used to control and enable the System Management Interrupt (SMI). SMI is a special hardware interrupt that allows the operating system or firmware to perform low-level system management tasks, such as power management, thermal monitoring, etc. The priority of SMI is higher than all other types of interrupts, including the Non-Maskable Interrupt (NMI), and it runs under the independent System Management Mode (SMM).

[0033] Exemplarily, the functions of the SMI_EN register, which is the enable register in the embodiment of the present application, can include but are not limited to: 1. Enable / disable the SMI enable bit; when a specific enable bit is set, a specific type of SMI can be allowed to occur. Disable bit: Clearing the corresponding enable bit can prohibit some types of SMI, thus avoiding unnecessary system management operations. 2. Configure the SMI trigger condition; Different sources of SMI: The SMI_EN register can be configured to control which hardware events or software commands trigger the SMI. For example, an alarm for too high temperature, a low battery state, or other hardware monitoring signals. Software trigger: Software can manually trigger the SMI by writing a specific value to an I / O port or a memory-mapped register.

[0034] For this embodiment, the use of the SMI_EN register, which is the enable register in the embodiment of the present application, can include but are not limited to: 1. Access the SMI_EN register; Permission requirements: Modifying the SMI_EN register usually requires a high level of permission because it involves underlying operations of the system. This can generally only be done by the BIOS or the Unified Extensible Firmware Interface (UEFI), or an operating system module with the corresponding permissions. Location: The specific location of the SMI_EN register depends on the processor architecture and motherboard design, and it is located within the I / O space or the memory-mapped address space. 2. Read and modify; Read the current state: To understand which types of SMI are currently enabled, the content of the SMI_EN register needs to be read first. Modify the enable state: Modify the specific bit in the register according to the requirements. For example, if you want to enable the SMI related to a certain function, set the corresponding bit to 1; if you want to disable it, set it to 0.

[0035] It should be noted that by detecting the enable register, it can be determined whether the System Management Interrupt (SMI) has been temporarily disabled due to the loading of other software under the operating system (OS).

[0036] Exemplarily, in an operating system (OS), some software does cause the system management interrupt (SMI) to be temporarily disabled. The following are some reasons and scenarios, including but not limited to: 1. Operating system kernel control over SMI; Reason: The operating system kernel module temporarily disables SMI to avoid interfering with critical tasks or improving performance. Scenario: Tasks with high real-time requirements: In a real-time operating system (RTOS) or tasks that require low latency, the triggering of SMI introduces unpredictable latency. Therefore, the operating system temporarily disables SMI. Performance optimization: Frequent triggering of SMI affects CPU performance, so SMI is disabled. 2. Driver or firmware operations; Reason: The driver or underlying firmware modifies the SMI control register (such as SMI_EN), thereby temporarily disabling SMI. Scenario: Device initialization: When initializing a hardware device, the driver disables SMI to prevent hardware conflicts. Debug mode: When debugging hardware or firmware, SMI is disabled to better capture the system state. 3. Security-related software; Reason: Security software disables SMI to prevent malicious code from using SMI for attacks. Scenario: Preventing Rootkit attacks: The high-privilege feature of SMI makes it a potential security threat. Security software disables SMI to prevent malicious code from running in SMM. Virtualization environment: In a virtualization environment, the host disables the SMI of the guest to ensure isolation and security. 4. BIOS / UEFI settings; Reason: BIOS / UEFI settings affect the behavior of SMI, and these settings are dynamically adjusted by the operating system or software. Scenario: Power management strategy: The power-saving strategy disables SMI to reduce unnecessary interruptions. Advanced configuration: Advanced users or administrators can adjust the BIOS / UEFI settings through software tools, thereby affecting the enabled state of SMI.

[0037] For this embodiment, the count register can be a register in the processor that counts the number of system management interrupt triggers (SMI_COUNT). Specifically, the SMI_COUNT register, which is the count register in the embodiment of the present application, is a hardware register dedicated to recording the number of SMI triggers. Whenever SMI is triggered, the value of this register automatically increments. By reading the value of this register, the trigger frequency and number of SMI in the system can be understood.

[0038] In some examples, the main uses of the count register may include, but are not limited to: 1. Performance monitoring: Frequently triggered SMIs can have a negative impact on system performance (such as increasing latency). By monitoring SMI_COUNT, it is possible to analyze whether the trigger frequency of SMIs is too high. 2. Troubleshooting: If the system exhibits anomalies (such as high latency, unresponsiveness, etc.), checking SMI_COUNT can determine whether it is related to the frequent triggering of SMIs. 3. Debugging and optimization: During the development or debugging phase, engineers can use SMI_COUNT to determine which operations or events cause SMIs to be triggered, thereby optimizing the system design. 4. Security analysis: The high-privilege nature of SMIs makes them a target for attacks. By monitoring SMI_COUNT, potential security issues can be discovered.

[0039] Exemplarily, the working principle of SMI_COUNT, which is the count register in the embodiments of the present application, may include, but are not limited to: 1. Automatic increment mechanism: Each time an SMI is triggered, the processor automatically increments the value of the SMI_COUNT register by 1. This register is typically a read-only register, and users cannot directly modify its value. 2. Access method: Register location: SMI_COUNT is usually located in a specific I / O address space or memory-mapped area, and the specific location depends on the processor architecture and motherboard design. Reading method: The value of SMI_COUNT can be read through assembly language or a low-level driver. 3. Overflow handling: If the bit width of the SMI_COUNT register is limited (such as 32 bits or 64 bits), when the count value reaches the maximum, an overflow occurs. It will be cleared and continue counting when overflow occurs, or additional flag bits are provided to indicate the overflow status.

[0040] Exemplarily, the application scenarios of SMI_COUNT may include, but are not limited to: 1. Performance tuning: If the value of SMI_COUNT grows too fast, it indicates that hardware events (such as power management, thermal monitoring) frequently trigger SMIs. At this time, the BIOS settings can be adjusted or the hardware configuration can be optimized. 2. Fault diagnosis: When the system exhibits performance problems or abnormal behaviors, by comparing the SMI_COUNT values at different time periods, it can be determined whether there are excessive SMI triggers. 3. Security auditing: If the growth rate of SMI_COUNT is abnormal, it indicates that there may be malware attempting to use SMIs for attacks. Combined with other security tools, potential threats can be further analyzed. 4. BIOS / firmware development: During the development or debugging of BIOS / firmware, SMI_COUNT is an important debugging tool that can help developers verify the behavior of SMI handlers.

[0041] It should be noted that by detecting the enable counter, it can be determined whether there is any other system management interrupt function (SMI function) occupying the system management interrupt (SMI) channel, resulting in the failure of the software system management interrupt function (SWSMI function) to execute successfully.

[0042] Step 103: When it is determined that the hardware detection result is normal, perform a fault detection on the software information of the system management interrupt in the basic input / output system chip to obtain a software detection result.

[0043] In the embodiment of the present application, if it is determined that there is no abnormality through the fault detection of the enable register and the count register in the processor, then it is necessary to perform a fault detection on the software information of the system management interrupt. Specifically, it can be to detect whether there are errors in the program code of the system management interrupt.

[0044] It should be noted that by performing a fault detection on the software information, it can be determined whether the software system management interrupt function (SW SMI function) fails to execute normally when it reaches the basic input / output system (BIOS) end, and whether it is caused by the system management interrupt function (SMI function) under the basic input / output system (BIOS).

[0045] Exemplarily, under the basic input / output system (BIOS), if the system management interrupt function (SMI function) causes a system management interrupt exception, this will affect the normal operation of the system and lead to system crashes, performance degradation, or other unexpected behaviors. The specific reasons may include, but are not limited to: 1. BIOS firmware problems; BIOS version incompatibility: Using a BIOS version that does not match the hardware will cause the SMI function to fail to execute correctly. Firmware error: There are errors in the BIOS or related firmware, affecting the triggering or processing process of the SMI. 2. Hardware failures; Motherboard or chipset problems: If there are hardware failures in the motherboard or chipset, it will interfere with the normal operation of the SMI. Unstable power supply: An unstable power supply causes hardware errors, thus triggering an SMI exception. Overheating or voltage fluctuations: High temperature or voltage fluctuations also cause hardware errors, resulting in abnormal SMI processing. 3. Configuration errors; Incorrect BIOS settings: BIOS settings (such as power management policies) cause SMI exceptions. Conflicting device settings: The settings of the device conflict with the SMI function, resulting in an exception. 4. Improper triggering of external events; Frequent triggering of SMI: If the hardware or software frequently triggers the SMI, it will exceed the system processing capacity, resulting in an exception. External hardware signal problems: External events such as pressing the power button or a thermal sensor alarm trigger SMI exceptions.

[0046] Step 104: Generate the fault location information of the system management interrupt based on the hardware detection result and the software detection result.

[0047] It should be noted that in the embodiments of the present application, the fault location information of the system management interrupt can be generated based on the hardware detection result and / or the software detection result. For example, during the hardware detection process, if an abnormality is detected in the enable register, the fault location information of the system management interrupt can be generated based on the detection result of the enable register; if an abnormality is detected in the count register, the fault location information of the system management interrupt can be generated based on the detection result of the count register; during the software detection process, if an abnormality is detected in the software information, the fault location information of the system management interrupt can be generated based on the software detection result.

[0048] Compared with the related art, in this embodiment, when the system management interrupt is abnormal, an exception location program pre-embedded in the program of the system management interrupt can be called. Through the exception location program, the enable register and the count register in the processor can be detected to determine whether there is an abnormality in the hardware part of the processor, which in turn causes the system management interrupt to be abnormal; the software information of the system management interrupt can also be detected to analyze whether there is an abnormality in the program code of the system management interrupt, which in turn causes the system management interrupt to be abnormal. According to the results of the hardware detection and the software detection, the location information that causes the system management interrupt to be abnormal is determined, that is, the enable register, the count register, or the software program information. In addition, the manpower and time investment required for R & D personnel to determine the fault location after the system management interrupt is abnormal can be reduced.

[0049] Further, as a refinement and extension of the above embodiment, the embodiments of the present application provide a fault detection method, as Figure 3 shown, the method includes: Step 201: In response to an exception occurring during the process of the processor triggering the system management interrupt, call the exception location program.

[0050] Among them, the exception location program is embedded in the program of the system management interrupt.

[0051] In the embodiments of the present application, an exception location program can be embedded in the program of the system management interrupt. A program tool running under an operating system (OS), that is, the exception location program in the embodiments of the present application, can be set, and this program tool can be embedded into the program process triggered by the system management interrupt (SMI). When a problem occurs that the software system management interrupt (SW SMI) cannot be executed normally, this program tool, that is, the exception location program in the embodiments of the present application, is called.

[0052] Step 202: Based on the exception location program, perform a fault detection on the enable register to obtain a first detection result.

[0053] Optionally, step 202 may specifically include: reading the enable register in the exception location program to determine the byte information corresponding to the enable register; determining whether the loading of the software in the processor affects the system management interrupt according to the byte information, and determining a first detection result based on the determination result.

[0054] In the embodiments of the present application, the byte composition of the enable register may include: A typical enable register is usually an 8-bit, 16-bit or 32-bit register. Each bit represents the enabling state of a specific function or module, where 1 represents Enable and 0 represents Disable. The functions of the byte information may include: Function control: Each bit controls a specific function or module. Configuration flexibility: By setting different bit combinations, multiple functions can be flexibly enabled or disabled.

[0055] Optionally, when performing "reading the enable register in the exception location program to determine the byte information corresponding to the enable register", it may include but is not limited to: reading the enable register in the exception location program to determine the target byte corresponding to the enable register, and detecting whether the target byte is set; if it is determined that the target byte is set, generating first byte information indicating that there is no setting exception for the target byte; if it is determined that the target byte is not set, generating second byte information indicating that there is a setting exception for the target byte.

[0056] In some examples, the current value of the enable register is read through I / O operations or memory mapping to understand which functions are enabled or disabled. According to the read value of the enable register, the status of each bit is analyzed to determine which types of SMIs have been enabled or disabled. Exemplarily, if byte (Bit) 2 (software SMI enable) is 1, it means that software-triggered SMI is allowed. If Bit2 is 0, it means that software-triggered SMI is prohibited.

[0057] In the embodiments of the present application, if the target bytes are byte (Bit) 0 and byte (Bit) 5, the software system management interrupt control and enable register (SMI_EN) in the processor may be read to parse byte (Bit) 0 and byte (Bit) 5, and check whether byte (Bit) 0 and byte (Bit) 5 are set. If byte (Bit) 0 and byte (Bit) 5 are set, first byte information indicating that there is no setting exception for the target byte needs to be generated; on the contrary, if byte (Bit) 0 and byte (Bit) 5 are not set, second byte information indicating that there is a setting exception for the target byte is generated.

[0058] Note that in the software system management interrupt control and enable register (SMI_EN) in the processor, Bit0 is a reserved bit and is usually not used. Bit5 is the hardware event SMI enable, which controls whether to allow hardware events to trigger SMI.

[0059] For this embodiment, Bit0: Function: Reserved bit, usually not used. Default value: Usually 0, indicating that this bit does not enable any function. Bit5: Function: Controls whether to allow hardware events to trigger SMI. If Bit5 = 1, then hardware events are allowed to trigger SMI. If Bit5 = 0, then hardware events are prohibited from triggering SMI. Use: Used to manage SMI triggered by hardware signals (such as power button press, over-temperature alarm, etc.).

[0060] Exemplarily, by performing a bitwise AND operation (&) to check whether a specific bit is 1. Assume that the value read from the SMI_EN register is 0x21 (binary 00100001), then: Bit0: Is set (value is 1), indicating that the reserved bit is set. Bit5: Is set (value is 1), indicating that the hardware event SMI is enabled.

[0061] Optionally, when executing "judging whether the loading of software in the processor affects the system management interrupt based on the byte information and determining the first detection result based on the judgment result", it may include but is not limited to: If it is determined based on the second byte information that the loading affects the system management interrupt, then the first detection result is determined to be that there is an abnormality in the enable register.

[0062] In the embodiment of the present application, if the generated second byte information, that is, the target byte is not set, it can be determined that the loading affects the system management interrupt. Then, it can be judged that there is an abnormality in the enable register, and the abnormal position information can be generated based on the enable register.

[0063] Exemplarily, if the target bytes are byte (Bit) 0 and byte (Bit) 5 and the second byte information is generated, it can be judged that there is software loading under the operating system (OS) that temporarily disables the system management interrupt (SMI).

[0064] Optionally, the method of this embodiment further includes: When the first detection result is determined to be that there is an abnormality in the enable register, setting the target byte in the enable register.

[0065] In the embodiment of the present application, setting a byte bit (Bit) (setting it to 1) is usually achieved through a bitwise OR operator. The function of the bitwise OR operator is to perform a logical OR operation on each bit of the two operands. For example, if any one bit is 1, the result bit is 1. If both bits are 0, the result bit is 0.

[0066] In some examples, the steps of setting a byte may include but are not limited to: 1. Determine the bits to be set: Determine the bit numbers to be set (counting from 0). Construct a mask where the bits to be set are 1 and the rest are 0. 2. Perform a bitwise OR operation: Perform a bitwise OR operation on the original value and the mask to obtain a new byte value. 3. Write back to the register or store: If operating on a hardware register, write the new value back to the register.

[0067] Step 203: When it is determined that the first detection result is normal, perform a fault detection on the count register based on the exception location program to obtain a second detection result.

[0068] Wherein, the first detection result and the second detection result constitute the hardware detection result.

[0069] Optionally, step 203 may specifically include: When it is determined that the enable register in the first detection result is normal, trigger the system management interrupt again and determine the response of the count register; Based on the re - trigger situation and the response situation, determine whether the system management interrupt channel in the processor is occupied, and determine the second detection result based on the judgment result.

[0070] In the embodiments of the present application, triggering the system management interrupt again when it is determined that the enable register in the first detection result is normal means triggering the system management interrupt again when the system management interrupt is abnormal.

[0071] In some examples, the scenarios of triggering the SMI again may include but are not limited to: 1. Debugging and diagnosis: If there is an error in the SMI handler, more context information can be captured by triggering the SMI again for debugging. For example, in embedded system or BIOS development, the register status or memory content can be checked by re - triggering the SMI. 2. Fault recovery: It is necessary to trigger the SMI again to attempt to restore the system to a normal state. For example, when the SMI exception is caused by a temporary hardware error (such as signal interference), re - triggering the SMI helps to restore normal operation. 3. Testing and verification: During the development or testing phase, multiple SMI triggers need to be simulated to verify the stability and robustness of the system.

[0072] For this embodiment, the method for triggering the SMI again may include but is not limited to: 1. Software-triggered SMI: Most processors support manually triggering the SMI by writing to a specific I / O port (such as 0xB2 or 0x92). 2. Hardware-triggered SMI: Hardware events (such as power button press, over-temperature alarm, etc.) can automatically trigger the SMI. In a debugging environment, the SMI can be triggered by simulating hardware events. 3. Using debugging tools: Using hardware debugging tools (such as JTAG interface) to forcibly trigger the SMI. The debugging tools can directly intervene in the behavior of the processor and are suitable for use in the development and testing phases.

[0073] Optionally, when performing "triggering the system management interrupt again and determining the response of the count register", it may include but is not limited to: writing the target function number to the target port in the input / output port, triggering the system management interrupt again, and reading the response value of the count register.

[0074] Optionally, when performing "judging whether the system management interrupt channel in the processor is occupied based on the re-triggering situation and the response situation, and determining the second detection result based on the judgment result", it may include but is not limited to: judging whether there are consecutive system management interrupts in the processor based on the re-triggering situation and the response value; if it is determined that there are consecutive system management interrupts in the processor, it is determined that the system management interrupt channel is occupied, and the second detection result is determined as an abnormality in the count register.

[0075] In the embodiment of the present application, trigger the system management interrupt (SMI) again by writing the function number to the I / O port with the port number 0xB2 (IO Port 0xB2), to see if it can be successfully triggered, and read the value of the register (SMI_COUNT) that counts the number of system management interrupt triggers in the processor, to see if there are dense system management interrupts (SMIs) generated, so as to exclude whether there is any other system management interrupt function (SMI function) occupying the system management interrupt (SMI) channel, resulting in the failure of the software system management interrupt function (SW SMI function) to be successfully executed.

[0076] In some examples, a dense system management interrupt refers to a system management interrupt that is frequently triggered within a short period of time. Due to the high priority of the SMI and its operation in the system management mode, dense SMIs will have a significant impact on the system's performance, stability, and response time.

[0077] For this embodiment, the reasons for generating intensive system management interrupts may include but are not limited to: 1. Hardware failures; Power issues: Unstable power supply or frequent voltage fluctuations frequently trigger SMIs related to power management. Temperature monitoring: If the system has poor heat dissipation, too high temperature will frequently trigger heat monitoring SMIs. Hardware errors: Hardware failures such as bus errors and memory errors also frequently trigger SMIs. 2. BIOS / firmware configurations; Unreasonable settings: BIOS settings (such as overly sensitive power management policies) cause SMIs to be frequently triggered. Bugs or design flaws: Bugs in the BIOS or firmware cause unnecessary SMI triggers. 3. Software problems; Driver conflicts: Underlying drivers interfere with the normal processing of SMIs, resulting in frequent triggers. Operating system modules: Modules in the operating system (such as power management services) frequently request SMIs. 4. Malware Rootkit attacks: Malware takes advantage of the high-privilege characteristics of SMIs for covert operations, resulting in frequent triggers.

[0078] In the embodiment of the present application, SMI_COUNT is a register that records the number of SMI triggers. By reading this register, the trigger frequency of SMIs can be understood; specifically, if the system management interrupt can be triggered again and the response value is low, it can be determined that there are no consecutive multiple system management interrupts in the processor.

[0079] Exemplarily, if the system management interrupt cannot be triggered again or the response value is high, it can be determined that there are consecutive multiple system management interrupts in the processor.

[0080] Optionally, the method of this embodiment further includes: when the second detection result is determined to be that there is an abnormality in the count register, repeatedly trigger the system management interrupt.

[0081] Optionally, after performing "judging whether the system management interrupt channel in the processor is occupied based on the re-triggering situation and the response situation, and determining the second detection result based on the judgment result", the method of the embodiment further includes: if it is determined that there are no consecutive multiple system management interrupts in the processor, it is determined that the system management interrupt channel is not occupied, and the second detection result is determined to be that there is no abnormality in the count register.

[0082] In some examples, in the system management interrupt (SMI) channel, there is a situation where multiple system management interrupt functions (SMI functions) share the same channel. If a certain SMI function occupies the channel or resources, it will cause other SMI functions (such as the software system management interrupt function SW SMI function) to be unable to execute successfully.

[0083] In the embodiments of the present application, the reasons for SMI channel conflicts may include, but are not limited to: 1. Sharing the SMI handler; SMI is a high-priority interrupt triggered by hardware or requested by software, and all SMIs will enter the same system management mode (SMM). In SMM, there is usually a unified SMI handler to distribute and process different SMI functions (such as power management, thermal monitoring, debugging, etc.). If an SMI function occupies the SMI handler resources for a long time (for example, it does not exit correctly or enters an infinite loop), other SMI functions cannot be processed. (2) Resource competition; SMI functions require exclusive access to specific hardware resources (such as memory mapped areas, registers, etc.). If these resources are occupied, other SMI functions will fail. For example, SW SMI depends on registers or memory buffers, and these resources are occupied by other SMI functions. 3. Priority issues; SMI functions have a higher priority, resulting in lower-priority SMI functions (such as SW SMI) being delayed or ignored. For example, hardware-triggered SMIs (such as over-temperature alarms) take precedence over software-triggered SMIs. 4. Abnormal states; If an exception occurs during the execution of an SMI function (such as a stack overflow, address error, etc.), the entire SMI handler will crash, thereby affecting other SMI functions.

[0084] For this embodiment, the methods for detecting whether other SMI functions are occupying the channel may include, but are not limited to: 1. Checking the SMI trigger frequency; Use the SMI_COUNT register or other logging mechanisms to count the number and type of SMI triggers. If it is found that the SMI trigger frequency is abnormally high, it means that an SMI function is frequently occupying the channel. 2. Analyzing the log information; Embed a logging function in the SMI handler to record the reason and context information for each SMI trigger. 3. Using debugging tools; Use hardware debugging tools (such as JTAG or logic analyzers) to monitor the SMI trigger signal. Check for multiple consecutive SMI trigger events and the trigger source. 4. Checking resource occupancy; Check whether the resources related to SMI (such as registers, memory buffers) are occupied.

[0085] Step 204, in the case where it is determined that the hardware detection result is normal, perform a fault detection on the software information of the system management interrupt in the basic input / output system chip to obtain a software detection result.

[0086] Optionally, step 204 may specifically include: in the case where the second detection result is determined to be that the count register is normal, determine the software information of the system management interrupt in the basic input / output system; perform a fault detection on the software information to obtain a software detection result.

[0087] Optionally, when performing "determining software information of system management interrupts in the basic input / output system", it may include, but is not limited to: determining the total entry function, processor list function, script function of system management interrupts in the basic input / output system, and the serial port printing information of the basic input / output system.

[0088] In the embodiments of the present application, the total entry function of SMI is the code that the processor first executes when entering the system management mode (SMM). It is responsible for initializing the SMI processing environment and distributing specific SMI function handlers; before entering SMI, the total entry function of SMI must save all register and stack contents to avoid interfering with normal operations. The total entry function of SMI can also call corresponding function handlers according to the SMI trigger source (such as hardware events, software requests, etc.).

[0089] In some examples, in a multi-processor system, the processor list function is used to enumerate all available processors and determine which processors need to participate in SMI processing. The processor list function can use the CPUID instruction to obtain the identification information of the processors, and can also distinguish different processors through the local APIC ID of the Advanced Programmable Interrupt Controller (APIC).

[0090] For this embodiment, in the BIOS implementation, the script function is used to define and execute specific SMI processing logic. The script function is provided by the ACPI table or other configuration files. The script function needs to parse and execute a series of predefined commands. The script function allows dynamic adjustment of the SMI processing logic without modifying the firmware code.

[0091] As an optional method, during the BIOS development and debugging process, serial port printing is an important debugging tool. Key information during the SMI processing can be output to the serial port terminal for easy problem analysis; before using serial port printing, the serial port registers must be correctly initialized; serial port printing can output debugging information in real time for easy tracking of problems during the SMI processing.

[0092] Optionally, when performing "performing fault detection on the software information to obtain a software detection result", it may include, but is not limited to: judging whether the system management interrupt in the basic input / output system is abnormal based on the total entry function, processor list function, script function, and serial port printing information; determining the software detection result based on the judgment result.

[0093] Optionally, when performing "judging whether the system management interrupt in the basic input / output system is abnormal based on the total entry function, the processor list function, the script function, and the serial port print information; and determining the software detection result based on the judgment result", it may include but is not limited to: adding serial port print information to the total entry function and the processor list function respectively; calling the serial port collection tool for local area network serial transmission through the script function to collect the fault serial port information of the system management interrupt in the basic input / output system; in the case of collecting the fault serial port information, determining that the system management interrupt in the basic input / output system is abnormal, and saving the fault serial port information in a predetermined log file.

[0094] In the embodiment of the present application, add the basic input / output system (BIOS) serial port print information to the total entry function of the software system management interrupt (SW SMI) at the basic input / output system (BIOS) end and the software system management interrupt processor list (SW SMI Handler List) function, call the local area network serial transmission (Sol) serial port collection tool in the script function, print and save the basic input / output system (BIOS) serial port information, to see whether it reaches the basic input / output system (BIOS) end when the software system management interrupt function (SWSMI function) cannot be executed normally, and whether it is caused by the system management interrupt function (SMI function) under the basic input / output system (BIOS).

[0095] In some examples, the reading and collection of information in the embodiment of the present application will be integrated and printed into a log file (.log) document, and finally the document will be saved to the root directory.

[0096] In some examples, the BIOS serial port print can be used to output debug information, specifically including: enabling serial port debugging: enabling the serial port redirection or debug option in the BIOS settings. Reading the output: using a serial port terminal program (such as PuTTY, TeraTerm) to connect to the specified serial port to view the output log information. Custom printing: if necessary, add custom print statements (such as SerialPrint("Debug Info")) to the code segment suspected of having problems to further narrow down the problem range.

[0097] Optionally, when performing "determining the software detection result based on the judgment result", it may include but is not limited to: in the case of determining that the system management interrupt in the basic input / output system is abnormal, determining that the software detection result is that the software information is abnormal.

[0098] Step 205, generate the fault location information of the system management interrupt according to the hardware detection result and the software detection result.

[0099] For this embodiment, to solve the problem that when the SW SMI function fails to execute properly, quickly locate where the problem occurs so that it can be directly and quickly provided to the corresponding engineer for problem solving, the present invention proposes a method that can automatically locate the problem of the abnormal execution of the SW SMI. This method provides a relatively simple program code, which can be embedded into the program that triggers the SW SMI function and is only called when the SW SMI Function execution goes wrong and will not be called when the SW SMI function is executed normally. After the SW SMI function execution goes wrong, it will be directly executed and the parsed file will be directly generated and saved in the document. This method is simple and fast. When the SW SMI function execution error occurs, the problem can be located and analyzed according to the saved document. This method can quickly locate the problem of the reason why the SW SMI Function cannot be triggered normally, and can reduce the manpower and time investment of R & D personnel.

[0100] It should be noted that the server architecture of the Intel platform is used as an example in the embodiments of the present application, but this method is not limited to the servers of the Intel platform, nor is it limited to the server system only. The fault monitoring method provided in the embodiments of the present application can still be used for fault detection on the server systems of other platforms or other computer systems.

[0101] To illustrate the specific implementation process of this embodiment, the following specific application examples are given, such as Figure 4 shown, but not limited to this: Set up a program tool that runs under an operating system (OS), and embed this program tool into the program process triggered by the system management interrupt (SMI). When there is a problem that the software system management interrupt (SW SMI) cannot be executed normally, this program tool; in this program tool, first read the software system management interrupt control and enable register (SMI_EN) in the processor to check whether the corresponding byte (Bit) of this register is set, so as to rule out whether the system management interrupt (SMI) is temporarily disabled due to the loading of other software under the operating system (OS); then trigger the system management interrupt (SMI) by writing the function number through the I / O port with the port number 0xB2 (IOPort 0xB2), check whether it can be successfully triggered, and read the value of the register (SMI_COUNT) that counts the number of system management interrupt triggers in the processor to check whether there are intensive system management interrupts (SMI) generated, so as to rule out whether there are other system management interrupt functions (SMI function) occupying the system management interrupt (SMI) channel, resulting in the failure of this software system management interrupt function (SW SMI function) to be successfully executed; add basic input / output system (BIOS) serial port printing information in the total entry function of the software system management interrupt (SW SMI) and the software system management interrupt processor list (SW SMI Handler List) function at the basic input / output system (BIOS) end. Call the local area network serial transmission (Sol) serial port collection tool in the script function to print and save the basic input / output system (BIOS) serial port information, so as to check whether it reaches the basic input / output system (BIOS) end when the software system management interrupt function (SW SMI function) cannot be executed normally, and whether it is caused by the system management interrupt function (SMI function) under the basic input / output system (BIOS); all the above read and collected information will be integrated and printed into a log file (.log) document and finally saved to the root directory.

[0102] Compared with the related technology, in this embodiment, when there is an exception in the system management interrupt, an exception location program pre-embedded in the program of the system management interrupt can be called. Through the exception location program, the enable register and the count register in the processor can be detected to determine whether there is an exception in the hardware part of the processor, which in turn causes the system management interrupt exception; the software information of the system management interrupt can also be detected to analyze whether there is an exception in the program code of the system management interrupt, which in turn causes the system management interrupt exception. According to the results of the hardware detection and the software detection, the location information that causes the system management interrupt exception can be determined, that is, the enable register, the count register or the software program information. In addition, it can also reduce the human and time investment required for R & D personnel to determine the fault location after the system management interrupt exception.

[0103] An embodiment of the present application further provides a fault detection device, which includes: a call module 31, a detection module 32, and a generation module 33.

[0104] The call module 31 is configured to call an exception location program in response to an exception occurring during the process of the processor triggering a system management interrupt. The exception location program is embedded in the system management interrupt program; The detection module 32 is configured to perform fault detection on the enable register and the count register in the processor in sequence based on the exception location program to obtain a hardware detection result of the processor; The detection module 32 is further configured to perform fault detection on the software information of the system management interrupt in the basic input / output system chip to obtain a software detection result when it is determined that there is no exception in the hardware detection result; The generation module 33 is configured to generate fault location information of the system management interrupt based on the hardware detection result and the software detection result.

[0105] In some examples of this embodiment, the detection module 32 is specifically configured to perform fault detection on the enable register based on the exception location program to obtain a first detection result; when it is determined that there is no exception in the first detection result, perform fault detection on the count register based on the exception location program to obtain a second detection result. The first detection result and the second detection result constitute the hardware detection result.

[0106] In some examples of this embodiment, the detection module 32 is further specifically configured to read the enable register in the exception location program to determine the byte information corresponding to the enable register; determine whether the loading of the software in the processor affects the system management interrupt based on the byte information, and determine the first detection result based on the judgment result.

[0107] In some examples of this embodiment, the detection module 32 is further specifically configured to read the enable register in the exception location program to determine the target byte corresponding to the enable register, and detect whether the target byte is set; if it is determined that the target byte is set, generate a first byte information indicating that there is no setting exception for the target byte; if it is determined that the target byte is not set, generate a second byte information indicating that there is a setting exception for the target byte.

[0108] In some examples of this embodiment, the detection module 32 is further specifically configured to determine that the loading situation affects the system management interrupt based on the second byte information, and then determine that the first detection result is that the enable register has an exception.

[0109] In some examples of this embodiment, the detection module 32 is further specifically configured to set the target byte in the enable register when the first detection result is determined to indicate that the enable register has an abnormality.

[0110] In some examples of this embodiment, the detection module 32 is further specifically configured to determine that the first detection result indicates that the enable register has no abnormality if it is determined based on the first byte information that the loading condition has no impact on the system management interrupt.

[0111] In some examples of this embodiment, the detection module 32 is further specifically configured to, when the first detection result is determined to indicate that the enable register has no abnormality, trigger the system management interrupt again and determine the response condition of the count register; determine whether the system management interrupt channel in the processor is occupied based on the re-triggering condition and the response condition, and determine the second detection result based on the determination result.

[0112] In some examples of this embodiment, the detection module 32 is further specifically configured to write the target function number to the target port in the input / output port, trigger the system management interrupt again, and read the response value of the count register.

[0113] In some examples of this embodiment, the detection module 32 is further specifically configured to determine whether the processor has consecutive system management interrupts based on the re-triggering condition and the response value; if it is determined that the processor has consecutive system management interrupts, determine that the system management interrupt channel is occupied, and determine that the second detection result indicates that the count register has an abnormality.

[0114] In some examples of this embodiment, the detection module 32 is further specifically configured to repeatedly trigger the system management interrupt when the second detection result is determined to indicate that the count register has an abnormality.

[0115] In some examples of this embodiment, the detection module 32 is further specifically configured to determine that the system management interrupt channel is not occupied and determine that the second detection result indicates that the count register has no abnormality if it is determined that the processor does not have consecutive system management interrupts.

[0116] In some examples of this embodiment, the detection module 32 is further specifically configured to determine the software information of the system management interrupt in the basic input / output system when the second detection result is determined to indicate that the count register has no abnormality; perform a fault detection on the software information to obtain a software detection result.

[0117] In some examples of this embodiment, the detection module 32 is further specifically configured to determine the total entry function, the processor list function, the script function of the system management interrupt in the basic input / output system, and the serial port print information of the basic input / output system.

[0118] In some examples of this embodiment, the detection module 32 is further specifically configured to determine whether the system management interrupt in the basic input / output system is abnormal based on the total entry function, the processor list function, the script function, and the serial port print information; and determine the software detection result based on the determination result.

[0119] In some examples of this embodiment, the detection module 32 is further specifically configured to add serial port print information to the total entry function and the processor list function respectively; call the serial port collection tool for local area network serial transmission through the script function to collect the fault serial port information of the system management interrupt in the basic input / output system; in the case of collecting the fault serial port information, determine that the system management interrupt in the basic input / output system is abnormal, and save the fault serial port information in a predetermined log file.

[0120] In some examples of this embodiment, the detection module 32 is further specifically configured to, in the case of determining that the system management interrupt in the basic input / output system is abnormal, determine that the software detection result is that the software information is abnormal.

[0121] It should be noted that for other corresponding descriptions of each functional unit involved in the fault detection device provided in this embodiment, reference can be made to Figure 2 the corresponding description in, which will not be elaborated here.

[0122] Based on the above method as Figure 2 shown, correspondingly, this embodiment further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above method as Figure 2 shown is implemented.

[0123] Based on the above method as Figure 2 shown, correspondingly, this embodiment further provides a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, the above method as Figure 2 shown is implemented.

[0124] Based on such an understanding, the technical solution of this application can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of various implementation scenarios of this application.

[0125] Based on the above as Figure 2The method shown, as well as the virtual device embodiments, to achieve the above object, the embodiments of the present application further provide an electronic device, such as a personal computer, a server, the device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the above as Figure 2 the method shown.

[0126] In some embodiments, the above-mentioned physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may include a standard wired interface, a wireless interface (such as a WI-FI interface), etc. in some embodiments.

[0127] Those skilled in the art can understand that the above-mentioned physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or a combination of components, or different component arrangements.

[0128] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, and communication between other hardware and software in the information processing physical device.

[0129] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or can be implemented by hardware. By applying the solution of this embodiment, compared with the related art, in the case of a system management interrupt exception, this embodiment can call an exception location program pre-embedded in the system management interrupt program, and through the exception location program, the enable register and count register in the processor can be detected to determine whether there is an exception in the hardware part of the processor, which in turn causes the system management interrupt exception; it is also possible to detect the software information of the system management interrupt to analyze whether there is an exception in the program code of the system management interrupt, which in turn causes the system management interrupt exception, and determine the location information that causes the system management interrupt exception according to the results of the hardware detection and the software detection, that is, the enable register, the count register, or the software program information. In addition, it is also possible to reduce the human and time investment required for R & D personnel to determine the fault location after the system management interrupt exception.

[0130] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0131] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A fault detection system, characterized in that, including: a system disk, a processor, and a basic input / output system chip, wherein the processor is respectively connected to the system disk and the basic input / output system chip; in response to an exception occurring during the processor triggering a system management interrupt, the processor calls an exception location program embedded in the program of the system management interrupt in the system disk; the processor sequentially performs fault detection on an enable register and a count register in the processor based on the exception location program to obtain a hardware detection result of the processor; when the processor determines that there is no exception in the hardware detection result, the processor performs fault detection on software information of the system management interrupt in the basic input / output system chip to obtain a software detection result; the processor generates fault location information of the system management interrupt according to the hardware detection result and the software detection result.

2. A fault detection method, characterized in that including: in response to an exception occurring during the processor triggering a system management interrupt, calling an exception location program embedded in the program of the system management interrupt; sequentially performing fault detection on an enable register and a count register in the processor based on the exception location program to obtain a hardware detection result of the processor; when it is determined that there is no exception in the hardware detection result, performing fault detection on software information of the system management interrupt in the basic input / output system chip to obtain a software detection result; generating fault location information of the system management interrupt according to the hardware detection result and the software detection result.

3. The method according to claim 2, wherein Sequentially performing fault detection on an enable register and a count register in the processor based on the exception location program to obtain a hardware detection result of the processor, including: performing fault detection on the enable register based on the exception location program to obtain a first detection result; when it is determined that there is no exception in the first detection result, performing fault detection on the count register based on the exception location program to obtain a second detection result, and the first detection result and the second detection result constitute the hardware detection result.

4. The method according to claim 3, wherein Performing fault detection on the enable register based on the exception location program to obtain a first detection result, including: reading the enable register in the exception location program to determine byte information corresponding to the enable register; judging whether the loading of software in the processor affects the system management interrupt according to the byte information, and determining the first detection result based on the judgment result.

5. The method according to claim 4, characterized in that Reading the enable register in the exception location program to determine byte information corresponding to the enable register, including: reading the enable register in the exception location program to determine a target byte corresponding to the enable register, and detecting whether the target byte is set; if it is determined that the target byte is set, generating first byte information indicating that there is no setting exception for the target byte; if it is determined that the target byte is not set, generating second byte information indicating that there is a setting exception for the target byte.

6. The method according to claim 5, wherein Determining whether the loading situation of the software in the processor affects the system management interrupt based on the byte information, and determining the first detection result based on the determination result, includes: If it is determined according to the second byte information that the loading situation affects the system management interrupt, then determine the first detection result as that the enable register is abnormal.

7. The method according to claim 6, wherein The method further includes: When the first detection result is determined as that the enable register is abnormal, set the target byte in the enable register.

8. The method according to claim 5, wherein Determining whether the loading situation of the software in the processor affects the system management interrupt based on the byte information, and determining the first detection result based on the determination result, includes: If it is determined according to the first byte information that the loading situation does not affect the system management interrupt, then determine the first detection result as that the enable register is not abnormal.

9. The method according to claim 8, wherein When it is determined that the first detection result is not abnormal, performing a fault detection on the count register in the processor based on the abnormality location program to obtain a second detection result, includes: When the first detection result is determined as that the enable register is not abnormal, trigger the system management interrupt again and determine the response situation of the count register; Judging whether the system management interrupt channel in the processor is occupied based on the re-triggering situation and the response situation, and determining the second detection result based on the determination result.

10. The method according to claim 9, characterized in that, Triggering the system management interrupt again and determining the response situation of the count register, includes: Writing a target function number to a target port in the input / output port, triggering the system management interrupt again, and reading the response value of the count register.

11. The method according to claim 10, wherein Judging whether the system management interrupt channel in the processor is occupied based on the re-triggering situation and the response situation, and determining the second detection result based on the determination result, includes: Judging whether there are consecutive multiple system management interrupts in the processor based on the re-triggering situation and the response value; If it is determined that there are consecutive multiple system management interrupts in the processor, then determine that the system management interrupt channel is occupied, and determine the second detection result as that the count register is abnormal.

12. The method according to claim 11, wherein The method further includes: When the second detection result is determined as that the count register is abnormal, repeatedly trigger the system management interrupt.

13. The method according to claim 11, wherein After judging whether there are consecutive multiple system management interrupts in the processor based on the re-triggering situation and the response value, the method further includes: If it is judged that there are no consecutive multiple system management interrupts in the processor, then determine that the system management interrupt channel is not occupied, and determine the second detection result as that the count register is not abnormal.

14. The method according to claim 13, wherein When it is determined that the hardware detection result is not abnormal, performing a fault detection on the software information of the system management interrupt in the basic input / output system chip to obtain a software detection result, includes: When determining that the second detection result indicates that there is no abnormality in the count register, determine the software information of the system management interrupt in the basic input / output system; Perform a fault detection on the software information to obtain the software detection result.

15. The method according to claim 14, wherein Determining the software information of the system management interrupt in the basic input / output system includes: Determine the total entry function, the processor list function, the script function of the system management interrupt in the basic input / output system, and the serial port printing information of the basic input / output system.

16. The method according to claim 15, characterized in that, The performing a fault detection on the software information to obtain the software detection result includes: Based on the total entry function, the processor list function, the script function, and the serial port printing information, determine whether the system management interrupt is abnormal in the basic input / output system; Determine the software detection result based on the judgment result.

17. The method according to claim 16, characterized in that, The determining whether the system management interrupt is abnormal in the basic input / output system based on the total entry function, the processor list function, the script function, and the serial port printing information includes: Add the serial port printing information to the total entry function and the processor list function respectively; Call the serial port collection tool for local area network serial transmission through the script function to collect the fault serial port information of the system management interrupt in the basic input / output system; When the fault serial port information is collected, determine that there is an abnormality in the system management interrupt in the basic input / output system, and save the fault serial port information in a predetermined log file.

18. The method according to claim 17, characterized in that, The determining the software detection result based on the judgment result includes: When it is determined that there is an abnormality in the system management interrupt in the basic input / output system, determine that the software detection result indicates that the software information is abnormal.

19. A computer-readable storage medium, on which a computer program is stored, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1 to 18.

20. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein, The processor, when executing the computer program, implements the method according to any one of claims 1 to 18.

21. A computer program product having a computer program stored thereon, characterized in that, The computer program product, when executed by a processor, implements the method according to any one of claims 1 to 18.

Citation Information

Patent Citations

  • Method for positioning and detecting abnormal interruption sources used in embedded software

    CN106201892A

  • Memory CE fault processing method, system and related device

    CN111008091A

  • Fault processing method, computer system, baseboard management controller and system

    CN113407391A

  • Fault processing method, device and system for PCI (Peripheral Component Interconnect) equipment

    CN118245269A

  • System management interrupt source including a programmable counter and power management system employing the same

    US5606713A