Hardware detection process exception handling method and device and electronic equipment

By splitting the hardware detection process into independent execution units and performing status monitoring, the system failure spread caused by hardware detection process exceptions is solved, and the stable operation and security of the hardware detection process are achieved.

CN120492200APending Publication Date: 2025-08-15INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510570796.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, when an abnormality occurs during operation of the hardware detection process, it cannot be fundamentally handled, resulting in hardware failures in the system being unable to be detected in time and spread, affecting system stability and security.

Method used

Split the hardware detection process into multiple independent execution units, and monitor the status of each unit, and carry out targeted processing when abnormalities are found to ensure that the process returns to normal operation.

Benefits of technology

By refining and isolating the hardware detection process, the spread of hardware failures is avoided, the stable operation of the hardware detection process is ensured, and the hardware security and detection quality of computing devices are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492200A_ABST
    Figure CN120492200A_ABST
Patent Text Reader

Abstract

The invention discloses a hardware detection process exception handling method, a hardware detection process exception handling device and electronic equipment, and relates to the technical field of computers, a hardware detection process is split into a plurality of independent execution units according to various detection functions supported by the hardware detection process, and state monitoring is carried out on each execution unit in the running process, so that the hardware detection process is more accurate. When a certain execution unit is abnormal, the abnormal execution unit is processed in a targeted manner, so that the abnormality of the process is processed fundamentally, the hardware detection process can be recovered to a normal operation state in time, and the situation that hardware faults in a system are diffused due to the fact that the hardware faults cannot be detected in time is avoided; and a foundation is laid for improving the hardware safety of the to-be-detected computing equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device, and electronic device for handling hardware detection process exceptions. Background Art

[0002] Currently, computing devices such as servers generally perform hardware status detection on field replaceable units (FRUs) based on a baseboard management controller (BMC).

[0003] In related technologies, BMC performs hardware status detection based on a hardware detection process. However, exceptions are inevitable during the operation of the hardware detection process. The current exception handling method is generally to restart the hardware detection process by starting a daemon process after determining that the hardware detection process has an exception. However, the restarted hardware detection process may have the same exception, that is, the exception of the process cannot be fundamentally handled, which causes the hardware detection process to be unable to operate normally for a long time, resulting in the spread of hardware faults in the system due to lack of timely detection. Summary of the Invention

[0004] The present application provides a method, device and electronic device for handling hardware detection process exceptions, so as to at least solve the problem in the related art that the process exceptions cannot be fundamentally handled.

[0005] This application provides a method for handling hardware detection process exceptions, including:

[0006] Get the configuration file of the hardware detection process;

[0007] Splitting the hardware detection process into a plurality of independent execution units according to the plurality of detection functions supported by the hardware detection process represented by the configuration file; wherein a single execution unit corresponds to at least one detection function;

[0008] During the operation of the hardware detection process, the status of each execution unit is monitored to obtain a status monitoring result of each execution unit;

[0009] The execution unit that is abnormal as a result of the status monitoring is regarded as the abnormal execution unit, and a corresponding abnormal processing operation is performed on the abnormal execution unit.

[0010] The present application also provides a hardware detection process exception handling device, comprising:

[0011] The acquisition module is used to obtain the configuration file of the hardware detection process;

[0012] a splitting module, configured to split the hardware detection process into a plurality of independent execution units according to the plurality of detection functions supported by the hardware detection process as represented by the configuration file; wherein a single execution unit corresponds to at least one detection function;

[0013] A monitoring module, configured to monitor the status of each execution unit during the running of the hardware detection process and obtain a status monitoring result of each execution unit;

[0014] The processing module is used to take the abnormal execution unit as the abnormal execution unit according to the status monitoring result, and perform corresponding abnormal processing operations on the abnormal execution unit.

[0015] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned hardware detection process exception handling methods when executing the computer program.

[0016] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned hardware detection process exception handling methods are implemented.

[0017] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned hardware detection process exception handling methods when executed by a processor.

[0018] Through this application, the hardware detection process is split into multiple independent execution units based on the multiple detection functions it supports, and the status of each execution unit is monitored separately during operation. When an exception occurs in an execution unit, the abnormal execution unit is processed in a targeted manner to fundamentally handle the process exception, ensuring that the hardware detection process can be restored to normal operating status in a timely manner, avoiding the spread of hardware failures in the system due to lack of timely detection, and laying the foundation for improving the hardware security of the computing device to be detected. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1 This is a structural diagram of the hardware detection process exception handling system based on the embodiment of the present application;

[0021] Figure 2 A flowchart of a method for handling hardware detection process exceptions provided in an embodiment of the present application;

[0022] Figure 3 A schematic diagram of the structure of an exemplary hardware detection process exception handling system provided in an embodiment of the present application;

[0023] Figure 4 A schematic diagram of the overall process of the hardware detection process exception handling method provided in an embodiment of the present application;

[0024] Figure 5 A schematic diagram of the structure of a hardware detection process exception handling device provided in an embodiment of the present application;

[0025] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0026] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0027] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0028] In a BMC-based server management system, FRU management is a crucial part. BMC supports system health and stability by monitoring and managing various aspects of server hardware, including temperature, voltage, light status, hardware operating status, etc. The main responsibility of the FRU management process (hardware detection process) is to access the hardware and collect and process this hardware status information. If these processes crash, especially due to abnormal hardware access or data processing, it may cause external interface (such as IPMI, Redfish) query failures, thereby affecting the monitoring and management of system health. If these processes fail to recover or are unstable for a long time, it may cause system failures to spread and even affect the management of other hardware units.

[0029] In related technologies, redundant or daemon processes are often used to ensure high availability of the FRU management process. If a process crashes, the system automatically starts a backup process (redundant or daemon) to restore service, reducing the risk of system failure. However, this approach cannot precisely manage hardware access units within a single process and is difficult to prevent fault propagation.

[0030] In order to solve the above problems, the embodiments of the present application provide a method, device and electronic device for handling hardware detection process exceptions, the method comprising: obtaining a configuration file of the hardware detection process; splitting the hardware detection process into multiple independent execution units according to the multiple detection functions supported by the hardware detection process represented by the configuration file; wherein a single execution unit corresponds to at least one detection function; during the operation of the hardware detection process, monitoring the status of each execution unit to obtain the status monitoring results of each execution unit; treating the execution unit with the status monitoring result as an abnormal execution unit as the abnormal execution unit, and performing corresponding exception handling operations on the abnormal execution unit. The method provided by the above scheme, by splitting the hardware detection process into multiple independent execution units according to the multiple detection functions it supports, monitoring the status of each execution unit separately during the operation process, and when an execution unit is abnormal, performing targeted processing on the abnormal execution unit, so as to fundamentally handle the process abnormality, ensure that the hardware detection process can be restored to normal operation in a timely manner, avoid the spread of hardware failures in the system due to lack of timely detection, and lay the foundation for improving the hardware security of the computing device to be detected.

[0031] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0032] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the hardware detection process exception handling method depends, the specific application environment architecture or specific hardware architecture is described herein.

[0033] First, the structure of the hardware detection process exception handling system on which this application is based is described:

[0034] The hardware detection process exception handling method, device and electronic device provided in the embodiments of the present application are suitable for handling exceptions in the hardware detection process used to detect the hardware status of a computing device, so as to ensure that the hardware detection process can run continuously and stably. Figure 1The figure shows a schematic diagram of the structure of the hardware detection process exception handling system based on the embodiment of the present application, which mainly includes the hardware to be detected, the hardware detection process, and the hardware detection process exception handling device. Among them, the hardware detection process is used to perform hardware status detection on the hardware to be detected, and the hardware detection process exception handling device is used to determine whether there are abnormal execution units in the hardware detection process during the operation of the hardware detection process. If it is determined that there are abnormal execution units, the abnormal execution units are targeted and handled to fundamentally handle the process exception.

[0035] The present invention provides a method for handling hardware detection process exceptions, which is used to handle exceptions in a hardware detection process that detects the hardware status of a computing device, thereby ensuring that the hardware detection process can continue to run stably. The execution subject of the present invention is an electronic device, such as a server, a desktop computer, a laptop computer, a tablet computer, or other electronic device that can be used for hardware detection process exception handling.

[0036] like Figure 2 FIG. 1 is a flow chart of a method for handling hardware detection process exceptions according to an embodiment of the present application, the method comprising:

[0037] Step 201: Obtain a configuration file for the hardware detection process.

[0038] The hardware detection process, also known as the FRU process, is a process within the BMC that monitors the hardware status of computing devices. The BMC can initiate multiple hardware detection processes in parallel. The configuration file contains the parameters and settings required to run the hardware detection process. It records the various detection functions supported by the hardware detection process, as well as configuration information related to each detection function, such as the enable status of each execution unit, the number of retries upon failure, and the timeout period.

[0039] Step 202 : split the hardware detection process into multiple independent execution units according to the multiple detection functions supported by the hardware detection process represented by the configuration file.

[0040] Each execution unit corresponds to at least one detection function. The execution unit is the fundamental unit for implementing hardware detection tasks. It can independently perform hardware status detection operations such as hardware access, data collection and processing.

[0041] Specifically, the single hardware detection process can be divided into multiple independent execution units based on the various detection functions recorded in the configuration file, thereby achieving detailed and isolated detection functions. For example, if a hardware detection process needs to detect multiple functions such as hardware temperature, voltage, and operating status, the execution code of these functions can be split into corresponding execution units through splitting, and each execution unit can run independently.

[0042] Step 203: During the execution of the hardware detection process, the status of each execution unit is monitored to obtain the status monitoring result of each execution unit.

[0043] Specifically, after the hardware detection process is started and begins running, the operating status of each execution unit is continuously monitored in real time. The monitoring content includes at least whether the execution unit starts normally, whether errors occur during the execution of the detection task, and whether the task is completed within the specified time.

[0044] Step 204: The execution unit whose status monitoring result indicates an abnormal execution unit is regarded as an abnormal execution unit, and corresponding abnormal processing operations are performed on the abnormal execution unit.

[0045] Specifically, when the status monitoring results of a particular execution unit indicate that its state does not meet normal standards, it is identified as an abnormal execution unit. For this abnormal execution unit, corresponding exception handling operations can be performed according to pre-set strategies to restore the abnormal execution unit to a normal state or to ensure that the abnormal execution unit does not affect the overall operation of the hardware detection process. This ensures that the entire hardware detection process can complete some detection tasks even if some execution units are abnormal, thus ensuring the normal operation of the hardware system.

[0046] Based on the above embodiment, as an implementable approach, in one embodiment, during the execution of the hardware detection process, the status of each execution unit is monitored to obtain the status monitoring results of each execution unit, including:

[0047] Step 2031: During the hardware detection process, the execution units are sequentially traversed according to the execution order of the execution units indicated by the configuration file to collect the execution status information of each execution unit.

[0048] Step 2032: Monitor the status of each execution unit based on the execution status information of each execution unit to determine whether the status monitoring result of each execution unit is abnormal.

[0049] Specifically, during the hardware detection process, the execution units can be accessed one by one according to the execution order specified in the previously acquired configuration file. For example, if the configuration file specifies that the temperature detection unit should be executed first, followed by the voltage detection unit, then the temperature detection execution unit and the voltage detection execution unit will be operated in this order. When accessing each execution unit, various information related to the execution unit's operation is collected, thus obtaining the execution status information of the execution unit.

[0050] Specifically, after obtaining the execution status information of each execution unit, the information is analyzed and judged to determine the comparison between the collected status information and the pre-set normal status standard. For example, if the normal execution time range of an execution unit is 5 seconds, but the collected status information shows that the execution unit has been running for 10 seconds and has not yet completed the task, then the execution unit status can be determined to be abnormal.

[0051] By monitoring and assessing the status of execution units, problems that arise during their operation can be promptly identified. Once an anomaly is detected, appropriate measures can be quickly implemented, such as adjusting the execution strategy, retrying, or disabling the anomalous execution unit, to prevent the anomaly from spreading further. This real-time monitoring and anomaly assessment mechanism also helps improve the quality of hardware testing and ensure the accuracy of test results.

[0052] Specifically, in one embodiment, for any execution unit, it can be determined whether the execution unit has a preset fault based on the execution status information of the execution unit; when it is determined that any preset fault has occurred in the execution unit, the status monitoring result of the execution unit is determined to be abnormal.

[0053] The preset faults include at least hardware access timeout, bus conflict, data verification failure and sensor failure.

[0054] Specifically, when an execution unit accesses a hardware device, if the data read or write operation fails to be completed within the specified time, a hardware access timeout fault will be triggered. When multiple execution units attempt to transmit data on the bus at the same time, a bus conflict may occur. By monitoring the data transmission status in the execution status information of the execution unit, it can be determined whether a bus conflict has occurred. To ensure the accuracy and integrity of the data, verification is usually performed during the data transmission and processing process. If the data received by the execution unit is found to be inconsistent with the expectations after verification, that is, the data verification fails, it is determined that the execution unit has a data verification failure fault. The sensor is responsible for collecting various status data of the hardware device, such as temperature, pressure, etc. If the sensor itself fails, the returned data may be an invalid value (such as -1 or a value outside the reasonable range). The sensor failure will cause the execution unit to obtain incorrect hardware status information, thereby affecting the judgment of the operating status of the hardware device.

[0055] Among them, such as Figure 3As shown, it is a structural diagram of an exemplary hardware detection process exception handling system provided by an embodiment of the present application. The system includes a process manager (FRU manager), a hardware detection process and a hardware layer. The method provided by the embodiment of the present application can be applied to the process manager. The process manager includes a configuration management module, a process monitoring module, a fault handling engine and a state database. Among them, the configuration management module is used to parse and maintain the configuration files (YAML format) of all hardware detection processes and supports dynamic updates. The process monitoring module is used to monitor the process status through a heartbeat mechanism and event markers and trigger exception handling. The fault handling engine is used to analyze the fault type based on the rule engine and perform operations such as disabling units and restarting processes. The state database is used to store information such as execution unit status, historical logs and fault modes. The hardware detection process (FRU process) includes multiple execution units, each of which supports atomic start and stop, such as execution units 1 to 3 are read-in-place execution units, read-temperature execution units and light-up execution units respectively. Among them, the hardware detection process also includes a communication interface for interacting with the FRU manager through shared memory and sending event markers; it also includes a local cache for temporarily storing hardware data, that is, storing historical detection data in the local cache to avoid data loss due to manager response delays. The hardware detection process checks hardware status through the hardware layer. This layer includes the hardware driver interface, which is a standardized access interface (such as the I2C / GPIO package library) that isolates hardware differences. The hardware layer also includes a virtualization layer, also known as the hardware test interface, which provides a simulated interface for key hardware (such as sensors) for testing and fault injection.

[0056] Based on the above embodiment, as an implementable manner, in one embodiment, performing corresponding exception handling operations on the exception execution unit includes:

[0057] Step 2041, when the preset fault occurring in the abnormal execution unit is a hardware access timeout;

[0058] Step 2042: Filter the configuration file for the unit configuration information of the abnormal execution unit;

[0059] Step 2043: retry the abnormal execution unit according to the maximum retry count indicated by the unit configuration information;

[0060] Step 2044: When the number of retries of the abnormal execution unit exceeds the maximum number of retries, a unit disabling parameter is added to the unit configuration information of the abnormal execution unit to disable the abnormal execution unit in the hardware detection process;

[0061] Step 2045: If the preset fault of the abnormal execution unit is a bus conflict, retry the abnormal execution unit after a preset delay.

[0062] Step 2046 , when the number of retries of the abnormal execution unit exceeds the maximum number of retries, the bus priority is adjusted so that the bus responds to the abnormal execution unit first;

[0063] Step 2046 , when the preset fault of the abnormal execution unit is a data verification failure, the data of the abnormal execution unit is rolled back to replace the current data to be verified with the latest historical detection data cached locally by the abnormal execution unit;

[0064] Step 2047, when the preset fault of the abnormal execution unit is sensor failure, a unit disable parameter is added to the unit configuration information of the abnormal execution unit to disable the abnormal execution unit in the hardware detection process, and the hardware to be detected corresponding to the abnormal execution unit is marked as a fault state.

[0065] It should be noted that preset faults also include process crashes, hardware resource exhaustion, and configuration errors. After determining that a preset fault has occurred in the abnormal execution unit, a corresponding error code can be generated. The error code is used to represent the specific type of the preset fault. For example, the error code corresponding to a hardware access timeout is 0x1001, the error code corresponding to a bus conflict is 0x1002, the error code corresponding to a process crash is 0x1003, the error code corresponding to a data check failure is 0x1004, the error code corresponding to a sensor failure is 0x1005, the error code corresponding to hardware resource exhaustion is 0x1006, and the error code corresponding to a configuration error is 0x1007.

[0066] Specifically, for hardware access timeout failures, the unit configuration information of the abnormal execution unit experiencing the hardware access timeout can be filtered from the configuration file. The unit configuration information is the file content corresponding to the abnormal execution unit in the configuration file. By filtering the unit configuration information, the specific configuration for the abnormal execution unit can be quickly obtained. Retry processing is performed on the abnormal execution unit experiencing the hardware access timeout based on the maximum retry count specified in the filtered unit configuration information. The purpose of the retry is to attempt to reaccess the hardware device in the hope of successfully acquiring data and resuming the normal detection process. This operation is based on the assumption that the hardware access timeout may be caused by temporary hardware busyness or communication problems. Retrying may solve the problem and avoid abandoning the detection task due to a single failure. During the retry process for the abnormal execution unit, if the number of retries exceeds the maximum retry count specified in the configuration file, it indicates that multiple attempts have failed to successfully access the hardware. In this case, a unit disable parameter can be added to the unit configuration information of the abnormal execution unit. In this way, the abnormal execution unit will be disabled during subsequent hardware detection processes and will no longer participate in detection. This prevents the stability of the entire detection process from being affected by the continuously failing execution unit and also avoids invalid access to the faulty hardware.

[0067] Specifically, for bus conflict faults, when the preset fault of the abnormal execution unit is a bus conflict, since the bus conflict may be a temporary conflict caused by multiple devices competing for bus resources at the same time, a strategy of delaying the abnormal execution unit for a preset period of time before retrying the abnormal execution unit is adopted to reduce the possibility of conflict. If the number of retries exceeds the maximum number of retries when retrying the abnormal execution unit that has a bus conflict, it indicates that the bus conflict problem cannot be resolved by delaying the retry alone. Therefore, the bus priority can be adjusted to make the bus respond to the abnormal execution unit first. By changing the allocation priority of the bus resources, it is ensured that the execution unit can obtain bus resources first, thereby ensuring that the detection task can continue.

[0068] Specifically, for data verification failures, in order to ensure the accuracy of process data and the reliability of detection results, the abnormal execution unit can be subjected to a data rollback operation, that is, the latest historical detection data cached in the local cache is used to replace the current problematic data to be verified, so as to avoid subsequent detection results errors due to incorrect data, and ensure the consistency and accuracy of the detection process.

[0069] Specifically, for sensor failure, if the default fault of an abnormal execution unit is sensor failure, it means that the sensor on which the abnormal execution unit relies cannot normally provide hardware data. Therefore, a unit disable parameter is added to the unit configuration information of the abnormal execution unit, disabling the abnormal execution unit during the hardware detection process to prevent the use of incorrect sensor data for subsequent processing. At the same time, the hardware to be detected corresponding to the abnormal execution unit is marked as faulty, making it easier for subsequent operation and maintenance personnel to quickly locate and resolve the hardware fault and promptly repair or replace the faulty hardware.

[0070] Specifically, for configuration error failures, that is, configuration file parsing failure or invalid parameters, the configuration file can be restored to the default configuration and an alarm can be recorded to remind subsequent operation and maintenance personnel to handle the failure in a timely manner.

[0071] Exemplarily, a process configuration file includes at least a unique process identifier (process_id) for distinguishing different FRU processes (such as fans, PSUs, etc.), a brief description of the process (description) to enhance process readability, a list of execution unit parameters (units) for distinguishing different execution units, and unit configuration information for the execution unit, including at least the enabled state (enabled) of the execution unit, the maximum number of retries (retry_max) when an operation fails, and the timeout threshold (timeout_ms, in milliseconds). Adding a unit disable parameter to the unit configuration information of the abnormal execution unit changes the enabled state to the disabled state. The inter-process communication data structure is as follows:

[0072] Struct UnitEvent{

[0073] string process_id; / / FRU process ID

[0074] string unit_id; / / Execution unit ID

[0075] EventType type; / / Event type (START / END / FAIL)

[0076] int64 timestamp; / / timestamp

[0077] int64 errorCode; / / error code

[0078] }

[0079] That is, if a failure is detected in an execution unit during process execution, the aforementioned communication data is generated, using an error code to indicate the failure of the abnormal execution unit. The event type represents the execution status of the execution unit. START indicates that the execution unit has begun executing the task and started the timer. END indicates that the execution unit has completed the task and obtained the execution result of the execution unit. FAIL indicates that a failure occurred during the execution of the execution unit.

[0080] Furthermore, in one embodiment, when the hardware detection process meets a preset restart condition, a latest configuration file is obtained; and based on the latest configuration file, the hardware detection process is restarted.

[0081] In the latest configuration file, the unit configuration information of the abnormal execution unit includes a unit disabling parameter to disable the abnormal execution unit during the restart of the hardware detection process.

[0082] Specifically, after obtaining the latest configuration file, the hardware detection process can be restarted based on the latest configuration file. During the restart process, since the unit configuration information of the abnormal execution unit includes a unit disable parameter, certain abnormal execution units will be prohibited from participating in the hardware detection process after the restart based on the disable parameter. This ensures that the hardware detection process can run in a stable state after the restart, avoiding new failures caused by problems with the previous abnormal execution unit.

[0083] Specifically, in one embodiment, a heartbeat signal sent by the hardware detection process according to a preset period can be obtained; based on the acquisition of the heartbeat signal, it is determined whether the hardware detection process has crashed; if it is determined that the hardware detection process has crashed, it is determined that the hardware detection process meets the preset restart conditions.

[0084] It should be noted that the hardware detection process sends a heartbeat signal according to a preset period. If the acquisition of the heartbeat signal indicates that the hardware detection process has not sent a heartbeat signal for more than the preset period, it is determined that the hardware detection process has crashed.

[0085] Specifically, a process crash can prevent hardware detection from working properly, severely impacting the system's hardware status monitoring. To restore the hardware detection function and ensure the normal operation of the system, the hardware detection process can be restarted to resume working and ensure continued effective detection and monitoring of the hardware.

[0086] Specifically, in one embodiment, hardware resource consumption information corresponding to the hardware detection process can be obtained; when the hardware resource consumption information indicates that the system hardware resources are exhausted, it is determined that the hardware detection process meets the preset restart conditions, so as to release the hardware resources by restarting the hardware detection process.

[0087] It should be noted that the hardware detection process will occupy various system hardware resources, such as CPU and memory, during its operation.

[0088] Specifically, by monitoring the usage of system hardware resources in real time, the hardware resource consumption information corresponding to the hardware detection process is obtained, and the degree of hardware resource occupation by the current hardware detection process during operation is determined. If the hardware resource consumption information indicates that the system hardware resources are exhausted, that is, the usage of certain key hardware resources reaches or exceeds the upper limit of the system, then the hardware resources are determined to be exhausted. Restarting the hardware detection process can release the hardware resources it occupies. After the restart, the hardware detection process will re-run in the initial state. At this time, the hardware detection process will re-apply for the required hardware resources instead of continuing to occupy the previously exhausted resources, so that the system's hardware resources can be released.

[0089] For example, Figure 4As shown, it is a schematic diagram of the overall process of the hardware detection process exception handling method provided by an embodiment of the present application. In the startup phase, the FRU process is first initialized. After the FRU process is started, the configuration file (such as fru_fan.conf) is read, the metadata of the execution unit (such as the unique identifier of the execution unit unit_id, the startup status enabled and the maximum number of retries retry_max, etc.) is parsed, and a registration request is sent to the FRU process manager, carrying the process ID and the list of supported functional units. After receiving the registration request, the FRU manager records the metadata of the FRU process and returns the list of currently disabled units. The FRU process starts the heartbeat thread and sends the survival status and resource occupancy rate to the FRU manager regularly (such as every 5 seconds). The FRU manager records the real-time status of each FRU process and its execution unit to maintain the global status table. During the running phase, the FRU process traverses the execution units in the order in the configuration file and checks the enabled status (enabled) of each unit. For enabled units, a UNIT_START event is sent to the FRU manager, carrying the unit ID and time The time stamp is set and a timeout timer is started. The unit operation is performed (e.g., reading temperature sensor data via I2C). The collected data is processed (e.g., unit conversion and data formatting). After completion, a UNIT_END event is sent to the FRU manager, carrying the result data or an error code (e.g., 0x1001 for hardware access timeout). If the execution fails, retries are repeated according to the retry_max value until it is exhausted, triggering a UNIT_FAIL event. During the exception handling phase, after receiving a UNIT_END or UNIT_FAIL event, the FRU manager parses the error code (e.g., 0x1001). The error code is matched to the fault type (e.g., 0x1001 for hardware access timeout), and the response strategy in the fault classification and strategy table is referenced. This table can be continuously expanded based on subsequent needs, and the corresponding strategy is executed. The fault classification and strategy table is maintained based on at least steps 2041 to 2047 above. The FRU manager updates the unit status (e.g., marking read_temp as disabled) and persists the status change to the database, which updates the configuration file to ensure configuration consistency after restart.

[0090] Specifically, in actual applications, in addition to supporting the above-mentioned automatic updates, the process configuration file also supports manual updates, which increases the debugging means for on-site personnel. Specific functions include:

[0091] 1) Configuration Hot Update: Operations personnel submit a new configuration file (e.g., increase retry_max) through the FRU Manager interface. After the FRU Manager verifies the configuration, it notifies the FRU process to reload the configuration, which takes effect without restarting.

[0092] 2) Status query: The external system queries the status of the FRU process and execution unit through the FRU manager interface.

[0093] 3) Manual intervention and repair: Operations and maintenance personnel locate the faulty hardware (such as replacing a damaged sensor) based on the error log and status dashboard. After repair, they manually enable the disabled unit through the manager.

[0094] Specifically, in one embodiment, historical data of each execution unit under different hardware environments and operating conditions can also be collected, including execution time, number of retries, frequency of error code occurrence, and hardware status information, and this information can be used as execution unit status information. This data is used to train a machine learning model, and when the hardware detection process is running, the model analyzes and predicts based on the execution unit status information obtained in real time. If the model predicts that a certain execution unit has a high probability of hardware access timeout or other preset failures in the future, measures can be taken in advance, such as increasing the number of retries for the execution unit in advance, disabling it in advance, or modifying the bus priority in advance, so as to avoid failures in the execution unit affecting the overall operation of the process.

[0095] Specifically, while the hardware detection process is running, the model uses real-time execution unit status information as input to predict failures. The model outputs the probability and failure category of each execution unit failing within a certain timeframe. If the predicted failure probability of a particular execution unit exceeds a preset threshold, preventive measures are implemented, and operations personnel are notified to perform maintenance on that execution unit in advance to prevent failures.

[0096] The hardware detection process exception handling method provided by the embodiment of the present application obtains the configuration file of the hardware detection process; according to the multiple detection functions supported by the hardware detection process represented by the configuration file, the hardware detection process is split into multiple independent execution units; wherein, a single execution unit corresponds to at least one detection function; during the operation of the hardware detection process, the status of each execution unit is monitored to obtain the status monitoring results of each execution unit; the execution unit with the status monitoring result being abnormal is regarded as the abnormal execution unit, and the abnormal execution unit is subjected to corresponding exception handling operations. The method provided by the above scheme splits the hardware detection process into multiple independent execution units according to the multiple detection functions it supports, monitors the status of each execution unit separately during the operation process, and when an execution unit is abnormal, the abnormal execution unit is handled in a targeted manner to fundamentally handle the process abnormality, avoid a single fault causing a process-level crash, ensure that the hardware detection process can be restored to normal operation in a timely manner, and avoid the spread of hardware faults in the system due to lack of timely detection, thereby laying a foundation for improving the hardware security of the computing device to be detected. In addition, by dynamically disabling the faulty unit based on real-time error code classification and skipping the abnormal operation by hot updating the configuration file, functional degradation is achieved without interrupting the core service. In addition, by dynamically loading configuration files at runtime (such as adjusting the number of retries and timeout thresholds), it supports rapid on-site repair of problems without replacing the BMC firmware, improving operation and maintenance efficiency while reducing on-site maintenance costs.

[0097] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0098] An embodiment of the present application further provides a hardware detection process exception handling device, which is used to execute the hardware detection process exception handling method provided in the above embodiment.

[0099] like Figure 5 FIG. 5 is a schematic diagram of a hardware detection process exception handling device according to an embodiment of the present invention. The hardware detection process exception handling device 50 comprises: an acquisition module 501 , a splitting module 502 , a monitoring module 503 and a processing module 504 .

[0100] Among them, the acquisition module is used to obtain the configuration file of the hardware detection process; the splitting module is used to split the hardware detection process into multiple independent execution units according to the multiple detection functions supported by the hardware detection process represented by the configuration file; wherein, a single execution unit corresponds to at least one detection function; the monitoring module is used to monitor the status of each execution unit during the operation of the hardware detection process and obtain the status monitoring results of each execution unit; the processing module is used to treat the execution unit with an abnormal status monitoring result as an abnormal execution unit as an abnormal execution unit, and perform corresponding exception handling operations on the abnormal execution unit.

[0101] Specifically, in one embodiment, the monitoring module is specifically configured to:

[0102] During the execution of the hardware detection process, traversing each execution unit in sequence according to the execution order of the execution units represented by the configuration file to collect execution status information of each execution unit;

[0103] The status of each execution unit is monitored according to the execution status information of each execution unit to determine whether the status monitoring result of each execution unit is abnormal.

[0104] Specifically, in one embodiment, the monitoring module is specifically configured to:

[0105] For any of the execution units, judging whether a preset fault occurs in the execution unit according to the execution status information of the execution unit;

[0106] In the case where it is determined that any preset fault occurs in the execution unit, determining that the state monitoring result of the execution unit is abnormal;

[0107] The preset faults include at least hardware access timeout, bus conflict, data verification failure and sensor failure.

[0108] Specifically, in one embodiment, the processing module is specifically configured to:

[0109] In the case where the preset fault occurring in the abnormal execution unit is a hardware access timeout;

[0110] Filtering the unit configuration information of the abnormal execution unit in the configuration file;

[0111] retrying the abnormal execution unit according to the maximum retry number represented by the unit configuration information;

[0112] When the number of retries of the abnormal execution unit exceeds the maximum number of retries, adding a unit disabling parameter to the unit configuration information of the abnormal execution unit to disable the abnormal execution unit in the hardware detection process;

[0113] In the case where the preset fault occurring in the abnormal execution unit is a bus conflict, retrying the abnormal execution unit after a preset delay;

[0114] When the number of retries of the abnormal execution unit exceeds the maximum number of retries, adjusting the bus priority so that the bus responds to the abnormal execution unit first;

[0115] When the preset fault occurring in the abnormal execution unit is a data verification failure, the abnormal execution unit is subjected to data rollback to replace the current data to be verified with the latest historical detection data cached locally by the abnormal execution unit;

[0116] When the preset fault occurring in the abnormal execution unit is sensor failure, a unit disable parameter is added to the unit configuration information of the abnormal execution unit to disable the abnormal execution unit in the hardware detection process, and the hardware to be detected corresponding to the abnormal execution unit is marked as a fault state.

[0117] Specifically, in one embodiment, the device further includes a restart module, configured to:

[0118] When the hardware detection process meets the preset restart condition, obtaining the latest configuration file;

[0119] Restarting the hardware detection process based on the latest configuration file;

[0120] Wherein, in the latest configuration file, the unit configuration information of the abnormal execution unit includes a unit disabling parameter, so as to disable the abnormal execution unit during the process of restarting the hardware detection process.

[0121] Specifically, in one embodiment, the restart module is further configured to:

[0122] Obtaining a heartbeat signal sent by the hardware detection process according to a preset period;

[0123] Determining whether a process crash occurs in the hardware detection process based on the acquisition of the heartbeat signal;

[0124] When it is determined that the hardware detection process crashes, it is determined that the hardware detection process meets a preset restart condition.

[0125] Specifically, in one embodiment, the restart module is further configured to:

[0126] Obtaining hardware resource consumption information corresponding to the hardware detection process;

[0127] In a case where the hardware resource consumption information indicates that the system hardware resources are exhausted, it is determined that the hardware detection process meets a preset restart condition, so as to release the hardware resources by restarting the hardware detection process.

[0128] For the description of the features in the embodiment corresponding to the hardware detection process exception handling device, please refer to the relevant description of the embodiment corresponding to the hardware detection process exception handling method, which will not be repeated here.

[0129] The embodiment of the present application also provides an electronic device, such as Figure 6 As shown, it is a structural diagram of an electronic device provided in an embodiment of the present application, including a processor 10 and a memory 20, in which a computer program is stored. The processor 10 is configured to run the computer program to execute the steps in any of the above-mentioned hardware detection process exception handling method embodiments.

[0130] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned hardware detection process exception handling method embodiments when running.

[0131] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0132] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned hardware detection process exception handling method embodiments are implemented.

[0133] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned hardware detection process exception handling method embodiments.

[0134] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0135] The above is a detailed introduction to a hardware detection process exception handling method, device, and electronic device provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A method for handling hardware detection process exceptions, characterized in that: include: Get the configuration file of the hardware detection process; Splitting the hardware detection process into a plurality of independent execution units according to the plurality of detection functions supported by the hardware detection process represented by the configuration file; wherein a single execution unit corresponds to at least one detection function; During the operation of the hardware detection process, the status of each execution unit is monitored to obtain a status monitoring result of each execution unit; The execution unit that is abnormal as a result of the status monitoring is regarded as the abnormal execution unit, and a corresponding abnormal processing operation is performed on the abnormal execution unit.

2. The method for handling hardware detection process exceptions according to claim 1, wherein: During the running of the hardware detection process, the status of each execution unit is monitored to obtain the status monitoring result of each execution unit, including: During the execution of the hardware detection process, traversing each execution unit in sequence according to the execution order of the execution units represented by the configuration file to collect execution status information of each execution unit; The status of each execution unit is monitored according to the execution status information of each execution unit to determine whether the status monitoring result of each execution unit is abnormal.

3. The method for handling hardware detection process exceptions according to claim 2, wherein: The step of monitoring the status of each execution unit according to the execution status information of each execution unit to determine whether the status monitoring result of each execution unit is abnormal includes: For any of the execution units, judging whether a preset fault occurs in the execution unit according to the execution status information of the execution unit; In the case where it is determined that any preset fault occurs in the execution unit, determining that the state monitoring result of the execution unit is abnormal; The preset faults include at least hardware access timeout, bus conflict, data verification failure and sensor failure.

4. The method for handling hardware detection process exceptions according to claim 1, wherein: The performing corresponding exception handling operations on the exception execution unit includes: In the case where the preset fault occurring in the abnormal execution unit is a hardware access timeout; Filtering the unit configuration information of the abnormal execution unit in the configuration file; retrying the abnormal execution unit according to the maximum retry number represented by the unit configuration information; When the number of retries of the abnormal execution unit exceeds the maximum number of retries, adding a unit disabling parameter to the unit configuration information of the abnormal execution unit to disable the abnormal execution unit in the hardware detection process; In the case where the preset fault occurring in the abnormal execution unit is a bus conflict, retrying the abnormal execution unit after a preset delay; When the number of retries of the abnormal execution unit exceeds the maximum number of retries, adjusting the bus priority so that the bus responds to the abnormal execution unit first; When the preset fault occurring in the abnormal execution unit is a data verification failure, the abnormal execution unit is subjected to data rollback to replace the current data to be verified with the latest historical detection data cached locally by the abnormal execution unit; When the preset fault occurring in the abnormal execution unit is sensor failure, a unit disable parameter is added to the unit configuration information of the abnormal execution unit to disable the abnormal execution unit in the hardware detection process, and the hardware to be detected corresponding to the abnormal execution unit is marked as a fault state.

5. The method for handling hardware detection process exceptions according to claim 4, wherein: The method further comprises: When the hardware detection process meets the preset restart condition, obtaining the latest configuration file; Restarting the hardware detection process based on the latest configuration file; Wherein, in the latest configuration file, the unit configuration information of the abnormal execution unit includes a unit disabling parameter, so as to disable the abnormal execution unit during the process of restarting the hardware detection process.

6. The method for handling hardware detection process exceptions according to claim 5, wherein: The method further comprises: Obtaining a heartbeat signal sent by the hardware detection process according to a preset period; Determining whether a process crash occurs in the hardware detection process based on the acquisition of the heartbeat signal; When it is determined that the hardware detection process crashes, it is determined that the hardware detection process meets a preset restart condition.

7. The method for handling hardware detection process exceptions according to claim 5, wherein: The method further comprises: Obtaining hardware resource consumption information corresponding to the hardware detection process; In a case where the hardware resource consumption information indicates that the system hardware resources are exhausted, it is determined that the hardware detection process meets a preset restart condition, so as to release the hardware resources by restarting the hardware detection process.

8. A hardware detection process exception handling device, characterized in that: include: The acquisition module is used to obtain the configuration file of the hardware detection process; a splitting module, configured to split the hardware detection process into a plurality of independent execution units according to the plurality of detection functions supported by the hardware detection process as represented by the configuration file; wherein a single execution unit corresponds to at least one detection function; A monitoring module, configured to monitor the status of each execution unit during the execution of the hardware detection process and obtain a status monitoring result of each execution unit; The processing module is used to take the abnormal execution unit as the abnormal execution unit according to the status monitoring result, and perform corresponding abnormal processing operations on the abnormal execution unit.

9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the hardware detection process exception handling method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the hardware detection process exception handling method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Method, system and device for detecting dynamic scheduling of software process

    CN120653407A

  • A method, system and apparatus for detecting dynamic scheduling of software processes

    CN120653407B

  • Conflict processing method and device, storage medium and program product

    CN120780528A

  • Conflict processing method, device, storage medium, and program product

    CN120780528B

  • Operating system loading method and electronic equipment

    CN120892260A